数据来源:papers.cool/arxiv/cs.AI · 生成时间:2026/8/7 17:00:05
本报告聚合本周 arXiv cs.AI 的 Agent 相关论文,跨天去重后深度分析。
本周总览
本周去重后共 17 篇论文,覆盖 1 天数据。上周 12 篇,环比增加 5 篇。
研究方向分布
| 方向 | 本周 | 上周 | 变化 |
|---|---|---|---|
| 规划推理 | 7 | 2 | +5 |
| 其他 | 5 | 5 | 0 |
| 评估基准 | 4 | 2 | +2 |
| 记忆系统 | 4 | 0 | +4 |
| 安全对齐 | 1 | 1 | 0 |
| 工具使用 | 1 | 0 | +1 |
| 多智能体 | 1 | 2 | -1 |
应用场景分布
| 场景 | 论文数 | 占比 |
|---|---|---|
| 决策支持 | 2 | 12% |
| 信息检索与问答 | 2 | 12% |
| 代码开发 | 1 | 6% |
| 企业自动化 | 1 | 6% |
| 科学研究 | 1 | 6% |
核心论文解读
1. CoPlan: A Trustworthy Co-Intelligence Interface for Care Planning through Role-Based Contestable Argument Graphs
- arXiv: 2608.05107
- 方向: 规划推理 · 评估基准
- 场景: 决策支持
- 关键词:
coplancarecontestableplanningclinicaltrustworthyinterfacehumanargumentagents
2. ContextWeave: A Real-World Workflow Benchmark
- arXiv: 2608.04830
- 方向: 记忆系统 · 评估基准
- 场景: 企业自动化
- 关键词:
contextweavememoryrecallworkflowworkspacemisleadingworkflowstasksrelevance568
3. EviGraph: Evidence-Guided Autonomous Research Agents
- arXiv: 2608.04738
- 方向: 其他
- 场景: 科学研究、信息检索与问答
- 关键词:
evigraphresearchevidenceclaimautonomousmanuscriptsnanoresearchinconsistenciesagentsgraph
4. A/B Agent: A Self-Evolving Agent for Strategy Iteration in Industrial A/B Testing
- arXiv: 2608.04625
- 方向: 记忆系统 · 评估基准
- 场景: 决策支持
- 关键词:
strategyagentstrategiesrecommendationindustrialhistoricalragiterationexperiencetree
5. OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling
- arXiv: 2608.05141
- 方向: 其他
- 场景: 代码开发
- 关键词:
octolongcontextcodelongcontextslmsagentictokensrepositorymid
6. ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment
- arXiv: 2608.05102
- 方向: 其他
- 场景: 信息检索与问答
- 关键词:
backtrackedanswerabseekeragentsabccreditgrporewardsassignmentsearch
7. Item Response Theory for AI Safety
- arXiv: 2608.05086
- 方向: 安全对齐 · 评估基准
- 关键词:
irtbenchmarkssafetypsychometricitemsitemsandbagtoolkitsandbaggingpsychometrically
8. Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools
- arXiv: 2608.04719
- 方向: 记忆系统 · 规划推理
- 关键词:
canarycapabilitymiragestoolstooldecoyssusceptibilitycsrhostedprovider
9. Calibrating Artificial Guilt: Neurally Grounded Reward Shaping for Prosocial Multi-Agent Reinforcement Learning
- arXiv: 2608.04663
- 方向: 工具使用 · 多智能体
- 关键词:
guiltprosocialneurallyshapingsocialcalibratedagenthumanrewardartificial
10. Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning
- arXiv: 2608.05144
- 方向: 规划推理
- 关键词:
argusruntimeagenticbenchsweownedselfverificationrwkv6missions
研究趋势
主导方向:规划推理(7 篇),较上周(2 篇)上升。
上升: 规划推理(2→7)、评估基准(2→4)、记忆系统(0→4)、工具使用(0→1)
下降: 工程架构(2→0)、多智能体(2→1)
技术演进脉络
规划推理(7 篇)
- Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning
- CoPlan: A Trustworthy Co-Intelligence Interface for Care Planning through Role-Based Contestable Argument Graphs
- WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models
- 及另外 4 篇
其他(5 篇)
- OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling
- ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment
- EviGraph: Evidence-Guided Autonomous Research Agents
- 及另外 2 篇
评估基准(4 篇)
- CoPlan: A Trustworthy Co-Intelligence Interface for Care Planning through Role-Based Contestable Argument Graphs
- Item Response Theory for AI Safety
- ContextWeave: A Real-World Workflow Benchmark
- 及另外 1 篇
记忆系统(4 篇)
- Hierarchical Graph Memory for LLM Agents with Path-level Localization and Rewrite
- ContextWeave: A Real-World Workflow Benchmark
- Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools
- 及另外 1 篇
安全对齐(1 篇)
工具使用(1 篇)
多智能体(1 篇)
工程实践启示
- 工具使用方向 1 篇,function calling 与工具链持续演进。
- 记忆系统方向 4 篇,RAG 与长期记忆方案不断优化。
- 多智能体方向 1 篇,协作模式从简单分工走向复杂协调。
- 安全方向 1 篇,Agent 安全从外部围栏走向内化机制。
下周关注
持续热点:规划推理(本周 7 篇,上周 2 篇)、其他(本周 5 篇,上周 5 篇)、评估基准(本周 4 篇,上周 2 篇)
新出现方向:记忆系统(4 篇)、工具使用(1 篇)
附录:本周论文完整列表
去重后共 17 篇。
2026-08-06(17 篇)
- Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning — planning
- OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling — other
- CoPlan: A Trustworthy Co-Intelligence Interface for Care Planning through Role-Based Contestable Argument Graphs — planning, evaluation
- ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment — other
- Hierarchical Graph Memory for LLM Agents with Path-level Localization and Rewrite — memory
- Item Response Theory for AI Safety — safety, evaluation
- WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models — planning
- ContextWeave: A Real-World Workflow Benchmark — memory, evaluation
- Privileged, but Biased: How PI-Conditioned Teachers Break Self-Distillation — planning
- Fewer Tokens, Smaller Cache: Reward-Coordinated Efficient Reasoning — planning
- EviGraph: Evidence-Guided Autonomous Research Agents — other
- When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning — planning
- Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools — memory, planning
- Calibrating Artificial Guilt: Neurally Grounded Reward Shaping for Prosocial Multi-Agent Reinforcement Learning — tool, multi_agent
- A/B Agent: A Self-Evolving Agent for Strategy Iteration in Industrial A/B Testing — memory, evaluation
- Joint UAV Flight and Opportunistic Routing under Reinforcement Learning for Delay-Tolerant Networks — other
- What Is a Skill Worth? Structure-Aware Shapley Valuation of Agent Skills — other
本报告由 OpenClaw 自动生成,基于 agent-papers-research 每日数据聚合。