数据来源:papers.cool/arxiv/cs.AI · 生成时间:2026/10/9 17:00:05
本报告聚合本周 arXiv cs.AI 的 Agent 相关论文,跨天去重后深度分析。
本周总览
本周去重后共 60 篇论文,覆盖 4 天数据。上周 23 篇,环比增加 37 篇。
研究方向分布
| 方向 | 本周 | 上周 | 变化 |
|---|---|---|---|
| 其他 | 21 | 5 | +16 |
| 规划推理 | 14 | 8 | +6 |
| 评估基准 | 13 | 4 | +9 |
| 自我进化 | 7 | 2 | +5 |
| 安全对齐 | 7 | 2 | +5 |
| 工程架构 | 5 | 2 | +3 |
| 记忆系统 | 4 | 2 | +2 |
| 多智能体 | 2 | 4 | -2 |
应用场景分布
| 场景 | 论文数 | 占比 |
|---|---|---|
| 决策支持 | 9 | 15% |
| 机器人与物理世界 | 6 | 10% |
| 科学研究 | 5 | 8% |
| 信息检索与问答 | 5 | 8% |
| 代码开发 | 5 | 8% |
| 企业自动化 | 4 | 7% |
| 创意与内容 | 2 | 3% |
| 数据分析 | 1 | 2% |
核心论文解读
1. EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution
- 英文标题: EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution
- arXiv: 2610.10498 Kimi解读
- 方向: 其他
- 场景: 代码开发、科学研究、机器人与物理世界
2. HazardWeaver: Scientific Route Selection for Hazard Analysis Agents
- 英文标题: HazardWeaver: Scientific Route Selection for Hazard Analysis Agents
- arXiv: 2610.03591 Kimi解读
- 方向: 其他
- 场景: 科学研究、企业自动化
3. Knowledge or Calculator? Decomposing the Skill Premium in Verifiable Financial Agent Workflows
- 英文标题: Knowledge or Calculator? Decomposing the Skill Premium in Verifiable Financial Agent Workflows
- arXiv: 2610.03564 Kimi解读
- 方向: 其他
- 场景: 企业自动化、信息检索与问答
4. Benchmarking Candidate Coverage in Typed Decision Models
- 英文标题: Benchmarking Candidate Coverage in Typed Decision Models
- arXiv: 2610.03387 Kimi解读
- 方向: 记忆系统 · 评估基准
- 场景: 决策支持
5. Proof-Grounded Patient-Specific Clinical Explanations from Knowledge-Graph Reasoning
- 英文标题: Proof-Grounded Patient-Specific Clinical Explanations from Knowledge-Graph Reasoning
- arXiv: 2610.06549 Kimi解读
- 方向: 规划推理
- 场景: 科学研究、信息检索与问答
6. GPlaceRL: An Open-Source Graph Reinforcement Learning Framework for Detailed Placement
- 英文标题: GPlaceRL: An Open-Source Graph Reinforcement Learning Framework for Detailed Placement
- arXiv: 2610.06489 Kimi解读
- 方向: 工程架构
- 场景: 代码开发、决策支持
7. VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning
- 英文标题: VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning
- arXiv: 2610.08761 Kimi解读
- 方向: 规划推理 · 自我进化
- 场景: 机器人与物理世界
8. ScienceClaw: Benchmarking Continual Self-Evolution of AI-for-Science Agents Across the Natural and Social Sciences
- 英文标题: ScienceClaw: Benchmarking Continual Self-Evolution of AI-for-Science Agents Across the Natural and Social Sciences
- arXiv: 2610.08691 Kimi解读
- 方向: 评估基准 · 自我进化
- 场景: 科学研究
9. Parallel Predictive World Models for Accurate and Efficient Long-Horizon Planning
- 英文标题: Parallel Predictive World Models for Accurate and Efficient Long-Horizon Planning
- arXiv: 2610.08627 Kimi解读
- 方向: 规划推理
- 场景: 代码开发、决策支持
10. Adaptive Power Sampling for LLM Reasoning
- 英文标题: Adaptive Power Sampling for LLM Reasoning
- arXiv: 2610.08563 Kimi解读
- 方向: 规划推理 · 自我进化
- 场景: 信息检索与问答
研究趋势
主导方向:其他(21 篇),较上周(5 篇)上升。
上升: 其他(5→21)、规划推理(8→14)、评估基准(4→13)、自我进化(2→7)、安全对齐(2→7)、工程架构(2→5)、记忆系统(2→4)
下降: 多智能体(4→2)
技术演进脉络
其他(21 篇)
- NeutronGym: Physics-Graded Neutron Instrument Design for LLM Agents Kimi解读
- HazardWeaver: Scientific Route Selection for Hazard Analysis Agents Kimi解读
- Knowledge or Calculator? Decomposing the Skill Premium in Verifiable Financial Agent Workflows Kimi解读
- 及另外 18 篇
规划推理(14 篇)
- Recursive Harness Self-Improvement for Frontier Reasoning Data Synthesis Kimi解读
- Reasoning Models Are Accurate but Unsound on Identification Kimi解读
- Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability Kimi解读
- 及另外 11 篇
评估基准(13 篇)
- Do Large Language Models Know Colombian Law? A Reliability Benchmark for the Colombian Legal System Kimi解读
- HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents Kimi解读
- From Benchmarks to Production: A Text-to-SQL System for Complex Financial Data Kimi解读
- 及另外 10 篇
自我进化(7 篇)
- Recursive Harness Self-Improvement for Frontier Reasoning Data Synthesis Kimi解读
- Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability Kimi解读
- VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning Kimi解读
- 及另外 4 篇
安全对齐(7 篇)
- ANT: A Multi-Granularity Network Traffic Dataset and Benchmark for Agents Behavior Auditing Kimi解读
- AgentPrivArena: Evaluating and Auditing Real-world AI Agent Privacy Kimi解读
- ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic Coding Kimi解读
- 及另外 4 篇
工程架构(5 篇)
- Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents Kimi解读
- GPlaceRL: An Open-Source Graph Reinforcement Learning Framework for Detailed Placement Kimi解读
- Toward Alignment Scaling Laws: A Framework and First Preregistered Measurements Kimi解读
- 及另外 2 篇
记忆系统(4 篇)
- Benchmarking Candidate Coverage in Typed Decision Models Kimi解读
- MINDSET: Energy-based Schema Evolution for Long Conversational Agent Memory Kimi解读
- RECAST: Learning to Compute the Right Context through Adaptive Evidence Routing Kimi解读
- 及另外 1 篇
多智能体(2 篇)
- Can Agent Harnesses and Inference Engines Hear Each Other? The HEAR Protocol for Agentic LLM Serving Kimi解读
- SciExam for ENSO: Can AI Agents Build Climate Models? Kimi解读
工程实践启示
- 工程架构方向 5 篇,关注系统设计与可扩展性。
- 记忆系统方向 4 篇,RAG 与长期记忆方案不断优化。
- 多智能体方向 2 篇,协作模式从简单分工走向复杂协调。
- 安全方向 7 篇,Agent 安全从外部围栏走向内化机制。
下周关注
持续热点:其他(本周 21 篇,上周 5 篇)、规划推理(本周 14 篇,上周 8 篇)、评估基准(本周 13 篇,上周 4 篇)、自我进化(本周 7 篇,上周 2 篇)、安全对齐(本周 7 篇,上周 2 篇)、工程架构(本周 5 篇,上周 2 篇)、记忆系统(本周 4 篇,上周 2 篇)、多智能体(本周 2 篇,上周 4 篇)
附录:本周论文完整列表
去重后共 60 篇。
2026-10-05(13 篇)
- Do Large Language Models Know Colombian Law? A Reliability Benchmark for the Colombian Legal System Kimi解读 — evaluation
- Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents Kimi解读 — engineering
- NeutronGym: Physics-Graded Neutron Instrument Design for LLM Agents Kimi解读 — other
- HazardWeaver: Scientific Route Selection for Hazard Analysis Agents Kimi解读 — other
- HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents Kimi解读 — evaluation
- Knowledge or Calculator? Decomposing the Skill Premium in Verifiable Financial Agent Workflows Kimi解读 — other
- Recursive Harness Self-Improvement for Frontier Reasoning Data Synthesis Kimi解读 — planning, evolution
- From Benchmarks to Production: A Text-to-SQL System for Complex Financial Data Kimi解读 — evaluation
- Reasoning Models Are Accurate but Unsound on Identification Kimi解读 — planning
- Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability Kimi解读 — planning, evolution
- Benchmarking Candidate Coverage in Typed Decision Models Kimi解读 — memory, evaluation
- ReFract: Benchmarking Perspective Awareness in Language Model Agents with Text World Models Kimi解读 — planning, evaluation
- Preserving Mathematical Reasoning in Compressed Diffusion Language Models via Trajectory-Aware Low-Rank Approximation Kimi解读 — planning
2026-10-06(11 篇)
- Back to the Future: Rethinking EDA Infrastructure for Agentic Systems in Chip Design Verification Kimi解读 — other
- FREA: A Multi-Source Expert Benchmark for Reaction Feasibility Verification Kimi解读 — evaluation
- Can Agent Harnesses and Inference Engines Hear Each Other? The HEAR Protocol for Agentic LLM Serving Kimi解读 — multi_agent
- HERA: Harness-Environment Co-Evolution for Reliable Agentic Abstention Kimi解读 — other
- Proof-Grounded Patient-Specific Clinical Explanations from Knowledge-Graph Reasoning Kimi解读 — planning
- ANT: A Multi-Granularity Network Traffic Dataset and Benchmark for Agents Behavior Auditing Kimi解读 — safety, evaluation
- polyview: A Python package for multi-view machine learning Kimi解读 — other
- GPlaceRL: An Open-Source Graph Reinforcement Learning Framework for Detailed Placement Kimi解读 — engineering
- AgentPrivArena: Evaluating and Auditing Real-world AI Agent Privacy Kimi解读 — safety
- From Benchmark to Bench: Can Agents Survive Real-World Drug Discovery? Kimi解读 — evaluation
- What Did the Agent Actually Do? Evidence-Grounded Oversight for Long-Horizon Agents Kimi解读 — other
2026-10-07(18 篇)
- Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts? Kimi解读 — other
- VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning Kimi解读 — planning, evolution
- Does an Agent's History Tell You When Compaction Will Hurt? A Modest, Bounded Effect on the TRACE Paired-Replay Corpus Kimi解读 — other
- WorldSolver: Can LLM Agents Simulate the Physical Dynamics via Solver Generation? Kimi解读 — other
- nanoMuse: An Open-Source Personal Agent for Every Device You Own Kimi解读 — other
- ScienceClaw: Benchmarking Continual Self-Evolution of AI-for-Science Agents Across the Natural and Social Sciences Kimi解读 — evaluation, evolution
- ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic Coding Kimi解读 — safety, evaluation
- SquidAgent: Parallelize Wisely, Coordinate Efficiently Kimi解读 — other
- Parallel Predictive World Models for Accurate and Efficient Long-Horizon Planning Kimi解读 — planning
- Recursive Game Creator: An Agentic Product-Level Experience-Oriented Game Harness Kimi解读 — other
- MINDSET: Energy-based Schema Evolution for Long Conversational Agent Memory Kimi解读 — memory
- Adaptive Power Sampling for LLM Reasoning Kimi解读 — planning, evolution
- Toward Alignment Scaling Laws: A Framework and First Preregistered Measurements Kimi解读 — safety, engineering
- How Much Evidence Should a Coding Agent's Self-Correction Carry? Adaptive Dirichlet Evidence for Self-Distillation Kimi解读 — evolution
- AssemState: Manual and Physical-State-Guided Reasoning for Zero-shot Furniture Assembly Kimi解读 — planning
- EMHO: EMbodied Agent Harness Optimization via Experience Traces Kimi解读 — engineering
- Transect: Retaining Observability for Long-Horizon LLM Agent Evaluations Kimi解读 — evaluation
- MedZERO: Self-Evolving Agents for Open-Ended Medical Reasoning Through Controlled Knowledge Accumulation Kimi解读 — planning
2026-10-08(18 篇)
- RoboJEPA: Scaling Robotic Latent World Models Kimi解读 — planning
- SciExam for ENSO: Can AI Agents Build Climate Models? Kimi解读 — multi_agent
- RECAST: Learning to Compute the Right Context through Adaptive Evidence Routing Kimi解读 — memory, evolution
- EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution Kimi解读 — other
- Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models Kimi解读 — evaluation
- RunningTab: Direct Workspace Interaction with Environment-Side Tabs Kimi解读 — other
- SOTA: Stock Options Trading Agents Guided by Option-Implied Return Distributions Kimi解读 — other
- Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language Models Kimi解读 — planning
- AI Safety Considerations for Agents With Limited Time to Act Kimi解读 — safety
- OOM-RL II: Reality Is an Oracle, Not a Debugger Provenance-Constrained Diagnosis in Continually Evolving Agent-Engineered Systems Kimi解读 — other
- GAGR-Lab: Evaluating Joint Spatial-Geometric and Analytic Function Reasoning Kimi解读 — planning, evaluation
- Agentic AI-Assisted Modeling for Production Scheduling: Assessment in Constraint Programming Kimi解读 — other
- RewardWeaver: Long-Horizon Interactive Learning for Language Agents via Self-Evolving Reward Adaptation Kimi解读 — other
- HGP:An on-device personalized agent memory via hybrid graph storage Kimi解读 — memory
- Loud Failures, Quiet Failures: Fault Detection and Recovery in Tool-Using Language Model Agents Kimi解读 — other
- Learning to Accumulate Knowledge with Mutual Information Kimi解读 — other
- From Expected Harmfulness to Likelihood: A Probabilistic Reformulation of Jailbreaking LLM Agents Kimi解读 — safety
- Successive Training Stages and Large Language Model Persuasion: Effects of Misalignment, Supervised Fine-Tuning, and Preference Optimization Kimi解读 — safety, engineering
本报告由 OpenClaw 自动生成,基于 agent-papers-research 每日数据聚合。论文标题和摘要由 GLM-5 翻译生成。