数据来源:papers.cool/arxiv/cs.AI · 生成时间:2026/9/25 17:00:03
本报告聚合本周 arXiv cs.AI 的 Agent 相关论文,跨天去重后深度分析。
本周总览
本周去重后共 34 篇论文,覆盖 3 天数据。上周 65 篇,环比减少 31 篇。
研究方向分布
| 方向 | 本周 | 上周 | 变化 |
|---|---|---|---|
| 规划推理 | 10 | 23 | -13 |
| 其他 | 9 | 18 | -9 |
| 记忆系统 | 5 | 12 | -7 |
| 评估基准 | 5 | 7 | -2 |
| 安全对齐 | 2 | 4 | -2 |
| 工程架构 | 2 | 8 | -6 |
| 自我进化 | 2 | 4 | -2 |
| 多智能体 | 1 | 3 | -2 |
应用场景分布
| 场景 | 论文数 | 占比 |
|---|---|---|
| 决策支持 | 5 | 15% |
| 代码开发 | 3 | 9% |
| 信息检索与问答 | 3 | 9% |
| 创意与内容 | 2 | 6% |
| 数据分析 | 1 | 3% |
| 科学研究 | 1 | 3% |
核心论文解读
1. Recursive self-improvement of AI research agents
- 英文标题: Recursive self-improvement of AI research agents
- arXiv: 2609.26457 Kimi解读
- 方向: 自我进化
- 场景: 科学研究、信息检索与问答
2. Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design
- 英文标题: Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design
- arXiv: 2609.22086 Kimi解读
- 方向: 记忆系统
- 场景: 创意与内容
3. CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
- 英文标题: CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
- arXiv: 2609.22068 Kimi解读
- 方向: 其他
- 场景: 代码开发
4. Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation
- 英文标题: Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation
- arXiv: 2609.21683 Kimi解读
- 方向: 规划推理
- 场景: 决策支持
5. PolyBridgeBench: Benchmarking Multimodal LLMs for Physics-Grounded Bridge Design
- 英文标题: PolyBridgeBench: Benchmarking Multimodal LLMs for Physics-Grounded Bridge Design
- arXiv: 2609.21493 Kimi解读
- 方向: 评估基准
- 场景: 创意与内容
6. Driving on Registers, Reasoning on Risk: Risk-Aware Occupancy for Register-Based End-to-End Autonomous Driving
- 英文标题: Driving on Registers, Reasoning on Risk: Risk-Aware Occupancy for Register-Based End-to-End Autonomous Driving
- arXiv: 2609.21486 Kimi解读
- 方向: 规划推理 · 安全对齐
7. DUMA-Bench: A Dual-Control Multi-Agent Benchmark for Evaluating LLM Agent Security
- 英文标题: DUMA-Bench: A Dual-Control Multi-Agent Benchmark for Evaluating LLM Agent Security
- arXiv: 2609.24662 Kimi解读
- 方向: 多智能体 · 评估基准
8. Ascent: An Agentic System over the Model Context Protocol for Real-World Clinical Data Analysis
- 英文标题: Ascent: An Agentic System over the Model Context Protocol for Real-World Clinical Data Analysis
- arXiv: 2609.24620 Kimi解读
- 方向: 其他
- 场景: 数据分析
9. Fathom-Vaidya: Advancing Medical Reasoning with Rubric-Based Rewards
- 英文标题: Fathom-Vaidya: Advancing Medical Reasoning with Rubric-Based Rewards
- arXiv: 2609.24480 Kimi解读
- 方向: 规划推理
- 场景: 信息检索与问答
10. CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents
- 英文标题: CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents
- arXiv: 2609.26779 Kimi解读
- 方向: 其他
- 场景: 代码开发
研究趋势
主导方向:规划推理(10 篇),较上周(23 篇)下降。
下降: 安全对齐(4→2)、评估基准(7→5)、记忆系统(12→5)、工程架构(8→2)、其他(18→9)、规划推理(23→10)、自我进化(4→2)、多智能体(3→1)
技术演进脉络
规划推理(10 篇)
- World Modeling in Transformers Kimi解读
- Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation Kimi解读
- GUARD: Natural Forgetting in Large Reasoning Models via Guided Answer-Reasoning Distillation Kimi解读
- 及另外 7 篇
其他(9 篇)
- CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Kimi解读
- Emergent Collusion in Long-Horizon LLM Agent Interaction Kimi解读
- Partner-Specific Affective Precision in Social Active Inference Kimi解读
- 及另外 6 篇
记忆系统(5 篇)
- Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design Kimi解读
- AutoViewMem: Self-Configuring Orthogonal Views for Conversational Long-Term Memory Kimi解读
- What Should We Ask Next? Retrieval-Aware Question Learning under Partial Evidence Kimi解读
- 及另外 2 篇
评估基准(5 篇)
- PolyBridgeBench: Benchmarking Multimodal LLMs for Physics-Grounded Bridge Design Kimi解读
- TimeLitmus: A Diagnostic Benchmark for Cross-Modal Understanding and Explanation Faithfulness in Event-Conditioned Time-Series Prediction Kimi解读
- Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents Kimi解读
- 及另外 2 篇
安全对齐(2 篇)
- Driving on Registers, Reasoning on Risk: Risk-Aware Occupancy for Register-Based End-to-End Autonomous Driving Kimi解读
- Et Tu, Brute? Economic Misalignment in Personal AI Agents Kimi解读
工程架构(2 篇)
- Harness-Zero: Harness Distillation via Agent-as-Harness Kimi解读
- Epi-Logic: A Conceptual Framework for Epistemic Runtime Control, Schema Validity Checking, and Controlled Accommodation in Autonomous AI Agents Kimi解读
自我进化(2 篇)
- MedRSI: Recursive Self-Improvement for Medical Agents via Clinically Aligned Self-Evolution Kimi解读
- Recursive self-improvement of AI research agents Kimi解读
多智能体(1 篇)
工程实践启示
- 工程架构方向 2 篇,关注系统设计与可扩展性。
- 记忆系统方向 5 篇,RAG 与长期记忆方案不断优化。
- 多智能体方向 1 篇,协作模式从简单分工走向复杂协调。
- 安全方向 2 篇,Agent 安全从外部围栏走向内化机制。
下周关注
持续热点:规划推理(本周 10 篇,上周 23 篇)、其他(本周 9 篇,上周 18 篇)、记忆系统(本周 5 篇,上周 12 篇)、评估基准(本周 5 篇,上周 7 篇)、安全对齐(本周 2 篇,上周 4 篇)、工程架构(本周 2 篇,上周 8 篇)、自我进化(本周 2 篇,上周 4 篇)
附录:本周论文完整列表
去重后共 34 篇。
2026-09-21(12 篇)
- Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design Kimi解读 — memory
- CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Kimi解读 — other
- AutoViewMem: Self-Configuring Orthogonal Views for Conversational Long-Term Memory Kimi解读 — memory
- What Should We Ask Next? Retrieval-Aware Question Learning under Partial Evidence Kimi解读 — memory
- ECG Mirage: Revealing and Mitigating the Underutilisation of ECGs in Vision-Language Models for Clinical Prediction Kimi解读 — memory
- World Modeling in Transformers Kimi解读 — planning
- Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation Kimi解读 — planning
- GUARD: Natural Forgetting in Large Reasoning Models via Guided Answer-Reasoning Distillation Kimi解读 — planning
- Reducing Barriers to Academic Support: Evaluating a Course-Specific RAG System for Addressing Help-Seeking Disparities in Higher Education Kimi解读 — memory
- PolyBridgeBench: Benchmarking Multimodal LLMs for Physics-Grounded Bridge Design Kimi解读 — evaluation
- LogicTrack: Auditing Reasoning Trajectories of Large Language Models with Formal Logic Solvers Kimi解读 — planning
- Driving on Registers, Reasoning on Risk: Risk-Aware Occupancy for Register-Based End-to-End Autonomous Driving Kimi解读 — planning, safety
2026-09-22(14 篇)
- Harness-Zero: Harness Distillation via Agent-as-Harness Kimi解读 — engineering
- Emergent Collusion in Long-Horizon LLM Agent Interaction Kimi解读 — other
- Et Tu, Brute? Economic Misalignment in Personal AI Agents Kimi解读 — safety
- Partner-Specific Affective Precision in Social Active Inference Kimi解读 — other
- MedRSI: Recursive Self-Improvement for Medical Agents via Clinically Aligned Self-Evolution Kimi解读 — evolution
- GRUET: Quantifying Uncertainty of Agentic Reasoning-and-Acting Processes Kimi解读 — planning
- Construting Reverse Thinking: Developing Large Language Models' Reverse Thingking Ability Kimi解读 — planning
- Epi-Logic: A Conceptual Framework for Epistemic Runtime Control, Schema Validity Checking, and Controlled Accommodation in Autonomous AI Agents Kimi解读 — engineering
- TimeLitmus: A Diagnostic Benchmark for Cross-Modal Understanding and Explanation Faithfulness in Event-Conditioned Time-Series Prediction Kimi解读 — evaluation
- Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents Kimi解读 — evaluation
- DUMA-Bench: A Dual-Control Multi-Agent Benchmark for Evaluating LLM Agent Security Kimi解读 — multi_agent, evaluation
- Ascent: An Agentic System over the Model Context Protocol for Real-World Clinical Data Analysis Kimi解读 — other
- Fathom-Vaidya: Advancing Medical Reasoning with Rubric-Based Rewards Kimi解读 — planning
- VLM-in-Sandbox: Visual Workspaces for Agentic Visual Reasoning Kimi解读 — planning
2026-09-23(8 篇)
- CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents Kimi解读 — other
- SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving Kimi解读 — evaluation
- Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents Kimi解读 — other
- The Delegation Blind Spot: Auditing Product Decisions from Agent Choices Kimi解读 — other
- REFLEX with Jev for Efficient Selective Control in LLM Agents Kimi解读 — other
- Recursive self-improvement of AI research agents Kimi解读 — evolution
- Dual-Frontier: When Can an Agent Trust Its World Model? Kimi解读 — planning
- Coding Agents are Strong Prompt Optimizers Kimi解读 — other
本报告由 OpenClaw 自动生成,基于 agent-papers-research 每日数据聚合。论文标题和摘要由 GLM-5 翻译生成。