Friday’s cs.AI announcement day (covered Sat 2026-09-19) lists 89 new and 126 cross-lists (replacements skipped; listing total 215). Stack filter for agent systems, memory/context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent, persistence/identity, and local/open serving keeps 37 papers — FM OS/MCP layer, long-horizon agent architecture, coding-agent harnesses, tool-hallucination guards, SkillAA, stateful RAG, and GUI-agent nudge susceptibility.
Suparna Bhattacharya; Tarun Kumar; Cong Xu; Satish Kumar Mopur; … arXiv: 2609.19203
AI applications have shifted from single, monolithic foundation models (FM) to compound agentic systems. Yet today's stacks remain fragmented: even as protocols (e.g., MCP, A2A) ease tool/agent connectivity, each framework embeds an implicit runtime for state, memory, budgets, and guardrails, making behavior non-portable and governance brittle.
Key insight: Foundation models need a self-evolving OS layer for memory, MCP tools, and agent scheduling—not one monolithic model call.
Erik Nijkamp; Anurag Koul; Egor Pakhomov; Bo Pang arXiv: 2609.19519
Language-model agents are increasingly asked to carry out work spanning days or weeks, such as an operations remediation or a research programme. Such a task outlives any context window, any process and any interval at which a person can attend. In this paper, we argue that a long-horizon agent must run continually without forgetting before it can learn continually.
Key insight: Long-horizon agents can be structured as levels, ticks, and cascaded intelligence rather than a single flat loop.
Run-Ze Fan; Zihao Zhang; Simin Ma; Yebowen Hu; … arXiv: 2609.20804
Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the effectiveness of individual components unclear.
Key insight: Harness design choices (context management, tools, scaffolding) measurably change coding-agent outcomes.
Haozhe Liu; Tian Ye; Sensen Gao; Qihang Cao; … arXiv: 2609.20519
As coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands from isolated predictions into long trajectories of reasoning, tool use, and feedback. Token efficiency therefore becomes important for scaling recursive self-improvement.
Key insight: Recursive auto-research loops become efficient when the agent harness itself is scaled, not only the model.
Yukun Zhang; Kemu Xu; Yishen Chen arXiv: 2609.20474
Agent harnesses supply planning guidance, organize execution, and check completion. We study how these components affect success, erroneous acceptance, and cost in two Retail experiments and an Airline pilot in $τ^2$-bench.
Key insight: Harnesses create value through planning information and release control in stateful LLM agents.
Laxmipriya Ganesh Iyer arXiv: 2609.19425
Tool-augmented large language model (LLM) agents fail in a way no tool-selection or tool-security method addresses: they call tools that do not exist and pass arguments no schema declares. Existing defenses either pick the right tool (selection) or constrain what an agent may do with real tools (gating), both of which presuppose the emitted call refers to a real tool at all.
Key insight: Closed-world resolution and MCP-aware checks reduce tool hallucination in LLM agents.
Ziqiao Shang; Ling-Yue Ge; Lan-Zhe Guo arXiv: 2609.20455
External skills provide domain procedures without parameter updates, but existing methods often edit skills directly from failed rollouts without structured routing from an observed failure to an editable location; existing skill graphs also underuse semantic boundaries, object addresses, and topological dependencies for skill retrieval, targeted updating, and scoped validation
Key insight: Skill graphs need attribution-guided updates with targeted validation and rollback.
Yinzhu Quan; Zefang Liu arXiv: 2609.19523
Web agents often revisit the same sites, yet most evaluations discard the procedures learned in earlier successful interactions. We introduce EconSkills, a skill library and evaluation framework that distills verified EconWebArena trajectories into parameterized standard operating procedures for retrieving live economic data.
Key insight: Skill transfer and retrieval for web agents can be studied on live economic data streams.
Yanzhang Ma; Zhenghan Tai; Hanwei Wu; Sizhe Guan; … arXiv: 2609.19680
Financial QA systems are typically improved before deployment through better retrieval, prompting, or agent coordination, leaving their reliability behavior fixed thereafter. In practice, new SEC-filing questions repeatedly expose heterogeneous errors in period, entity, evidence use, and calculation.
Key insight: Self-evolving multi-agent skill ops can specialize for long-document QA workflows.
Wenjie Liao; Liangjie Zhao; Zehong Cao arXiv: 2609.20089
Self-evolving methods reduce the need for human-annotated trajectories by allowing tool-using agents to generate their own training data. Yet existing methods typically separate trajectory generation from evaluation, relying on static verifiers that cannot adapt to emerging failure modes or self-consistency signals that may reinforce errors shared across trajectories.
Key insight: Unified player training improves tool-integrated reasoning under agentic RL.
Mingxuan Zhang; Xiaowen Wang; Anupma Sharan; Zhengyi Chen; … arXiv: 2609.20754
Effective troubleshooting agents in enterprise customer support depend on retrieving actionable guidance from similar historical cases, yet existing retrieval-augmented generation (RAG) systems treat support cases as static documents and overlook their multi-stage, stateful nature.
Key insight: Stateful RAG should retrieve at intermediate case-timeline states, not only final documents.
Sangam Lee; Wonjae Lee; Sunghwan Kim; Deogyong Kim; … arXiv: 2609.19656
Information retrieval is increasingly important as LLM agents tackle complex tasks involving diverse information needs. Because retrieval relies on an index that represents each document through index keys, retrieval quality depends heavily on how effectively these keys expose the knowledge contained in each document.
Key insight: Search indexes for agents can self-evolve as usage patterns change.
Guangzhe Zhang arXiv: 2609.20045
A memory can answer a current query correctly while discarding distinctions required by a later update. We investigate this failure with a paired-history audit: two histories have the same current answer, receive a shared future update, and require different subsequent answers.
Key insight: Context-compression updates that look correct now can be insufficient later—audit update sufficiency.
Z. C. Luo; J. C. Guo; W. J. He; S. Y. Wang; … arXiv: 2609.20130
Recent memory-augmented repository-level program repair methods reuse historical repair experiences to improve LLM-based issue resolution. However, our analysis reveals three limitations in existing repository-level memory retrieval. First, episodic memory is highly imbalanced across repositories, leaving low-resource repositories with little effective support.
Key insight: Repository-level repair agents benefit from adaptive orchestration of historical repair experiences.
Xinghong Fu; Aravinth Kulanthaivelu; Yutaro Yamada arXiv: 2609.19526
Coding agents can recursively modify their own implementations, forming a loop of self-improvement. While prior work shows this can boost performance on coding benchmarks, existing approaches are costly and compute-intensive. We introduce a simple, sample-efficient self-improvement framework that significantly improves coding performance under strict budget constraints.
Key insight: Coding agents that self-modify can improve via fast tree-search rather than naive recursive edits.
Zhexi Feng; Ruiyi Zhang; Yongbo Yang; Pengtao Xie arXiv: 2609.20050
A coding agent halfway through an issue has already read much of what a retriever ranks highest. Relevance is scored per passage, but sufficiency belongs to the set: a ranker can fill its budget with variants of one required fact and leave the decision unsupported.
Key insight: Coding agents need state-conditioned minimal sufficient evidence of progress, not just logs.
Nicholas J. Conn arXiv: 2609.19607
Coding-agent benchmarks are built for broad and comprehensive comparisons, not frequent development decisions. Individual runs vary, full suites are expensive, and the benchmark harness may differ from the harness used in practice. In a resampling analysis of DeepSWE's published trials, only 19.5% of tasks (22 of 113) had a fifth-percentile Pearson correlation of at least 0.
Key insight: Affordable A/B testing (DeltaSelect) makes coding-agent harness comparisons practical.
Alex Remedios; Simon Storf; Fabien Roger; John Hughes arXiv: 2609.19587
To keep coding agents from going off the rails, production systems now review each proposed action with a blocking monitor that can reject it before it runs (Auto Mode in Claude Code, Guardian in OpenAI's Codex). Prior evaluations of such monitors largely measure robustness to accidental harm or prompt injections from untrusted sources looking to hijack the agent.
Key insight: Blocking classifiers against malign coding agents need red-teaming in Auto Mode settings.
Tisha Chawla; Susheem Koul arXiv: 2609.20625
Large language model responses are non-deterministic, so failures in LLM agents are hard to reproduce: a failure depends on inference that is not bitwise reproducible, on tools that read changing state, and on a multi-step trajectory that a re-run rarely repeats.
Key insight: Cut-point replay enables regression testing of LLM agents without full episode reruns.
Yishuo Yuan; Yibo Wu; Yihan Zhang; Minyuan Sun; … arXiv: 2609.19759
The rapid advancement of large language models and single-agent harnesses has reshaped the landscape of autonomous systems, raising a critical question of when multi-agent collaboration offers genuine value. As individual agent capabilities continue to scale, multi-agent collaboration faces diminishing returns while incurring growing context overhead.
Key insight: Adding agents can hurt: multi-agent collaboration needs interference-aware design.
Albert Wu; Nicholas Roberts; Tzu-Heng Huang; Haoran Lin; … arXiv: 2609.19391
LLM coding agents now generate complex programs at a scale that makes thorough human review increasingly difficult, raising the risk of safety and security failures. Common approaches, including fuzz testing, static analysis, and LLM-as-a-Verifier, can detect many failures but struggle to cover all possible edge cases.
Key insight: Multi-agent auto-formalization can harden safety guarantees on agentic outputs.
Yuejin Xie; Yu Li; Dadi Guo; Qingyu Liu; … arXiv: 2609.19892
As agent systems become more widely used, multiple agent sessions increasingly run alongside pre-existing user tasks in the same environment, sharing resources with limited capacity or mutually exclusive states.
Key insight: ClashBench stresses agents in conflict settings that induce seize-and-harm behaviors.
Wonmi Choi; Minuk Park; Zhixiong Niu; Yongqiang Xiong; … arXiv: 2609.19947
LLM-based AI agents process user requests through iterative reasoning and tool execution, often involving the invocation of remote LLM APIs with local tool containers. This execution model can make the optimization of agent serving difficult because latency, local resource demand, and container bottlenecks inter-mix across requests.
Key insight: Agent latency bottlenecks are task-dependent (CPU, disk, memory); faster LLMs do not always help.
Yusheng Zheng; Chaokun Chang; Yu Mao; Tianyuan Wu; … arXiv: 2609.20301
AI agents increasingly orchestrate long-running activities with users, tools, and system resources for days and weeks. To improve agent quality, safety, and cost efficiency, developers need to determine where failures happen, what triggers unsafe effects, and which tasks consume the most budget, then optimize those tasks.
Key insight: Semantic profilers expose where long-horizon agents spend time and tokens.
Jun He; Deying Yu arXiv: 2609.20261
Autonomous agents derive concrete mutations from database reads, retrieved evidence, policy, beliefs, and delegated authority. Those inputs may change while reasoning is in progress. Database isolation orders the submitted transaction; agentic transaction processing determines whether a proposal satisfies an executable contract.
Key insight: When agents commit, cognitive serializability must span data, evidence, policy, and authority.
Song Zhang; Jiankang Yao; Hongtao Li; Xiaojun Zhang; … arXiv: 2609.20095
The Internet of Agents is expected to enable large numbers of autonomous agents to discover, verify, and collaborate with each other across heterogeneous platforms. However, current agent protocols mainly address tool invocation and inter-agent communication, leaving scalable agent registration, trustworthy identification, and capability-oriented discovery largely unresolved.
Key insight: Internet-of-Agents trust requires scalable registration, identity, and capability discovery.
Haya Halimeh; Sascha Kaltenpoth; Kevin Bösch; Oliver Müller arXiv: 2609.19843
LLM-based GUI agents increasingly act on behalf of users in digital environments that were designed with human users in mind. These graphical user interfaces were designed to support, but also deliberately steer, the behaviour and decisions of users.
Key insight: LLM GUI agents are susceptible to interface nudges designed for humans.
Mahsa Amani; Seungeon Lee; Abhisek Dash; Asmaa El Fraihi; … arXiv: 2609.19244
Conversational LLM agents increasingly rely on Web search, yet the end-to-end lifecycle of agentic search remains poorly understood.
Key insight: Conversational LLM agents follow distinctive web-search decision and strategy patterns.
Ambika Sharan; Grigory Chirkov; Soheil Abbasloo arXiv: 2609.19387
Agents are increasingly asked to design hardware, and increasingly reported to succeed. Such reports establish that a design improved; they cannot establish why. An agent that improves an accelerator may be reasoning about the machine, or may be searching competently over knobs whose meaning it never recovers -- and only the first transfers to the next architecture.
Key insight: AI agents often lack computer-architecture literacy needed for systems-level tool use.
Run Peng; Zinnia Nie; Jing Ding; Yinpei Dai; … arXiv: 2609.19610
Understanding humans over long horizons requires agents to infer not only what people need in the moment, but also how routines form, why they repeat, and when they change. We introduce SimLife, a scalable platform for simulating long-term household life with rich visual observations, ground-truth action logs, and synthetic dialogues with audio.
Key insight: Long-horizon human–agent partnership depends on shared pattern understanding over days.
Shambhavi Mishra; David Vazquez; Perouz Taslakian; Marco Pedersoli; … arXiv: 2609.19551
In an enterprise system, updating one field can set another, create a record, or start an approval. These effects are produced by business rules that are not built into the platform but written by each organization and revised over time. An agent working in such a system cannot predict the result of its own actions without knowing these rules.
Key insight: Continual world-model discovery must revise, extend, and retire rules as environments shift.
Xiaofei Yuan; Yan Zhang; Shaobo Qiao; Huangleshuai He; … arXiv: 2609.19654
Travel-planning agents generate itineraries that may become infeasible after acceptance because of flight cancellations, hotel unavailability, or attraction closures. Revising these itineraries involves full replanning, classical plan repair, and LLM-based travel-agent revision, whose differing task formulations and evaluation protocols hinder comparison.
Key insight: Travel (and personal) agents under preference change need a replan/repair/edit taxonomy.
Donghan Bian; Marie Puren; Florian Cafiero arXiv: 2609.19897
Historical archives pose a difficult retrieval problem for retrievalaugmented generation systems: documents are OCR-degraded, heterogeneous across genres and sources, and require strong source traceability for scholarly and institutional use. We introduce TRACE, a training-free agentic retrieval framework designed for accountable source discovery over historical corpora.
Key insight: Accountable agentic retrieval should preserve provenance for archive source discovery.
Pritish Mishra; Ishaan Kumar; Akshat Mandoli; Sudarshan Kamath arXiv: 2609.20152
Generally, most voice agents are cascaded systems, i.e., an ASR model transcribes the caller's audio, a language model reads the transcript and decides what to say and which backend tools to call, and a TTS model speaks the reply. Nearly all of the decision making happens in the language model, but existing evaluations measure it either too broadly or too narrowly.
Key insight: Cascaded voice agents should evaluate the language-model+tool core separately from ASR/TTS.
Nolan Smyth; Yorguin-Jose Mantilla-Ramos; Pascal Jr Tikeng Notsawo; Saskia Helbling; … arXiv: 2609.20812
Frontier coding agents are increasingly trusted to work autonomously for long periods, yet an agent's final response is often the only account of that work a user sees. We quantify the propensity of frontier agents to \emph{overclaim} task completion, a misrepresentation that can mislead the user.
Key insight: Frontier LLM agents show measurable overclaiming propensity under harness evaluation.
Nitish Dashora; Douglas Chen; Idan Shenfeld; John Marangola; … arXiv: 2609.20820
Complex robotic manipulation tasks frequently require a long-term memory of past events and actions. As conditioning on full histories renders policies prone to spurious correlations and degrades performance, many approaches to policy memory involve compressing historical information through expensive VLM queries in-the-loop to process only task-salient information.
Key insight: Lightweight workspace memory via saliency supervision can persist robotic (and agent) state cheaply.
Xuan Liu; Jingbin Qian arXiv: 2609.19636
Reinforcement learning now trains language-model agents that act over dozens of steps in live environments. The gains are large, and they are read as better decision-making. An agent in a closed loop writes its own inputs. Each observation follows from its own earlier actions, so the states it meets late in an episode are partly of its own making.
Key insight: Agentic RL gains should be attributed via checkpoint handoffs—reachability is not the same as solving.