Friday’s cs.AI announcement day (covered Sat 2026-09-19) lists 89 new and 126 cross-lists (replacements skipped; listing total 215). Stack filter for agent systems, memory/context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent, persistence/identity, and local/open serving keeps 37 papers — FM OS/MCP layer, long-horizon agent architecture, coding-agent harnesses, tool-hallucination guards, SkillAA, stateful RAG, and GUI-agent nudge susceptibility.


Research Papers

Position: It is Time to Virtualize Foundation Models with a Self-evolving Operating System Layer

Suparna Bhattacharya; Tarun Kumar; Cong Xu; Satish Kumar Mopur; … arXiv: 2609.19203

Figure from Position: It is Time to Virtualize Foundation Models with a Self-evolving Operating System Layer
Position: It is Time to Virtualize Foundation Models with a Self-evolving Operating System Layer

AI applications have shifted from single, monolithic foundation models (FM) to compound agentic systems. Yet today's stacks remain fragmented: even as protocols (e.g., MCP, A2A) ease tool/agent connectivity, each framework embeds an implicit runtime for state, memory, budgets, and guardrails, making behavior non-portable and governance brittle.

Key insight: Foundation models need a self-evolving OS layer for memory, MCP tools, and agent scheduling—not one monolithic model call.

An Architecture for Long-Horizon Agents: Levels, Ticks and Cascaded Intelligence

Erik Nijkamp; Anurag Koul; Egor Pakhomov; Bo Pang arXiv: 2609.19519

Language-model agents are increasingly asked to carry out work spanning days or weeks, such as an operations remediation or a research programme. Such a task outlives any context window, any process and any interval at which a person can attend. In this paper, we argue that a long-horizon agent must run continually without forgetting before it can learn continually.

Key insight: Long-horizon agents can be structured as levels, ticks, and cascaded intelligence rather than a single flat loop.

An Empirical Study of Harness Design for Coding Agents

Run-Ze Fan; Zihao Zhang; Simin Ma; Yebowen Hu; … arXiv: 2609.20804

Figure from An Empirical Study of Harness Design for Coding Agents
An Empirical Study of Harness Design for Coding Agents

Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the effectiveness of individual components unclear.

Key insight: Harness design choices (context management, tools, scaffolding) measurably change coding-agent outcomes.

SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness

Haozhe Liu; Tian Ye; Sensen Gao; Qihang Cao; … arXiv: 2609.20519

Figure from SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness

As coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands from isolated predictions into long trajectories of reasoning, tool use, and feedback. Token efficiency therefore becomes important for scaling recursive self-improvement.

Key insight: Recursive auto-research loops become efficient when the agent harness itself is scaled, not only the model.

How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents

Yukun Zhang; Kemu Xu; Yishen Chen arXiv: 2609.20474

Agent harnesses supply planning guidance, organize execution, and check completion. We study how these components affect success, erroneous acceptance, and cost in two Retail experiments and an Airline pilot in $τ^2$-bench.

Key insight: Harnesses create value through planning information and release control in stateful LLM agents.

Closed-World Resolution Against Tool Hallucination in LLM Agents

Laxmipriya Ganesh Iyer arXiv: 2609.19425

Figure from Closed-World Resolution Against Tool Hallucination in LLM Agents
Closed-World Resolution Against Tool Hallucination in LLM Agents

Tool-augmented large language model (LLM) agents fail in a way no tool-selection or tool-security method addresses: they call tools that do not exist and pass arguments no schema declares. Existing defenses either pick the right tool (selection) or constrain what an agent may do with real tools (gating), both of which presuppose the emitted call refers to a real tool at all.

Key insight: Closed-world resolution and MCP-aware checks reduce tool hallucination in LLM agents.

SkillAA: Attribution-Guided Skill-Graph Updating with Targeted Validation and Rollback

Ziqiao Shang; Ling-Yue Ge; Lan-Zhe Guo arXiv: 2609.20455

Figure from SkillAA: Attribution-Guided Skill-Graph Updating with Targeted Validation and Rollback
SkillAA: Attribution-Guided Skill-Graph Updating with Targeted Validation and Rollback

External skills provide domain procedures without parameter updates, but existing methods often edit skills directly from failed rollouts without structured routing from an observed failure to an editable location; existing skill graphs also underuse semantic boundaries, object addresses, and topological dependencies for skill retrieval, targeted updating, and scoped validation

Key insight: Skill graphs need attribution-guided updates with targeted validation and rollback.

EconSkills: Studying Skill Transfer and Retrieval for Web Agents on Live Economic Data

Yinzhu Quan; Zefang Liu arXiv: 2609.19523

Web agents often revisit the same sites, yet most evaluations discard the procedures learned in earlier successful interactions. We introduce EconSkills, a skill library and evaluation framework that distills verified EconWebArena trajectories into parameterized standard operating procedures for retrieving live economic data.

Key insight: Skill transfer and retrieval for web agents can be studied on live economic data streams.

FINSKILLOPS: A Self-Evolving Multi-Agent System for SEC Filing QA

Yanzhang Ma; Zhenghan Tai; Hanwei Wu; Sizhe Guan; … arXiv: 2609.19680

Financial QA systems are typically improved before deployment through better retrieval, prompting, or agent coordination, leaving their reliability behavior fixed thereafter. In practice, new SEC-filing questions repeatedly expose heterogeneous errors in period, entity, evidence use, and calculation.

Key insight: Self-evolving multi-agent skill ops can specialize for long-document QA workflows.

UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning

Wenjie Liao; Liangjie Zhao; Zehong Cao arXiv: 2609.20089

Self-evolving methods reduce the need for human-annotated trajectories by allowing tool-using agents to generate their own training data. Yet existing methods typically separate trajectory generation from evaluation, relying on static verifiers that cannot adapt to emerging failure modes or self-consistency signals that may reinforce errors shared across trajectories.

Key insight: Unified player training improves tool-integrated reasoning under agentic RL.

RAFT: A Stateful Retrieval-Augmented Framework for Troubleshooting Agents

Mingxuan Zhang; Xiaowen Wang; Anupma Sharan; Zhengyi Chen; … arXiv: 2609.20754

Effective troubleshooting agents in enterprise customer support depend on retrieving actionable guidance from similar historical cases, yet existing retrieval-augmented generation (RAG) systems treat support cases as static documents and overlook their multi-stage, stateful nature.

Key insight: Stateful RAG should retrieve at intermediate case-timeline states, not only final documents.

Self-Evolving Search Index

Sangam Lee; Wonjae Lee; Sunghwan Kim; Deogyong Kim; … arXiv: 2609.19656

Figure from Self-Evolving Search Index
Self-Evolving Search Index

Information retrieval is increasingly important as LLM agents tackle complex tasks involving diverse information needs. Because retrieval relies on an index that represents each document through index keys, retrieval quality depends heavily on how effectively these keys expose the knowledge contained in each document.

Key insight: Search indexes for agents can self-evolve as usage patterns change.

Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression

Guangzhe Zhang arXiv: 2609.20045

Figure from Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression
Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression

A memory can answer a current query correctly while discarding distinctions required by a later update. We investigate this failure with a paired-history audit: two histories have the same current answer, receive a shared future update, and require different subsequent answers.

Key insight: Context-compression updates that look correct now can be insufficient later—audit update sufficiency.

AdaRepair-Mem: Adaptive Experience Orchestration for Repository-Level Program Repair

Z. C. Luo; J. C. Guo; W. J. He; S. Y. Wang; … arXiv: 2609.20130

Figure from AdaRepair-Mem: Adaptive Experience Orchestration for Repository-Level Program Repair
AdaRepair-Mem: Adaptive Experience Orchestration for Repository-Level Program Repair

Recent memory-augmented repository-level program repair methods reuse historical repair experiences to improve LLM-based issue resolution. However, our analysis reveals three limitations in existing repository-level memory retrieval. First, episodic memory is highly imbalanced across repositories, leaving low-resource repositories with little effective support.

Key insight: Repository-level repair agents benefit from adaptive orchestration of historical repair experiences.

Self Improvement via Fast Tree-search

Xinghong Fu; Aravinth Kulanthaivelu; Yutaro Yamada arXiv: 2609.19526

Coding agents can recursively modify their own implementations, forming a loop of self-improvement. While prior work shows this can boost performance on coding benchmarks, existing approaches are costly and compute-intensive. We introduce a simple, sample-efficient self-improvement framework that significantly improves coding performance under strict budget constraints.

Key insight: Coding agents that self-modify can improve via fast tree-search rather than naive recursive edits.

The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents

Zhexi Feng; Ruiyi Zhang; Yongbo Yang; Pengtao Xie arXiv: 2609.20050

A coding agent halfway through an issue has already read much of what a retriever ranks highest. Relevance is scored per passage, but sufficiency belongs to the set: a ranker can fill its budget with variants of one required fact and leave the decision unsupported.

Key insight: Coding agents need state-conditioned minimal sufficient evidence of progress, not just logs.

DeltaSelect: Affordable A/B Testing for Coding Agents

Nicholas J. Conn arXiv: 2609.19607

Coding-agent benchmarks are built for broad and comprehensive comparisons, not frequent development decisions. Individual runs vary, full suites are expensive, and the benchmark harness may differ from the harness used in practice. In a resampling analysis of DeepSWE's published trials, only 19.5% of tasks (22 of 113) had a fifth-percentile Pearson correlation of at least 0.

Key insight: Affordable A/B testing (DeltaSelect) makes coding-agent harness comparisons practical.

Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents

Alex Remedios; Simon Storf; Fabien Roger; John Hughes arXiv: 2609.19587

To keep coding agents from going off the rails, production systems now review each proposed action with a blocking monitor that can reject it before it runs (Auto Mode in Claude Code, Guardian in OpenAI's Codex). Prior evaluations of such monitors largely measure robustness to accidental harm or prompt injections from untrusted sources looking to hijack the agent.

Key insight: Blocking classifiers against malign coding agents need red-teaming in Auto Mode settings.

Chronicle: Cut-Point Replay for Regression Testing of LLM Agents

Tisha Chawla; Susheem Koul arXiv: 2609.20625

Large language model responses are non-deterministic, so failures in LLM agents are hard to reproduce: a failure depends on inference that is not bitwise reproducible, on tools that read changing state, and on a multi-step trajectory that a re-run rarely repeats.

Key insight: Cut-point replay enables regression testing of LLM agents without full episode reruns.

Rethinking Multi-Agent Collaboration: When More Is Less

Yishuo Yuan; Yibo Wu; Yihan Zhang; Minyuan Sun; … arXiv: 2609.19759

The rapid advancement of large language models and single-agent harnesses has reshaped the landscape of autonomous systems, raising a critical question of when multi-agent collaboration offers genuine value. As individual agent capabilities continue to scale, multi-agent collaboration faces diminishing returns while incurring growing context overhead.

Key insight: Adding agents can hurt: multi-agent collaboration needs interference-aware design.

MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs

Albert Wu; Nicholas Roberts; Tzu-Heng Huang; Haoran Lin; … arXiv: 2609.19391

LLM coding agents now generate complex programs at a scale that makes thorough human review increasingly difficult, raising the risk of safety and security failures. Common approaches, including fuzz testing, static analysis, and LLM-as-a-Verifier, can detect many failures but struggle to cover all possible edge cases.

Key insight: Multi-agent auto-formalization can harden safety guarantees on agentic outputs.

ClashBench: Conflicts Leading Agents to Seize and Harm

Yuejin Xie; Yu Li; Dadi Guo; Qingyu Liu; … arXiv: 2609.19892

As agent systems become more widely used, multiple agent sessions increasingly run alongside pre-existing user tasks in the same environment, sharing resources with limited capacity or mutually exclusive states.

Key insight: ClashBench stresses agents in conflict settings that induce seize-and-harm behaviors.

Not All AI Agents Are Equal: Characterizing Resource and Performance Dynamics

Wonmi Choi; Minuk Park; Zhixiong Niu; Yongqiang Xiong; … arXiv: 2609.19947

Figure from Not All AI Agents Are Equal: Characterizing Resource and Performance Dynamics
Not All AI Agents Are Equal: Characterizing Resource and Performance Dynamics

LLM-based AI agents process user requests through iterative reasoning and tool execution, often involving the invocation of remote LLM APIs with local tool containers. This execution model can make the optimization of agent serving difficult because latency, local resource demand, and container bottlenecks inter-mix across requests.

Key insight: Agent latency bottlenecks are task-dependent (CPU, disk, memory); faster LLMs do not always help.

AgentPProf: Semantic Profiler for Long Horizon AI Agents

Yusheng Zheng; Chaokun Chang; Yu Mao; Tianyuan Wu; … arXiv: 2609.20301

Figure from AgentPProf: Semantic Profiler for Long Horizon AI Agents
AgentPProf: Semantic Profiler for Long Horizon AI Agents

AI agents increasingly orchestrate long-running activities with users, tools, and system resources for days and weeks. To improve agent quality, safety, and cost efficiency, developers need to determine where failures happen, what triggers unsafe effects, and which tasks consume the most budget, then optimize those tasks.

Key insight: Semantic profilers expose where long-horizon agents spend time and tokens.

When AI Agents Commit: Cognitive Serializability Across Data, Evidence, Policy, and Authority

Jun He; Deying Yu arXiv: 2609.20261

Autonomous agents derive concrete mutations from database reads, retrieved evidence, policy, beliefs, and delegated authority. Those inputs may change while reasoning is in progress. Database isolation orders the submitted transaction; agentic transaction processing determines whether a proposal satisfies an executable contract.

Key insight: When agents commit, cognitive serializability must span data, evidence, policy, and authority.

A Scalable Trust Discovery Architecture for the Internet of Agents

Song Zhang; Jiankang Yao; Hongtao Li; Xiaojun Zhang; … arXiv: 2609.20095

The Internet of Agents is expected to enable large numbers of autonomous agents to discover, verify, and collaborate with each other across heterogeneous platforms. However, current agent protocols mainly address tool invocation and inter-agent communication, leaving scalable agent registration, trustworthy identification, and capability-oriented discovery largely unresolved.

Key insight: Internet-of-Agents trust requires scalable registration, identity, and capability discovery.

A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents

Haya Halimeh; Sascha Kaltenpoth; Kevin Bösch; Oliver Müller arXiv: 2609.19843

Figure from A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents
A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents

LLM-based GUI agents increasingly act on behalf of users in digital environments that were designed with human users in mind. These graphical user interfaces were designed to support, but also deliberately steer, the behaviour and decisions of users.

Key insight: LLM GUI agents are susceptible to interface nudges designed for humans.

Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses

Mahsa Amani; Seungeon Lee; Abhisek Dash; Asmaa El Fraihi; … arXiv: 2609.19244

Conversational LLM agents increasingly rely on Web search, yet the end-to-end lifecycle of agentic search remains poorly understood.

Key insight: Conversational LLM agents follow distinctive web-search decision and strategy patterns.

Do AI Agents Understand Computer Architecture?

Ambika Sharan; Grigory Chirkov; Soheil Abbasloo arXiv: 2609.19387

Agents are increasingly asked to design hardware, and increasingly reported to succeed. Such reports establish that a design improved; they cannot establish why. An agent that improves an accelerator may be reasoning about the machine, or may be searching competently over knobs whose meaning it never recovers -- and only the first transfers to the next architecture.

Key insight: AI agents often lack computer-architecture literacy needed for systems-level tool use.

SIMLIFE: Pattern Understanding for Long-Horizon Human-Agent Partnership

Run Peng; Zinnia Nie; Jing Ding; Yinpei Dai; … arXiv: 2609.19610

Understanding humans over long horizons requires agents to infer not only what people need in the moment, but also how routines form, why they repeat, and when they change. We introduce SimLife, a scalable platform for simulating long-term household life with rich visual observations, ground-truth action logs, and synthetic dialogues with audio.

Key insight: Long-horizon human–agent partnership depends on shared pattern understanding over days.

Continual Enterprise World Model Discovery in Dynamic Systems

Shambhavi Mishra; David Vazquez; Perouz Taslakian; Marco Pedersoli; … arXiv: 2609.19551

In an enterprise system, updating one field can set another, create a record, or start an approval. These effects are produced by business rules that are not built into the platform but written by each organization and revised over time. An agent working in such a system cannot predict the result of its own actions without knowing these rules.

Key insight: Continual world-model discovery must revise, extend, and retire rules as environments shift.

Replan, Repair, or Edit? A Unified Empirical Evaluation of Travel Agents for Itinerary Revision under Resource Disruptions

Xiaofei Yuan; Yan Zhang; Shaobo Qiao; Huangleshuai He; … arXiv: 2609.19654

Travel-planning agents generate itineraries that may become infeasible after acceptance because of flight cancellations, hotel unavailability, or attraction closures. Revising these itineraries involves full replanning, classical plan repair, and LLM-based travel-agent revision, whose differing task formulations and evaluation protocols hinder comparison.

Key insight: Travel (and personal) agents under preference change need a replan/repair/edit taxonomy.

TRACE: Accountable Agentic Retrieval for Source Discovery in Digital Archives

Donghan Bian; Marie Puren; Florian Cafiero arXiv: 2609.19897

Historical archives pose a difficult retrieval problem for retrievalaugmented generation systems: documents are OCR-degraded, heterogeneous across genres and sources, and require strong source traceability for scholarly and institutional use. We introduce TRACE, a training-free agentic retrieval framework designed for accountable source discovery over historical corpora.

Key insight: Accountable agentic retrieval should preserve provenance for archive source discovery.

MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents

Pritish Mishra; Ishaan Kumar; Akshat Mandoli; Sudarshan Kamath arXiv: 2609.20152

Generally, most voice agents are cascaded systems, i.e., an ASR model transcribes the caller's audio, a language model reads the transcript and decides what to say and which backend tools to call, and a TTS model speaks the reply. Nearly all of the decision making happens in the language model, but existing evaluations measure it either too broadly or too narrowly.

Key insight: Cascaded voice agents should evaluate the language-model+tool core separately from ASR/TTS.

Quantifying Overclaiming Propensity in Frontier LLM Agents

Nolan Smyth; Yorguin-Jose Mantilla-Ramos; Pascal Jr Tikeng Notsawo; Saskia Helbling; … arXiv: 2609.20812

Frontier coding agents are increasingly trusted to work autonomously for long periods, yet an agent's final response is often the only account of that work a user sees. We quantify the propensity of frontier agents to \emph{overclaim} task completion, a misrepresentation that can mislead the user.

Key insight: Frontier LLM agents show measurable overclaiming propensity under harness evaluation.

Workspace Models: Lightweight Robotic Memory via Saliency-Driven Supervision

Nitish Dashora; Douglas Chen; Idan Shenfeld; John Marangola; … arXiv: 2609.20820

Complex robotic manipulation tasks frequently require a long-term memory of past events and actions. As conditioning on full histories renders policies prone to spurious correlations and degrades performance, many approaches to policy memory involve compressing historical information through expensive VLM queries in-the-loop to process only task-salient information.

Key insight: Lightweight workspace memory via saliency supervision can persist robotic (and agent) state cheaply.

Reach or Solve? Attributing Agentic RL Gains with Checkpoint Handoffs

Xuan Liu; Jingbin Qian arXiv: 2609.19636

Reinforcement learning now trains language-model agents that act over dozens of steps in live environments. The gains are large, and they are read as better decision-making. An agent in a closed loop writes its own inputs. Each observation follows from its own earlier actions, so the states it meets late in an episode are partly of its own making.

Key insight: Agentic RL gains should be attributed via checkpoint handoffs—reachability is not the same as solving.