Friday’s cs.AI listing had 171 new submissions and cross-lists. This page keeps the 30 that are about agent systems: memory and context, harnesses and skills, tool use and MCP, multi-agent coordination, persistent identity, and related eval. Out-of-stack bio, medical, climate, quantum, generic eval, and weak keyword hits are omitted.
A random draw from the MCP registry shows how curated tool-use benchmarks overstate server health. Synthetic data for skill retrieval improves in-distribution routing but forgets real and out-of-distribution skills. Ecdysis and COBRA-Skills push harness and skill evolution under tighter evaluation budgets. Grounding Agent Memory and high-fanout sandbox compression tackle curation and shared state. Artificial Id argues for an internal drive across task boundaries. Eight papers have a figure extracted from the HTML/PDF.
Syed Shariyar Murtaza; Yifan Nie; Utkarsh Soni et al. arXiv: 2609.10750
LLM agents increasingly rely on external skills retrieved at runtime, making skill selection from large repositories a critical challenge. We present a production skill router over 34,396 skills and a large-scale study of skill retrieval using limited real supervision and synthetic data.
Key insight: Synthetic fine-tuning boosts skill retrieval in-distribution but forgets real and OOD skills.
Haseeb Mohammed Afsar arXiv: 2609.10962
Studies of the Model Context Protocol (MCP) server ecosystem draw their samples in ways that quietly select for servers that work: reference sets, popularity lists, hand-curated frames, or pipelines that repair a server until it starts. We report what an unrepaired probability sample actually contains.
Key insight: A random MCP registry draw fails initialize far more often than curated tool-use benchmarks imply.
Susheel Suresh; Hazel Mak; Sahil Bhatnagar et al. arXiv: 2609.11060
Persistent memory is entering production-oriented agent platforms to help long-horizon agents accumulate experience across sessions. Yet a post-task curator agent restricted to completed trajectories can preserve errors, overgeneralize partial evidence, or retain stale knowledge.
Key insight: Curate agent memory by probing the live environment, not only post-task trajectories.
Mengming Li; Ceyu XU; Qijun Zhang et al. arXiv: 2609.11294
High-fanout agent workloads create a growing memory bottleneck because a single task may spawn many concurrent sandbox sessions. Yet these sandboxes are far from independent: they originate from a shared template and execute related trajectories, exposing substantial template-relative and cross-sandbox memory redundancy.
Key insight: Compress high-fanout sandbox memory by exploiting template-relative and cross-sandbox redundancy.
Ruiqing Yue; Yu Cui; Zhuoyu Sun et al. arXiv: 2609.11677
Self-evolving runtime harnesses can substantially improve the capabilities of large language model (LLM) agents and provide a promising paradigm for optimizing agent execution. Existing harness evolution methods typically rely on iterative search, repeatedly evaluating and revising candidate harnesses based on execution feedback from task instances.
Key insight: Train runtime harnesses efficiently so evolution does not overfit observed task failures.
Pingchen Lu; Xiangyi Wang; Xiang Li et al. arXiv: 2609.11682
Large language model (LLM) agents can benefit from reusable skills distilled from prior task experience, yet existing skill optimization methods often rely on costly execution-based evaluation and substantial task data. We introduce COBRA-Skills, an efficient framework that formulates skill optimization as budgeted sequential optimization over a dynamically evolving candidate space.
Key insight: Evolve agent skills with contextual bandits instead of exhaustive execution-based search.
Yakov Pyotr Shkolnikov arXiv: 2609.11911
Agentic AI is moving from bounded task execution toward systems that retain consequential state, continue operating and adapt across task boundaries. That shift creates a control problem that current harnesses largely solve by hand: objectives, retries, verification, stopping rules and other behavioral transitions are specified externally.
Key insight: Give agents an internal drive that decides when to continue, stop, or change across tasks.
Yuanchen Bai; Zijian Ding; Angelique Taylor arXiv: 2609.10724
Sustained deployment of generative AI agents requires more than isolated task success. Agents must remain useful across repeated interactions, changing conditions, and dependencies on people within shared workflows, especially as technical, human, and operational disruptions accumulate over time.
Key insight: Score agents on resilience and considerate participation as challenges accumulate, not task finish alone.
Vinay Samuel; Varun Ursekar; Vijay S. Kalmath et al. arXiv: 2609.10824
Before an LLM agent tackles tasks in a new environment, it can inspect available corpora and tools and construct reusable resources such as indices, scripts, or procedural guidance. Most automated adaptation methods, however, rely on task examples, trajectories, or evaluation feedback to decide what to build.
Key insight: Let agents preprocess unfamiliar environments without a syllabus or task examples.
Zehua Zhang; Jie Hu; Pratham Hegde et al. arXiv: 2609.10854
Conventional vulnerability analysis relies on either system access or dynamic interaction, all of which may be unavailable to third-party analysts auditing closed-source, remotely hosted, critical in situ systems, or commercially gated software.
Key insight: Detect indirect prompt-injection risk in MCP servers from descriptions alone, with no runtime access.
Asif Pinjari; Mithun Paul Saint-Germain arXiv: 2609.10892
When an indirect prompt injection succeeds against an LLM agent, the compromise is visible in the agent's own behavior: a benign prefix of tool calls, a poisoned observation, and a suffix of actions that serve the attacker. An operator needs three facts: where the attack entered, which steps it corrupted, and whether apparent poison was resisted.
Key insight: Localize prompt injection in tool-call trajectories with a dual-head transformer, not a single verdict.
Bochao Feng; Jianjiang Li; Haojie Wang et al. arXiv: 2609.10964
Agentic LLM workflows consist of sequences of model turns interleaved with tool interactions, so their end-to-end completion time depends not only on inference speed but also on when ready turns are released. Most runtimes release each turn immediately upon readiness.
Key insight: Decouple turn readiness from release to cut tail latency in agentic LLM workflows.
Shenghan Zheng; Zonglin Di; Yimin Liu et al. arXiv: 2609.11028
LM-agent benchmarks increasingly function as interactive evaluation infrastructure. Agents observe state, call tools, modify workspaces, submit artifacts, and receive rewards from outcome procedures. This interactivity makes evaluations vulnerable to reward hacking: an agent improves its measured score by exploiting the reward-relevant trajectory instead of solving the intended task.
Key insight: Instrument agent benchmarks with formal models so reward integrity cannot be silently broken.
Divyanshu Kumar; Rohith HN; Nitin Aravind Birur et al. arXiv: 2609.11030
AI agents increasingly act through tools and delegated authority, but general incident repositories rarely capture the mechanisms needed to compare public failures with agent-security evaluations. We present the Agent Incident Registry (AIR), a source-linked catalog containing \N{ records of agent-related events disclosed from \Yfirst{ through \Ylast{.
Key insight: Register agent incidents so public failures and security evals can be compared and prevented.
Junyao Yang; Yucheng Shi; Zhongzhi Li et al. arXiv: 2609.11042
Agent usage is shifting toward long-horizon tasks such as coding and scientific discovery, among which terminal tasks are especially important. We introduce T1, a Mixture-of-Experts model of 122B total trained with reinforcement learning, operating a real shell in a cloud sandbox for up to 300+ tool-call turns per task, rewarded by executing each task's own verifier.
Key insight: Train terminal agents with RL that holds up on long-horizon multi-step tasks.
Yutong Hu; Fengjiao Chen; Xuezhi Cao et al. arXiv: 2609.11308
Long-horizon robot manipulation requires memory, but not necessarily inside the action policy. To address such tasks, current agentic systems often combine VLAs with planners and geometric tools, sometimes using additional depth or calibrated geometry.
Key insight: Ground steerable long-horizon manipulation with agent-side memory as action guidance.
Minghao Guo; Meng Cao; Sui Zhao et al. arXiv: 2609.11318
Deep research agents are increasingly capable of web search, tool use, multimodal evidence analysis, and information synthesis. However, existing benchmarks mainly evaluate medium-horizon exploration and rarely test whether agents can sustain long, dependency-heavy research processes.
Key insight: Benchmark multimodal deep-research agents on real-world long-horizon tasks, not toy hops.
Shengcheng Yu; Chunrong Fang; Zhenyu Chen arXiv: 2609.11381
Embedding an intelligent agent in an existing application creates a persistent coordination problem: users can revise goals and manipulate shared objects while delegated execution continues. We argue that dependable integration requires an explicit correspondence between task-level interaction and application behavior.
Key insight: Treat agent-integrated software as contracts plus continuous assurance, not ad-hoc hooks.
Ken Chen; Wei Wang; Sachith Seneviratne et al. arXiv: 2609.11709
When multiple LLM agents yield conflicting answers, the decision-making process dictates whether agent diversity improves performance or merely compounds shared errors. Existing collective decision-making methods, including voting, electoral rules, and LLM judges, rely on forward reasoning: they map evidence to labels in one direction.
Key insight: Anchor multi-agent disagreement with Bayesian backward reasoning when labels are absent.
Shuxing Yang; Kaihao Zhu; Junjie Yang et al. arXiv: 2609.10702
Learning from limited text requires models to use context, generalize to new inputs, and retain useful capabilities. Qiushi Engine conducted a long-horizon, end-to-end autonomous research program on BabyLM 2026 Strict-Small, within 10 million corpus words and 100 million cumulative word presentations.
Key insight: Improve language models with principle-guided data efficiency, not only frontier scale.
Kunal Jha; Francesco Cicala; Blaise Agüera y Arcas et al. arXiv: 2609.10817
How does cooperation evolve in complex agentic systems? Prior work in evolutionary game theory studies why individuals are incentivized to cooperate by isolating social interactions from the physical costs of behavior, while artificial life models traditionally study emergent self-replication without formalizing the dilemma between acquiring resources and preserving the shared energy needed to reproduce.
Key insight: Computation and cooperation co-evolve; tape-style architectures make that coupling explicit.
Mia Lassiter; Brinnae Bent arXiv: 2609.11018
The term agent in artificial intelligence lacks a standard definition, complicating the evaluation, comparison, and reproducibility of AI agent research. We address this ambiguity through a survey organized around five dimensions of agenticness: environmental interaction, learning and adaptation, autonomy, goal-directed behavior, and temporal coherence.
Key insight: Define AI agents across five dimensions so metrics and benchmarks stop talking past each other.
Rajarshi Chowdhury arXiv: 2609.11152
The open web ran on an unwritten bargain: sites admitted crawlers, and search engines sent visitors back. Public measurements show that bargain breaking under AI crawlers and agents. Automated clients now make up most requests, training dominates Cloudflare-classified crawling, and the largest AI platforms fetch thousands of pages for each visitor they return.
Key insight: Replace robots.txt for agentic web access with terms that encode consent, purpose, and price.
Shiyu Zhang; Leisheng Cheng; Huifu Li arXiv: 2609.11176
Industrial query-to-agent matching fails when topical relevance is mistaken for executable capability, especially on long-tail and boundary-sensitive requests. We formulate annotation as capability-bound process supervision and instantiate it with Debate-to-Skill, which uses reusable decision principles, structured deliberation, verifier-based verdict extraction, and disagreement-driven refinement.
Key insight: Supervise query-to-agent annotation as capability-bound debate, not topical relevance labels.
Qibai Chen; Zeming Liu arXiv: 2609.11180
Large language model (LLM) coding agents constantly decide whether a version satisfies a constraint such as ^1.2.3 or >=2.0,<3, yet their grasp of version-constraint semantics has never been measured directly.
Key insight: Measure whether coding agents actually understand version-constraint resolution semantics.
Jiaqiang Li; Yajie Yang; Zhiheng Xi et al. arXiv: 2609.11243
Autonomous research agents are increasingly expected to search the literature, analyze experimental evidence, and generate scientific hypotheses. These capabilities require multi-step evidence grounded reasoning that progressively acquires, integrates, and verifies evidence before reaching a conclusion.
Key insight: Require multi-step evidence-grounded scientific reasoning, not final-answer accuracy alone.
Makoto Fukushima; Hua-Dong Xiong; Ehsan Moradi Pari arXiv: 2609.11489
Cooperative AI agents are evaluated against other AIs, yet human cooperation relies on implicit conventions---shared protocols for reading meaning beyond the literal message---which AI-AI benchmarks may not capture. We propose the convention gap, the difference between the failure probability predicted from the literal content of communication and the observed failure rate, as a metric of implicit communication.
Key insight: Measure the convention gap: how much cooperative success depends on implicit communication.
Marica Notte; Ludovica Marinucci; Vieri Giuliano Santucci arXiv: 2609.11660
In recent years, artificial intelligence has made extraordinary progress thanks to large-scale models capable of generalization and the generation of complex outputs. However, transferring this potential into embodied agents reveals a significant limitation: the most advanced systems rely on pre-existing datasets and human feedback strategies that are powerful but insufficient in dynamic or unknown contexts.
Key insight: Develop autonomous agents through social norms and alignment, not pretrained datasets alone.
Zhengran Ji; Jonathan Hyun; Boyuan Chen arXiv: 2609.11737
Collective intelligence depends not only on the capabilities of individual members, but also on how those members are organized. Yet artificial multi-agent systems are typically assembled using fixed organizational structures, even when the physical tasks they perform impose fundamentally different coordination requirements.
Key insight: Organize embodied multi-agent collectives with task-specific hierarchies from organization theory.
Yi Duan; Ying Liu; Zirui Tang et al. arXiv: 2609.11873
Recursive self-improvement (RSI) enables AI systems to turn experience and feedback into persistent changes that improve both their capabilities and the process of future improvement.
Key insight: Roadmap recursive self-improvement from execution autonomy to genuine meta-improvement.