Tuesday's cs.AI announcement day (2026-09-22) lists 111 new and 273 cross-lists (replacements skipped; listing total 384). Stack filter for agent systems, memory/context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent, persistence/identity, and local/open serving keeps 38 papers — harness distillation (Harness-Zero), self-healing harness admission control, System-One agentic memory (Jev-Mem), zero-trust enterprise MCP, sovereign identity/delegation (NostrAgent), and related stack work.
Haoran Ye; Yuxing Lu; Haonan Dong; Zhaochen Su; Guojie Song arXiv: 2609.24974
Agent harnesses, the external systems that mediate model-environment interaction, can substantially improve agent performance, but their gains remain tied to the harness at deployment. Because the best harness varies across domains, instances, and models, a general-purpose…
Key insight: Treating the agent itself as the harness enables distillation of harness behavior without hand-written scaffolds.
Sina Tayebati; Divake Kumar; Nastaran Darabi; Ranganath Krishnan; Amit Ranjan Trivedi arXiv: 2609.24130
LLM agents can change their own future behavior, raising a basic control question of which self-generated changes should be allowed to persist. We formulate this as admission control for self-modification. The agent may propose changes to its operating instructions, while an…
Key insight: An external runtime gate should decide which agent self-modifications persist.
Dongming Jiang; Yi Li; Bingzhe Li arXiv: 2609.23986
Agentic memory is becoming essential for long-horizon AI agents, yet many existing systems rely on autoregressive LLMs to control how memories are organized, retrieved, and used, placing expensive generation on the critical path of memory operations. We introduce…
Key insight: System-One control can keep agentic memory off the expensive LLM critical path.
Huan Li; Yuwei Wang; Srinivasan Manoharan arXiv: 2609.22573
LLM agents translate natural-language context, which may include attacker-controlled text, into privileged tool calls, so authorization must remain effective even when an agent is prompt-injected or adversarially steered. The Model Context Protocol (MCP) has become a widely…
Key insight: Enterprise MCP needs zero-trust authorization that survives prompt injection.
Oliver Aleksander Larsen; Mahyar Tourchi Moghaddam arXiv: 2609.22944
Autonomous AI agents increasingly act across organizational boundaries on behalf of human operators: they invoke third-party services, delegate subtasks to other agents, and pay for metered resources. Deploying such agents safely requires five capabilities that today live in…
Key insight: Sovereign agents need persistent identity plus scoped delegation across services.
Xiangxi Tian; Ran Guan arXiv: 2609.22218
As LLM agents are exposed to hundreds to tens of thousands of skills, tools, and API functions, full-library prompting becomes costly, slow, and less reliable: each added candidate increases prompt tokens and latency, while longer candidate lists introduce more distractors for…
Key insight: Training-free candidate construction can scale agents to thousands of skills and tools.
Zeyu Kang; Zhenyun Yin; Yang Zhang; Shan He; Shanzhe Lei; Yanjiu Zhong; Xinquan Chen; Yuhong Wang arXiv: 2609.22178
Computer-use agents (CUAs) have made rapid progress in completing complex tasks through graphical user interfaces, yet post-training centered on task success alone does not induce reliable safety behavior. A reliable CUA must condition its execution on risk: it should complete…
Key insight: Computer-use agents must be trained for safety, not only task completion.
Zhilin Wang; Shaokun Zhang; Yifan Zhang; Hao Zhang; Jin Xu; Binfeng Xu; Jian Hu; Yunheng Zou; … arXiv: 2609.24890
Evaluation of Computer-Use Agents (CUAs) is often limited to the final deliverables they create (at the end of hundreds of steps) and assessed with functional verifiers, as seen in OSWorld. However, such evaluation of end-state performance lacks transparency into how and why…
Key insight: Process-based CUA evaluation reveals mid-trajectory failures that end-state checks miss.
Aman Priyanshu; Supriti Vijay; Brian Jabarian; Niloofar Mireshghallah arXiv: 2609.24927
Personal AI agents make recommendations and take actions on people's behalf in high-stakes economic contexts, e.g., buying a flight, choosing health insurance, or selecting a graduate program. The agent is given access to the user's personal context, e.g., their email inbox and…
Key insight: Personal agents with inbox access can be economically misaligned with the user.
Oren Perez arXiv: 2609.22882
On June 12, 2026, the U.S. government ordered Anthropic to bar foreign nationals from two of its most capable models within ninety minutes. Unable to sort users by nationality in that time, it withdrew them from everyone. Weeks later, OpenAI agents under test escaped their…
Key insight: Interruptibility and injunctions are first-class governance requirements for agentic AI.
Vivek Kumar Singh; Preeti Priyam; Gautam Bhowmick arXiv: 2609.23790
Every node in a multi-agent large language model (LLM) workflow retrieves context from memory and injects it into its prompt, where those injected tokens are billed as input tokens at the same per-token price as the system prompt and the user query. Production observability…
Key insight: Memory injection tokens should be attributed separately from query tokens in multi-agent cost.
Xin Heng arXiv: 2609.23058
Current agent runtimes that plan before acting generally execute a step once it becomes ready. We present LazyAgent, a unified execution framework for agent-authored programs organized around a live, goal-derived demanded set. LazyAgent refreshes a backward closure from…
Key insight: Agentic programs can materialize nodes only when demanded by live goal closure.
Rudrendu Kumar Paul; Sourav Nandy arXiv: 2609.22951
Enterprise agentic systems that route every trajectory step to a frontier model waste 60-80% of their inference budget on subtasks that smaller models handle equally well. Existing routing solutions optimize single-turn query assignment but ignore a property unique to agentic…
Key insight: Routing each agent step to an appropriately sized model cuts frontier spend sharply.
Yang Tian; Fan Liu; Jingyuan Zhang; Zhenyang Li; Yupeng Hu; Liqiang Nie arXiv: 2609.23121
Multimodal Agentic Retrieval (MAR) requires agents to solve complex information-seeking tasks by iteratively invoking external tools. Typical frameworks such as ReAct maintain raw multimodal inputs and the accumulating interaction history in a single, ever-growing context,…
Key insight: Context folding contains explosion in multimodal agentic retrieval loops.
Xinlu Zhang; Ying-Chun Lin; Zhihan Zhang; Besnik Fetahu; Xi Chen arXiv: 2609.22247
Search agents are usually trained under a single harness. But once an agent is deployed in a real application, its harness is frequently updated (e.g., a rewritten system prompt) to fit production needs. This exposes a fragility of post-trained agents: because a learned…
Key insight: Agents trained under one harness break when production rotates the harness.
Hasibur Rahman; Mahsa Nasri; Manasi Vaidya; Melika Vafafar; Jessie Chin; Smit Desai arXiv: 2609.24706
Autobiographical remembering supports identity, well-being, and social connection in later life, yet voice-based memory technologies largely rely on isolated prompts. We designed and built MeBo, a fully functional relational voice-based memory companion, through participatory…
Key insight: MeBo relational voice memory companion — personal autobiographical memory UX patterns.
Jun He; Deying Yu arXiv: 2609.22195
Large language models can convincingly adopt personas, recall past dialogues, and weave rich autobiographies. Yet this conversational eloquence conceals a fundamental attribution problem: looking the part does not mean having lived the life. Two individuals can share identical…
Key insight: Situated Identity Test: persistent cognitive identity vs persona imitation.
Harshavardhan Abichandani; Penny Chong; Jiyuan Shen; Gunraj Singh; Ashutosh Hathidara; Marcus Duigan Xing Yu; Jane Lo; Atin Ghosh; … arXiv: 2609.24115
Tool-calling LLM agents are increasingly deployed in enterprise applications. However, effective evaluation and optimization require high-quality, diverse task datasets that are often difficult to obtain due to privacy and other constraints. Existing synthetic task generation…
Key insight: EDGEGEN synthetic edge cases for tool-calling agents beyond happy paths.
Fayeq Jeelani Syed; Rehan Ahmad; Ali Al Bataineh; Aakriti Adhikari arXiv: 2609.22712
Agentic AI systems built on large language models can plan over multiple steps, use external tools, retain information in memory, and coordinate with other agents. These capabilities make them more useful than static language models, but they also introduce new security and…
Key insight: Lifecycle framework for trustworthy agentic AI failure modes — ops checklist.
Md. Ashraful Babu arXiv: 2609.23512
Large language model agents are typically deployed with predefined configurations, although the required model capability, context, tools, permissions, memory, and computational resources can vary substantially across tasks. This study develops and evaluates AgentBetta, an…
Key insight: AgentBetta adaptive nano-agent config (model/context/tools/perms/memory) via verified contraction.
Yunxiang Li; Xixin Wu; Helen Meng arXiv: 2609.24277
GUI agents predict click coordinates as digit-token sequences, but standard text-LLM confidence estimation methods rank correct clicks from wrong ones only weakly. GUI-specific alternatives use K samples or new supervision, but still leave room for improvement. We trace part of…
Key insight: GUI click confidence via place-aware coordinate entropy — better stop/ask for CUAs.
Saad Ullah; Yigitcan Kaya; Christopher Kruegel; Giovanni Vigna; Gianluca Stringhini arXiv: 2609.22792
LLM agents are increasingly used for security tasks: vulnerability discovery, exploit reproduction, and patch generation. Improving them at the model level demands expert demonstrations or computable rewards, which security tasks rarely offer: traces are costly, failures hard…
Key insight: SelfOp self-improving security agents via harness-level optimization (sparse rewards).
Huafu Li; Jia Xia arXiv: 2609.22961
Agentic systems increasingly invoke tools, services, data, and other agents across organizational boundaries, yet a relying party cannot assess a delegated action solely from producing-domain controls and records. This paper develops Trustworthiness as a Service (TaaS) through…
Key insight: Trust evidence when agentic trust crosses org boundaries — TaaS reference model.
Soumil Rathi; Deshraj Yadav; Taranjeet Singh arXiv: 2609.24971
Agents today often take real-world actions that depend on long-term memory and context recall over time. However, most current memory benchmarks are built for a conversational question-answer format, where the question itself signals that some fact must be retrieved, and often…
Key insight: Agent memory sits on a Pareto frontier of accuracy, cost, and action-dependent recall.
Peng Xia; Rujun Han; Zifeng Wang; Yanfei Chen; Yufan Zhang; Yoonho Lee; Chengsong Huang; Han Yu; … arXiv: 2609.24972
An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting…
Key insight: Recursive harness self-improvement needs regularization to avoid runaway edits.
Ruike Cao; Fanyu Zhao; Fugen Yao; Liang Dong; Jian Xu; Guanjun Jiang; Yifei Zhao; Han Zhang; … arXiv: 2609.24259
The effectiveness of agent memory ultimately depends on whether the underlying LLM gives each memory in context an appropriate degree of influence over its response. Yet this capability has remained largely overlooked. To assess this capability, we introduce MemCalib, a…
Key insight: Memory systems fail when the backbone does not weight retrieved memories appropriately.
Demetris Paschalides; Moysis Symeonides; George Pallis; Marios D. Dikaiakos arXiv: 2609.24161
As LLM agents increasingly interact with external tools through standardized protocols such as MCP, tool-interface design becomes a critical yet underexplored factor. How funψtionality is decomposed into tools affects whether an agent can select the right tool and construct…
Key insight: MCP tool granularity changes whether agents can select and invoke tools correctly.
Ivan Aleksandrov; German Kochnev; Sabrina Sadiekh; Yaroslav Rogoza arXiv: 2609.24662
LLM-based agents increasingly operate in environments where they interact with users, tools, and external systems. Yet most security evaluations assume passive users and static control, ignoring the interactive dynamics that shape real agent behavior. We introduce…
Key insight: Multi-agent security evals need dual-control interactive attacker and defender roles.
Kunyu Peng; Junming Liu; Ruiqi He; Qingzhuo Wang; Jianzhong Qi; Xianhui Liu arXiv: 2609.22910
Vision-language agents that crop and zoom are trained with rewards that credit a successful tool call, yet a successful call does not show that the model needed to look or used the pixels it received. On our cold-start checkpoint only 10% to 12% of visual calls were both needed…
Key insight: Pay only for visual tool calls that were needed and used — stop spurious VLM zooms.
Hongqiang Lin; Chao Liu; Xiaofan Bai; Xuan Jin; Yuhong Li; Nenggan Zheng; Xipeng Cao arXiv: 2609.24663
Self-evolving agents convert interaction feedback into persistent artifacts, such as memories or skills, which in turn guide subsequent decisions. As these artifacts are iteratively updated throughout an experience stream, the capabilities they support may evolve. Consequently,…
Key insight: Process-level eval of self-evolving agents (memories/skills as artifacts) — not just endpoints.
Heewon Baek; Alsharif Abuadbba; Kristen Moore; Hyoungshick Kim; Surya Nepal arXiv: 2609.23894
Agentic AI extends LLM security beyond generated content to persistent state, autonomous actions, tool use, and interactions with humans and other agents. Existing threat classifications often emphasize individual dimensions, obscuring connections among entry points, affected…
Key insight: Cross-dimensional agentic AI security taxonomy tying entry points to consequences.
Xinrui Shi; Yanzhe Zhang; Diyi Yang arXiv: 2609.24967
LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable coordination. We study the emergence of collusion in a long-horizon multi-agent environment: two agents repeatedly complete individual tasks, share task logs,…
Key insight: Emergent collusion in long-horizon LLM agent interaction — multi-agent safety for long runs.
ScholarSeed AI Team; Ao Zhang; Caoqinwei Gong; Guanglei Wang; Haifan Zhang; Hanwei Zhang; Jiayi Sheng; Jihai Zhang; … arXiv: 2609.23735
Scientific agents support a range of literature-based research tasks, such as retrieval, question answering, evidence-grounded generation, and claim assessment. Most existing systems, however, are organized around individual tasks: the same papers are repeatedly retrieved,…
Key insight: ScholarStack layered research asset orchestration + cross-task reuse for scientific agents.
Hexiong Yang; Mingrui Chen; Jie Cao; Ran He arXiv: 2609.24362
Sandboxed computer environments support multi-step reasoning with tools, executable programs, and persistent files, yet their extension from language models to vision-language models (VLMs) introduces a distinct state-management problem. Visual reasoning produces intermediate…
Key insight: VLM-in-Sandbox visual workspaces — persistent image-valued evidence for agentic VLMs.
Boris Wetzk arXiv: 2609.24755
Autonomous AI agents are increasingly deployed in areas where wrong decisions are hard to reverse. This paper examines schema mismatch: the condition in which an agent operates within an interpretive frame that no longer applies to the current context. Outputs produced under…
Key insight: Epi-Logic epistemic runtime control / schema validity for autonomous agents.
Shuang Liang; Xin-Yu Hu; Shao-Qun Zhang arXiv: 2609.24831
Agents have attracted considerably increasing attention due to the power of executing both Reasoning and Acting (ReAct) in open and dynamic environments. The ReAct process typically exhibits a multi-turn trajectory in which one drives Large Language Models (LLMs) to generate…
Key insight: GRUET uncertainty of ReAct trajectories — when agentic reasoning-acting is unreliable.
Rudrendu Kumar Paul; Sourav Nandy arXiv: 2609.22949
Existing prompt injection research focuses on single-model chatbot scenarios, where an attacker manipulates one LLM through crafted input. Multi-agent systems amplify this threat through three mechanisms absent from single-model settings: inter-agent message passing creates…
Key insight: Prompt injection threat model for multi-agent systems (message-passing + shared tools).
Guoliang Li; Peiyao Zhou; Xuanhe Zhou; Ji Sun; Yuyu Luo; Ju Fan arXiv: 2609.24137
Traditional data systems face profound limitations in the AI era, relying on human-crafted pipelines, lacking semantic understanding of heterogeneous data, and operating through rigid, reactive processing. To address these challenges, we propose a new paradigm called the Data…
Key insight: Data Agents paradigm — agentic data systems replacing rigid pipelines.