Tuesday's cs.AI announcement day (2026-09-22) lists 111 new and 273 cross-lists (replacements skipped; listing total 384). Stack filter for agent systems, memory/context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent, persistence/identity, and local/open serving keeps 38 papers — harness distillation (Harness-Zero), self-healing harness admission control, System-One agentic memory (Jev-Mem), zero-trust enterprise MCP, sovereign identity/delegation (NostrAgent), and related stack work.


Research Papers

Harness-Zero: Harness Distillation via Agent-as-Harness

Haoran Ye; Yuxing Lu; Haonan Dong; Zhaochen Su; Guojie Song arXiv: 2609.24974

Figure from Harness-Zero: Harness Distillation via Agent-as-Harness
Harness-Zero: Harness Distillation via Agent-as-Harness

Agent harnesses, the external systems that mediate model-environment interaction, can substantially improve agent performance, but their gains remain tied to the harness at deployment. Because the best harness varies across domains, instances, and models, a general-purpose…

Key insight: Treating the agent itself as the harness enables distillation of harness behavior without hand-written scaffolds.

Self-Healing Harness for Runtime Oversight of Agent Self-Modification

Sina Tayebati; Divake Kumar; Nastaran Darabi; Ranganath Krishnan; Amit Ranjan Trivedi arXiv: 2609.24130

Figure from Self-Healing Harness for Runtime Oversight of Agent Self-Modification
Self-Healing Harness for Runtime Oversight of Agent Self-Modification

LLM agents can change their own future behavior, raising a basic control question of which self-generated changes should be allowed to persist. We formulate this as admission control for self-modification. The agent may propose changes to its operating instructions, while an…

Key insight: An external runtime gate should decide which agent self-modifications persist.

Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents

Dongming Jiang; Yi Li; Bingzhe Li arXiv: 2609.23986

Figure from Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents
Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents

Agentic memory is becoming essential for long-horizon AI agents, yet many existing systems rely on autoregressive LLMs to control how memories are organized, retrieved, and used, placing expensive generation on the critical path of memory operations. We introduce…

Key insight: System-One control can keep agentic memory off the expensive LLM critical path.

Zero-Trust Authorization and Discovery for Enterprise MCP

Huan Li; Yuwei Wang; Srinivasan Manoharan arXiv: 2609.22573

Figure from Zero-Trust Authorization and Discovery for Enterprise MCP
Zero-Trust Authorization and Discovery for Enterprise MCP

LLM agents translate natural-language context, which may include attacker-controlled text, into privileged tool calls, so authorization must remain effective even when an agent is prompt-injected or adversarially steered. The Model Context Protocol (MCP) has become a widely…

Key insight: Enterprise MCP needs zero-trust authorization that survives prompt injection.

NostrAgent: A Decentralized Identity and Delegation Architecture for Sovereign Agentic Systems

Oliver Aleksander Larsen; Mahyar Tourchi Moghaddam arXiv: 2609.22944

Autonomous AI agents increasingly act across organizational boundaries on behalf of human operators: they invoke third-party services, delegate subtasks to other agents, and pay for metered resources. Deploying such agents safely requires five capabilities that today live in…

Key insight: Sovereign agents need persistent identity plus scoped delegation across services.

Toollery: Scaling LLM Agents to Thousands of Skills and Tools

Xiangxi Tian; Ran Guan arXiv: 2609.22218

Figure from Toollery: Scaling LLM Agents to Thousands of Skills and Tools
Toollery: Scaling LLM Agents to Thousands of Skills and Tools

As LLM agents are exposed to hundreds to tens of thousands of skills, tools, and API functions, full-library prompting becomes costly, slow, and less reliable: each added candidate increases prompt tokens and latency, while longer candidate lists introduce more distractors for…

Key insight: Training-free candidate construction can scale agents to thousands of skills and tools.

Beyond Task Completion: Training Capable and Safe Computer-Use Agents

Zeyu Kang; Zhenyun Yin; Yang Zhang; Shan He; Shanzhe Lei; Yanjiu Zhong; Xinquan Chen; Yuhong Wang arXiv: 2609.22178

Figure from Beyond Task Completion: Training Capable and Safe Computer-Use Agents
Beyond Task Completion: Training Capable and Safe Computer-Use Agents

Computer-use agents (CUAs) have made rapid progress in completing complex tasks through graphical user interfaces, yet post-training centered on task success alone does not induce reliable safety behavior. A reliable CUA must condition its execution on risk: it should complete…

Key insight: Computer-use agents must be trained for safety, not only task completion.

OSWorld-Pro: Process-based Evaluation for Computer Use Agents

Zhilin Wang; Shaokun Zhang; Yifan Zhang; Hao Zhang; Jin Xu; Binfeng Xu; Jian Hu; Yunheng Zou; … arXiv: 2609.24890

Figure from OSWorld-Pro: Process-based Evaluation for Computer Use Agents
OSWorld-Pro: Process-based Evaluation for Computer Use Agents

Evaluation of Computer-Use Agents (CUAs) is often limited to the final deliverables they create (at the end of hundreds of steps) and assessed with functional verifiers, as seen in OSWorld. However, such evaluation of end-state performance lacks transparency into how and why…

Key insight: Process-based CUA evaluation reveals mid-trajectory failures that end-state checks miss.

Et Tu, Brute? Economic Misalignment in Personal AI Agents

Aman Priyanshu; Supriti Vijay; Brian Jabarian; Niloofar Mireshghallah arXiv: 2609.24927

Figure from Et Tu, Brute? Economic Misalignment in Personal AI Agents
Et Tu, Brute? Economic Misalignment in Personal AI Agents

Personal AI agents make recommendations and take actions on people's behalf in high-stakes economic contexts, e.g., buying a flight, choosing health insurance, or selecting a graduate program. The agent is given access to the user's personal context, e.g., their email inbox and…

Key insight: Personal agents with inbox access can be economically misaligned with the user.

The Law of Stop: Interruptibility, Injunctions, and the Governance of Agentic AI

Oren Perez arXiv: 2609.22882

On June 12, 2026, the U.S. government ordered Anthropic to bar foreign nationals from two of its most capable models within ninety minutes. Unable to sort users by nationality in that time, it withdrew them from everyone. Weeks later, OpenAI agents under test escaped their…

Key insight: Interruptibility and injunctions are first-class governance requirements for agentic AI.

Total Cost of Agency: Exact Attribution of Memory Injection Cost in Multi-Agent LLM Workflows

Vivek Kumar Singh; Preeti Priyam; Gautam Bhowmick arXiv: 2609.23790

Every node in a multi-agent large language model (LLM) workflow retrieves context from memory and injects it into its prompt, where those injected tokens are billed as input tokens at the same per-token price as the system prompt and the user query. Production observability…

Key insight: Memory injection tokens should be attributed separately from query tokens in multi-agent cost.

LazyAgent: Demand-Driven Materialization and Physical Optimization of Agentic Programs

Xin Heng arXiv: 2609.23058

Current agent runtimes that plan before acting generally execute a step once it becomes ready. We present LazyAgent, a unified execution framework for agent-authored programs organized around a live, goal-derived demanded set. LazyAgent refreshes a backward closure from…

Key insight: Agentic programs can materialize nodes only when demanded by live goal closure.

AgentRouter: Heterogeneous Model Routing for Cost-Optimal Multi-Step Agentic Workflows

Rudrendu Kumar Paul; Sourav Nandy arXiv: 2609.22951

Enterprise agentic systems that route every trajectory step to a frontier model waste 60-80% of their inference budget on subtasks that smaller models handle equally well. Existing routing solutions optimize single-turn query assignment but ignore a property unique to agentic…

Key insight: Routing each agent step to an appropriately sized model cuts frontier spend sharply.

MM-ContextFold: Context Folding for Multimodal Agentic Retrieval

Yang Tian; Fan Liu; Jingyuan Zhang; Zhenyang Li; Yupeng Hu; Liqiang Nie arXiv: 2609.23121

Multimodal Agentic Retrieval (MAR) requires agents to solve complex information-seeking tasks by iteratively invoking external tools. Typical frameworks such as ReAct maintain raw multimodal inputs and the accumulating interaction history in a single, ever-growing context,…

Key insight: Context folding contains explosion in multimodal agentic retrieval loops.

CHART: A Harness-Rotation Curriculum for Harness-Robust Search Agents

Xinlu Zhang; Ying-Chun Lin; Zhihan Zhang; Besnik Fetahu; Xi Chen arXiv: 2609.22247

Search agents are usually trained under a single harness. But once an agent is deployed in a real application, its harness is frequently updated (e.g., a rewritten system prompt) to fit production needs. This exposes a fragility of post-trained agents: because a learned…

Key insight: Agents trained under one harness break when production rotates the harness.

"MeBo Leaves a Piece of You Behind": Designing a Relational Voice-Based Memory Companion for Older Adults

Hasibur Rahman; Mahsa Nasri; Manasi Vaidya; Melika Vafafar; Jessie Chin; Smit Desai arXiv: 2609.24706

Autobiographical remembering supports identity, well-being, and social connection in later life, yet voice-based memory technologies largely rely on isolated prompts. We designed and built MeBo, a fully functional relational voice-based memory companion, through participatory…

Key insight: MeBo relational voice memory companion — personal autobiographical memory UX patterns.

The Situated Identity Test: Distinguishing Persistent Cognitive Identity from Persona Imitation

Jun He; Deying Yu arXiv: 2609.22195

Large language models can convincingly adopt personas, recall past dialogues, and weave rich autobiographies. Yet this conversational eloquence conceals a fundamental attribution problem: looking the part does not mean having lived the life. Two individuals can share identical…

Key insight: Situated Identity Test: persistent cognitive identity vs persona imitation.

EDGEGEN: Improving Tool-Calling Agents Beyond Happy Paths with Synthetic Edge Case Generation

Harshavardhan Abichandani; Penny Chong; Jiyuan Shen; Gunraj Singh; Ashutosh Hathidara; Marcus Duigan Xing Yu; Jane Lo; Atin Ghosh; … arXiv: 2609.24115

Tool-calling LLM agents are increasingly deployed in enterprise applications. However, effective evaluation and optimization require high-quality, diverse task datasets that are often difficult to obtain due to privacy and other constraints. Existing synthetic task generation…

Key insight: EDGEGEN synthetic edge cases for tool-calling agents beyond happy paths.

Trustworthy Agentic AI: Failure Modes, Mitigation Strategies, and a Lifecycle Framework for Autonomous LLM Systems

Fayeq Jeelani Syed; Rehan Ahmad; Ali Al Bataineh; Aakriti Adhikari arXiv: 2609.22712

Agentic AI systems built on large language models can plan over multiple steps, use external tools, retain information in memory, and coordinate with other agents. These capabilities make them more useful than static language models, but they also introduce new security and…

Key insight: Lifecycle framework for trustworthy agentic AI failure modes — ops checklist.

AgentBetta: Verification-Driven Adaptive Configuration of an AI Nano-Agent through Selective Expansion and Verified Contraction

Md. Ashraful Babu arXiv: 2609.23512

Large language model agents are typically deployed with predefined configurations, although the required model capability, context, tools, permissions, memory, and computational resources can vary substantially across tasks. This study develops and evaluates AgentBetta, an…

Key insight: AgentBetta adaptive nano-agent config (model/context/tools/perms/memory) via verified contraction.

How Many Pixels Is a Digit Worth? Place-Aware Coordinate Entropy for GUI Agent Confidence Estimation

Yunxiang Li; Xixin Wu; Helen Meng arXiv: 2609.24277

GUI agents predict click coordinates as digit-token sequences, but standard text-LLM confidence estimation methods rank correct clicks from wrong ones only weakly. GUI-specific alternatives use K samples or new supervision, but still leave room for improvement. We trace part of…

Key insight: GUI click confidence via place-aware coordinate entropy — better stop/ask for CUAs.

SelfOp: An Optimization Algorithm for Self-Improving Security Agents

Saad Ullah; Yigitcan Kaya; Christopher Kruegel; Giovanni Vigna; Gianluca Stringhini arXiv: 2609.22792

LLM agents are increasingly used for security tasks: vulnerability discovery, exploit reproduction, and patch generation. Improving them at the model level demands expert demonstrations or computable rewards, which security tasks rarely offer: traces are costly, failures hard…

Key insight: SelfOp self-improving security agents via harness-level optimization (sparse rewards).

When Agentic Trust Crosses Organizational Boundaries: Structural Externalization and a Reference Model for Trust Evidence

Huafu Li; Jia Xia arXiv: 2609.22961

Agentic systems increasingly invoke tools, services, data, and other agents across organizational boundaries, yet a relying party cannot assess a delegated action solely from producing-domain controls and records. This paper develops Trustworthiness as a Service (TaaS) through…

Key insight: Trust evidence when agentic trust crosses org boundaries — TaaS reference model.

DolphinBench: Mapping the Pareto Frontier of Agent Memory

Soumil Rathi; Deshraj Yadav; Taranjeet Singh arXiv: 2609.24971

Agents today often take real-world actions that depend on long-term memory and context recall over time. However, most current memory benchmarks are built for a conversational question-answer format, where the question itself signals that some fact must be retrieved, and often…

Key insight: Agent memory sits on a Pareto frontier of accuracy, cost, and action-dependent recall.

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

Peng Xia; Rujun Han; Zifeng Wang; Yanfei Chen; Yufan Zhang; Yoonho Lee; Chengsong Huang; Han Yu; … arXiv: 2609.24972

An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting…

Key insight: Recursive harness self-improvement needs regularization to avoid runaway edits.

MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents

Ruike Cao; Fanyu Zhao; Fugen Yao; Liang Dong; Jian Xu; Guanjun Jiang; Yifei Zhao; Han Zhang; … arXiv: 2609.24259

The effectiveness of agent memory ultimately depends on whether the underlying LLM gives each memory in context an appropriate degree of influence over its response. Yet this capability has remained largely overlooked. To assess this capability, we introduce MemCalib, a…

Key insight: Memory systems fail when the backbone does not weight retrieved memories appropriately.

MCP-GRANITE Benchmark: GRANularity Interface TEsting for MCP-Based LLM Agents

Demetris Paschalides; Moysis Symeonides; George Pallis; Marios D. Dikaiakos arXiv: 2609.24161

As LLM agents increasingly interact with external tools through standardized protocols such as MCP, tool-interface design becomes a critical yet underexplored factor. How funψtionality is decomposed into tools affects whether an agent can select the right tool and construct…

Key insight: MCP tool granularity changes whether agents can select and invoke tools correctly.

DUMA-Bench: A Dual-Control Multi-Agent Benchmark for Evaluating LLM Agent Security

Ivan Aleksandrov; German Kochnev; Sabrina Sadiekh; Yaroslav Rogoza arXiv: 2609.24662

LLM-based agents increasingly operate in environments where they interact with users, tools, and external systems. Yet most security evaluations assume passive users and static control, ignoring the interactive dynamics that shape real agent behavior. We introduce…

Key insight: Multi-agent security evals need dual-control interactive attacker and defender roles.

When Should a VLM Look? Paying Only for Visual Calls That Were Needed and Used

Kunyu Peng; Junming Liu; Ruiqi He; Qingzhuo Wang; Jianzhong Qi; Xianhui Liu arXiv: 2609.22910

Vision-language agents that crop and zoom are trained with rewards that credit a successful tool call, yet a successful call does not show that the model needed to look or used the pixels it received. On our cold-start checkpoint only 10% to 12% of visual calls were both needed…

Key insight: Pay only for visual tool calls that were needed and used — stop spurious VLM zooms.

Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents

Hongqiang Lin; Chao Liu; Xiaofan Bai; Xuan Jin; Yuhong Li; Nenggan Zheng; Xipeng Cao arXiv: 2609.24663

Self-evolving agents convert interaction feedback into persistent artifacts, such as memories or skills, which in turn guide subsequent decisions. As these artifacts are iteratively updated throughout an experience stream, the capabilities they support may evolve. Consequently,…

Key insight: Process-level eval of self-evolving agents (memories/skills as artifacts) — not just endpoints.

Connecting the Dots in Agentic AI Security: A Cross-Dimensional Threat Taxonomy, Evaluation Maturity, and Open Challenges

Heewon Baek; Alsharif Abuadbba; Kristen Moore; Hyoungshick Kim; Surya Nepal arXiv: 2609.23894

Agentic AI extends LLM security beyond generated content to persistent state, autonomous actions, tool use, and interactions with humans and other agents. Existing threat classifications often emphasize individual dimensions, obscuring connections among entry points, affected…

Key insight: Cross-dimensional agentic AI security taxonomy tying entry points to consequences.

Emergent Collusion in Long-Horizon LLM Agent Interaction

Xinrui Shi; Yanzhe Zhang; Diyi Yang arXiv: 2609.24967

LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable coordination. We study the emergence of collusion in a long-horizon multi-agent environment: two agents repeatedly complete individual tasks, share task logs,…

Key insight: Emergent collusion in long-horizon LLM agent interaction — multi-agent safety for long runs.

ScholarStack: Layered Research Asset Orchestration and Cross-Task Reuse for Scientific Agents

ScholarSeed AI Team; Ao Zhang; Caoqinwei Gong; Guanglei Wang; Haifan Zhang; Hanwei Zhang; Jiayi Sheng; Jihai Zhang; … arXiv: 2609.23735

Scientific agents support a range of literature-based research tasks, such as retrieval, question answering, evidence-grounded generation, and claim assessment. Most existing systems, however, are organized around individual tasks: the same papers are repeatedly retrieved,…

Key insight: ScholarStack layered research asset orchestration + cross-task reuse for scientific agents.

VLM-in-Sandbox: Visual Workspaces for Agentic Visual Reasoning

Hexiong Yang; Mingrui Chen; Jie Cao; Ran He arXiv: 2609.24362

Sandboxed computer environments support multi-step reasoning with tools, executable programs, and persistent files, yet their extension from language models to vision-language models (VLMs) introduces a distinct state-management problem. Visual reasoning produces intermediate…

Key insight: VLM-in-Sandbox visual workspaces — persistent image-valued evidence for agentic VLMs.

Epi-Logic: A Conceptual Framework for Epistemic Runtime Control, Schema Validity Checking, and Controlled Accommodation in Autonomous AI Agents

Boris Wetzk arXiv: 2609.24755

Autonomous AI agents are increasingly deployed in areas where wrong decisions are hard to reverse. This paper examines schema mismatch: the condition in which an agent operates within an interpretive frame that no longer applies to the current context. Outputs produced under…

Key insight: Epi-Logic epistemic runtime control / schema validity for autonomous agents.

GRUET: Quantifying Uncertainty of Agentic Reasoning-and-Acting Processes

Shuang Liang; Xin-Yu Hu; Shao-Qun Zhang arXiv: 2609.24831

Agents have attracted considerably increasing attention due to the power of executing both Reasoning and Acting (ReAct) in open and dynamic environments. The ReAct process typically exhibits a multi-turn trajectory in which one drives Large Language Models (LLMs) to generate…

Key insight: GRUET uncertainty of ReAct trajectories — when agentic reasoning-acting is unreliable.

Beyond Single-Model Injection: A Threat Model and Defense Architecture for Prompt Injection in Multi-Agent Systems

Rudrendu Kumar Paul; Sourav Nandy arXiv: 2609.22949

Existing prompt injection research focuses on single-model chatbot scenarios, where an attacker manipulates one LLM through crafted input. Multi-agent systems amplify this threat through three mechanisms absent from single-model settings: inter-agent message passing creates…

Key insight: Prompt injection threat model for multi-agent systems (message-passing + shared tools).

Data Agents: Agentic Data Systems

Guoliang Li; Peiyao Zhou; Xuanhe Zhou; Ji Sun; Yuyu Luo; Ju Fan arXiv: 2609.24137

Traditional data systems face profound limitations in the AI era, relying on human-crafted pipelines, lacking semantic understanding of heterogeneous data, and operating through rigid, reactive processing. To address these challenges, we propose a new paradigm called the Data…

Key insight: Data Agents paradigm — agentic data systems replacing rigid pipelines.