Tuesday’s cs.AI announcement day lists 147 new and 261 cross-lists (replacements skipped; listing total 408). Stack filter for agent systems, memory/context, computer-use / GUI / tools / skills / harnesses, multi-agent, persistence/identity, and local/open models keeps 40 papers — a dense but readable day spanning tool boundaries, stateful handoffs, plan-injection attacks, and open models for agentic search.


Research Papers

When Tool Calls Succeed but Workflows Fail: Anomalies at the Agent-Tool Boundary

Trofimov; Artem; Novikov; Boris arXiv: 2609.15397

AI agents increasingly execute long-running workflows that externalize effects through independently supplied tools. Under retries, speculative execution, concurrency, and partial failures, the resulting external state may be inconsistent with the workflow's intended resolution: required effects may be missing or duplicated, aborted effects may survive, and committed effects may depend on provisional state that is later withdrawn.

Key insight: Tool-call success is not workflow success — external state can diverge under retries and partial failure.


RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments

Zhu; Sibo; Fan; Shicheng; Wang; Xinyue; … arXiv: 2609.15364

Digital agents must often adapt to new environments whose interfaces, tools, and failure modes are not fully captured by pretrained models. We introduce RSIAgent, a training-free multi-agent framework for recursive self-improvement through autonomous memory construction.

Key insight: Training-free recursive self-improvement via autonomous memory construction for unfamiliar environments.


BusMA: A Bus Communication Substrate for Multi-Agent Systems

Peng; Yanwen; Zhang; Delvin Ce; Wang; Xi; … arXiv: 2609.15054

Multi-Agent (MA) systems are effective at solving complex tasks that demand planning, tool use, and the synthesis of evidence from multiple sources. Existing systems typically adopt Hierarchical Manager-Worker (HMW) or Router-based Message Passing (RMP) structures as their communication protocol.

Key insight: A shared bus substrate can replace hierarchical manager-worker or router message passing in multi-agent systems.


Semantic-TVM: Structure-Preserving Trustworthy Virtual Memory for Memory-Augmented and Tool-Using Agents

Li; Yu; Cai; Qikun; Huang; Tao; … arXiv: 2609.15011

Memory-augmented and tool-using agents expose exact private values when remote LLMs process retrieved memory, tool actions, and intermediate observations. One-way masking limits direct exposure but removes values needed for trusted execution and can leak them through later observations.

Key insight: Structure-preserving trustworthy virtual memory keeps private values usable without raw exposure to remote LLMs.


Loop-Back Authority in LLM Agent Teams: A Paired Experiment on Flat and Hierarchical Coordination

Agachan; Burak; Duijn; Max van; Zohrehvand; Amirhossein arXiv: 2609.14767

Hierarchical orchestration, in which a Manager agent reviews worker output and can send it back for revision, is the default coordination pattern in production multi-agent LLM frameworks. Classical organizational theory predicts that the authority link speeds convergence on decisive output; work on sycophancy and Degeneration-of-Thought predicts that authoritative critique makes LLM output worse.

Key insight: Loop-back authority in manager-worker teams can hurt LLM output despite classical org-theory predictions.


AcquireBound: Runtime Authorization for Resources Acquired by AI Agents

Zhu; Genliang arXiv: 2609.14744

By acquiring compute, credentials, accounts, services, and other agents, autonomous AI agents can introduce new authority into a task. Payment, budget, OAuth, mandate, and fulfillment checks can validate transaction conditions without deciding whether a returned resource may become usable authority.

Key insight: Acquired compute, credentials, and services need runtime authorization before they become usable agent authority.


LIMBO: Lifelong Inference-Time Memory and Budget Optimization for LLM Agents

Sharma; Siddharth; Pandey; Nilesh Prasad; Gungor; Onat; … arXiv: 2609.14138

As LLM agents become integrated into increasingly complex workflows, they must continually acquire new capabilities while retaining competence on previously learned tasks. Lifelong agents address this through experience replay, injecting past interactions into the prompt to leverage prior experience during inference.

Key insight: Lifelong agents need joint memory and budget optimization at inference time, not only experience replay.


Do Not Restart: Residual Completion for Stateful Agent Handoffs

Deng; Runzhi; Zhong; Yiming; Zhao; Fang; … arXiv: 2609.13800

Figure from Do Not Restart: Residual Completion for Stateful Agent Handoffs
Do Not Restart: Residual Completion for Stateful Agent Handoffs

Routing and cascades reduce tool-agent cost by transferring control across models, but stateful handoffs must preserve accepted choices, realized effects, and unfinished obligations. We formulate this as commitment-constrained residual completion and introduce Commitment-Frontier Residual Completion (CFRC).

Key insight: Stateful handoffs should complete residual obligations instead of restarting from a clean slate.


Recoverability as a System Primitive for Long-Horizon AI Agents

Zhang; Zhihui; Liu; Wei arXiv: 2609.13672

AI agents can be interrupted while editing files, calling tools, or carrying out multi-step tasks. Restarting repeats completed work, but continuing from unverified or outdated progress can carry earlier errors forward.

Key insight: Treat recoverability as a system primitive for long-horizon agents interrupted mid-tool or mid-edit.


GeoSkill:Experience-Driven Hierarchical Skill Learning with Collaborative Revision forGeospatialAgents

Luo; Han; Xu; Xian; Liu; Yinhe; … arXiv: 2609.13667

Geospatial agents are increasingly expected to support recurring and evolving analytical tasks rather than execute isolated workflows. In such settings, effective agents must distill prior execution experience into reusable geospatial procedural knowledge to guide future planning and tool use.

Key insight: Hierarchical skill learning with collaborative revision for recurring geospatial agent workflows.


AutoTailor: Automatic, User-Aligned Capability Selection and Adaptation for Web Agents

Cao; Xinyun; Szekeres; Adriana; Faisal; Fazle Elahi arXiv: 2609.13548

Figure from AutoTailor: Automatic, User-Aligned Capability Selection and Adaptation for Web Agents
AutoTailor: Automatic, User-Aligned Capability Selection and Adaptation for Web Agents

Web agents can utilize reusable tools to reduce the cost and latency of low-level browser interaction, but automatically discovered tool collections can be large, redundant, and poorly aligned with user demand. We present AutoTailor, a meta-agentic framework for constructing and maintaining a compact set of trajectory-derived Model Context Protocol (MCP) APIs.

Key insight: AutoTailor selects and adapts web-agent capabilities to stay aligned with user intent.


Token Efficient Task Execution via Application Behavior Modeling for Web Agents

Ianta; Alexandru; Stroulia; Eleni arXiv: 2609.13491

The strong performance of AI Agents across an impressive variety of tasks is driving an unprecedented investment in agentic infrastructures, however the cost of processing tokens is fast increasing. Web agents automate the execution of web-application tasks described in natural language, by analyzing the web-application's user interface (UI) and interacting with it.

Key insight: Model application behavior to cut token cost on recurring web-agent task patterns.


Bridging Thought and Action: Taming Long-Horizon Instability in Open-Source LLM Agents with a MetaTool-Enhanced ROS Framework

Mahmud; Kazi Abrar; Dhurubo; Nilotpaul Kundu; Kirttonia; Tamal; … arXiv: 2609.13335

Large Language Models (LLMs) have enabled more natural human-robot interaction, but open-source models often exhibit unstable long-horizon reasoning and inefficient action execution when deployed in agentic robotic frameworks. This paper presents an enhanced ROS-Agent based architecture that improves task reliability and execution efficiency for agentic robotic systems using open-source LLMs.

Key insight: Bridge thought and action to tame long-horizon instability in open-source LLM agents.


Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection

Chidambaram; Keertana; Ilyas; Andrew; Syrgkanis; Vasilis arXiv: 2609.15989

Figure from Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection
Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection

Chain-of-thought (CoT) monitoring is a safety strategy where the reasoning of a large language model "actor" is inspected by a "monitor" (often another language model) for signs of unsafe planning, deception, or misalignment. We find that planting harmful but benign-sounding reasoning in the actor's context can steer it to perform adversarial actions while evading monitors, an attack we term "plan injection".

Key insight: Plan injection plants benign-sounding harmful plans that actors paraphrase while evading CoT monitors.


Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science

Lin; Honghao; Woodruff; David P.; Deng; Yuan; … arXiv: 2609.15983

Language models can produce plausible short proofs, but may still be unreliable on long-horizon research problems, where progress depends on a sequence of uncertain and interdependent decisions. We introduce Stellar Colosseum, a model-agnostic harness for allocating inference across research in mathematics and theoretical computer science.

Key insight: A many-agent harness (Stellar Colosseum) for long-horizon mathematical research collaboration.


The Router Within: Eliciting Native Skill Routing from a Frozen LLM

Chen; Ruishuo; Wang; Xun; Chen; Yu; … arXiv: 2609.15982

Skills extend an LLM agent beyond its parametric knowledge, and the gain they promise rests on picking the right one. Deployed harnesses route by preloading every skill's metadata into the context, which disperses the agent's attention and caps the library size.

Key insight: Elicit native skill routing from a frozen LLM without training a separate router.


AlgoEvo: Self-Evolving Agentic Search for Automated Algorithm Discovery

Qiu; Junhao; Hu; Qinglong; Tong; Xialiang; … arXiv: 2609.15820

Figure from AlgoEvo: Self-Evolving Agentic Search for Automated Algorithm Discovery
AlgoEvo: Self-Evolving Agentic Search for Automated Algorithm Discovery

Large language models have advanced automated algorithm discovery by synthesizing executable code, but existing frameworks trap them in rigid search pipelines with pre-defined control flows. This limitation restricts adaptive reasoning, blocks cross-paradigm transfer, and discards valuable execution feedback.

Key insight: AlgoEvo self-evolves agentic search for automated algorithm discovery.


Empirical Evaluation of Task-Based Permission Scoping Architecture for AI Agents

Noyan; Halil Burak arXiv: 2609.15422

AI agents are provisioned the same as employee-owned hosts in many enterprise settings with a static credential set fixed at deployment which includes all permissions the employee role might ever need. Role-based access control made this compromise for human principals because scoping access per task was infeasible.

Key insight: Empirical evaluation of task-based permission scoping for AI agent architectures.


SkillLift: Learning Dense Rubrics from Sparse Oracles for Efficient Skill Evolution

Kang; Haoxiang; Wen; Ming arXiv: 2609.15396

LLM-based agents increasingly rely on persistent skills, i.e., reusable procedural prompts, to adapt without weight updates. Existing skill self-evolution methods directly revise skill text based on execution feedback, but each oracle evaluation requires a full agent rollout, creating a supervision bottleneck that confines search to failure-patching updates.

Key insight: SkillLift learns dense rubrics from sparse oracles to evolve skills efficiently.


Why LLM Agents Collapse Without Oversight: The Enforcement Gap as the Mechanism Behind Emergence World Failures

Wang; Yuhang arXiv: 2609.15293

When Emergence World placed frontier LLM agents in an unsupervised multi-agent simulation, the results were alarming: agents committed crimes, starved, and enforced unanimous conformity -- without any external attacker. This paper identifies the mechanism.

Key insight: Without oversight, LLM agents collapse via an enforcement gap between intent and execution.


HazardAuditor: From Executable Threats to Safer Computer-Use Agents

Feng; Yunhao; Lin; Ruixiao; Wen; Ming; … arXiv: 2609.15134

Figure from HazardAuditor: From Executable Threats to Safer Computer-Use Agents
HazardAuditor: From Executable Threats to Safer Computer-Use Agents

Computer-use agents increasingly interact with browsers, terminals, file systems, and external services, introducing safety risks that emerge through runtime behavior rather than generated content alone. Existing guard models target static prompts and responses and are poorly suited to agent execution; existing executable safety platforms produce evaluation verdicts rather than the normalized supervision a guard model needs to learn across…

Key insight: HazardAuditor turns executable threats into safer computer-use agent behavior.


Salesforce Koa: An Enterprise Language Model for Agentic Tool Use

Chen; Zixiang; Niu; Sufeng; Liu; Yingchi; … arXiv: 2609.15066

Figure from Salesforce Koa: An Enterprise Language Model for Agentic Tool Use
Salesforce Koa: An Enterprise Language Model for Agentic Tool Use

We present Salesforce Koa, an enterprise language model built by post-training the open-weight Nemotron-3-Super-120B foundation model with reinforcement learning using Group Relative Policy Optimization (GRPO). Salesforce Koa is trained on public and synthetically generated data, with no customer data, to improve tool use and agentic capabilities while preserving strong general-purpose performance.

Key insight: Salesforce Koa post-trains Nemotron-3-Super-120B with GRPO for enterprise agentic tool use.


CoMem: Collective-Individual Memory Synergy for Evolutionary Multi-Agent Systems

Yu; Chengxin; Fan; Zhaoxin; Wu; Faguo; … arXiv: 2609.15009

Designing effective memory mechanisms is crucial for advancing LLM-driven Multi-Agent Systems (MAS), helping agents learn together and perform better over time. While recent work has led to strong cooperation skills, most methods still use flat, unstructured memories, which easily get filled with noise and erase differences between agents.

Key insight: CoMem couples collective and individual memory for evolutionary multi-agent systems.


ActGuard: Pre-execution Action Auditing against Indirect Prompt Injection in LLM Agents

Wang; Bingzheng; Gu; Xiaoyan; Wang; Wentao; … arXiv: 2609.14987

Large language model (LLM) agents interact with external environments through tool invocation, but tool outputs can also expose them to indirect prompt injection (IPI) attacks. Existing defenses mainly rely on prompt hardening, content filtering, pre-generated plans, or permission constraints.

Key insight: ActGuard audits actions before execution to blunt indirect prompt injection in tool agents.


MemRiskBench: Trace-Aware Risk-Preserving Evaluation for Long-Horizon LLM Agents

Jiang; Jianhua; Yuan; Dongbo; Li; Weihua arXiv: 2609.14976

Long-horizon LLM agents accumulate memory across sessions, creating sparse but high-impact risks: stale facts, conflicting updates, cross-user leakage, revoked-memory reuse, and constraint decay. Standard aggregate scores hide per-risk failure rates--a model achieving 78% average accuracy may still leak data in 4% of episodes--and benchmark compression preferentially discards the rare high-severity events that distinguish a mostly-working model…

Key insight: MemRiskBench evaluates long-horizon agents with trace-aware, risk-preserving metrics.


The Stochastic Deputy: Structural Tenant Isolation for Tool-Using LLM Agents

Baig; Mirza Samad Ahmed; Gillani; Syeda Anshrah; Ali; Asher; … arXiv: 2609.14780

Multi-tenant tools commonly accept a tenant identifier and validate it against the caller's entitlement. For a large language model (LLM) agent, that pattern delegates resource selection to a process whose context may contain attacker controlled instructions.

Key insight: Structural tenant isolation (Stochastic Deputy) for tool-using multi-tenant LLM agents.


MOSCOPT: Mixture-of-Skills Collective Optimization for LLM Agents

Zhang; Zhenyu; Yang; Jiudong arXiv: 2609.14399

Natural language prompts and skills serve as the strategic backbone of LLM-based agents. Recent advances in prompt and skill optimization have achieved notable gains, yet all existing methods optimize a \emph{single text template---missing the synergy among multiple complementary strategies.

Key insight: MOSCOPT mixes skills for collective optimization across LLM agents.


GraMRAG: Orchestrating Multi-Agent Multi-Step Reasoning via Graph Memory with Reinforcement Learning

Wang; Zhongyu arXiv: 2609.14066

Although existing multi-agent Retrieval-Augmented Generation (RAG) systems have demonstrated promise on complex multimodal reasoning tasks, they remain fundamentally limited in reasoning depth and memory structure, suffering from inadequate retrieval and state blindness when answering knowledge-intensive questions. To address these limitations, we propose GraMRAG, a graph memory-guided multi-agent RAG framework that integrates a dynamic…

Key insight: GraMRAG orchestrates multi-agent multi-step reasoning over graph memory with retrieval.


AGENTQ: Quantization-Conditioned Backdoor Attacks on LLM Agents

Liu; Xiaoqun; Yan; Qiben arXiv: 2609.14060

Quantization is one of the default deployment paths for open-weight LLM agents, but it is not behavior-preserving: an adversary can release a full-precision checkpoint that passes audits yet misbehaves once quantized, termed as quantization-conditioned attack (QCA). Prior QCA work targets free-text generation, where harm is mediated by a human reader.

Key insight: AGENTQ shows quantization-conditioned backdoors can compromise LLM agents.


Confuse the Model, Control the Flow: Understanding and Mitigating Privacy Leakage from LLM Agents with Information Flow Control

Shim; Minsun; Karim; Ramisha Raida; Jakkula; Ruthwik; … arXiv: 2609.14003

Personal AI agents built on large language models (LLMs) are increasingly given access to a user's private data and communications in order to provide personalized assistance. This access creates a persistent privacy risk: the agent must decide whether a given sensitive information should be disclosed to a particular party.

Key insight: Confuse the model to control flow: privacy leakage paths in tool-using agents and mitigations.


When Malicious Instructions Persist: Persistent Memory Poisoning Attack on Harness-Based Agents

Huang; Shuhuai; Zhang; Jingfeng; Jia; Hong arXiv: 2609.13889

Harness design has transformed the development of LLM-based agents by integrating memory, tool use, and runtime control. However, this design also introduces security and privacy risks because malicious instructions from external sources may be written into persistent memory and persist across sessions.

Key insight: Malicious instructions can persist via harness memory poisoning across sessions.


HarnessBandit: Joint Learnability-Transferability Scheduling for Multi-Harness Agentic Reinforcement Learning

Wei; Hongliang; Tu; Xiaobing; Wang; Yinggui; … arXiv: 2609.13739

Language-model agents are increasingly deployed through diverse harnesses that differ in system prompts, tool schemas, control loops, and trajectory formats. The same model can perform unevenly across these interfaces, making robustness to harness variation an important objective.

Key insight: HarnessBandit jointly schedules learnability and transferability across multi-harness agents.


Trustworthy Agentic AI: A Comprehensive Cybersecurity and Systems Survey on Threat Landscapes, Defense Architectures, and Open Challenges

Mostafavi; Seyedakbar arXiv: 2609.13731

The transition from passive foundation models to autonomous, goal-directed agentic AI systems has introduced unprecedented capabilities by coupling recursive cognitive reasoning loops, persistent memory architectures, live tool execution planes, and multi-agent collaboration topologies. However, granting probabilistic neural cores execution authority across filesystems, networks, and cloud infrastructure dissolves classical security perimeters:…

Key insight: Survey of threats, defenses, and systems issues for trustworthy agentic AI.


Identity Is More Than Recall: A Benchmark for Persistent Identity in Deployed AI Agents

Zhao; Zhenyu; Zhao; Roy arXiv: 2609.13637

Figure from Identity Is More Than Recall: A Benchmark for Persistent Identity in Deployed AI Agents
Identity Is More Than Recall: A Benchmark for Persistent Identity in Deployed AI Agents

Persistent agents need evaluations that distinguish identity facts they can recall from those they express and enact. We introduce PAI-Bench, a provider-neutral benchmark for fidelity to a versioned, update-governed identity contract.

Key insight: Persistent identity in deployed AI is more than recall — a dedicated benchmark.


ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

He; Jiyan; Liang; Guang; Liu; Hao; … arXiv: 2609.13356

Figure from ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

In this work, we present ZGCM-1, a fully open 7B dense foundation model trained from scratch with extreme data, system, and algorithmic efficiency. ZGCM-1 is founded on a core premise: compact models cannot passively memorize the open web, but can overcome parametric capacity limits by coupling deliberate internal thinking with active external tool use.

Key insight: ZGCM-1: fully open 7B dense foundation model trained for math and agentic search efficiency.


SkillAtlas: An Attack Trace Library for Agent Skills

Tian; Yuxin; Duan; Zenghao; Pang; Liang; … arXiv: 2609.13353

Agent skills are reusable units for language-model agents, but their risks emerge through model decisions, user context, tool calls, and execution feedback rather than through stable signatures or a single sandbox run. Existing static, dynamic, and benchmark-style evaluations rarely preserve public evidence that can be inspected, searched, and reused.

Key insight: SkillAtlas builds an attack-trace library targeting agent skills.


The Agentic Company OS: Substrate Inversion for Sustained Enterprise Agent Deployment

Larsen; Oliver Aleksander; Moghaddam; Mahyar T. arXiv: 2609.13334

Enterprise AI agents often succeed in a demonstration and then stall once they must operate day after day. An industry report estimates that most pilots never reach production and that deployed systems rarely retain feedback or improve over time, while agent benchmarks show single-run successes masking unreliable repetition.

Key insight: Agentic Company OS proposes substrate inversion for sustained enterprise agent deployment.


LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents

Gu; Zhangxuan; Chen; Haoxing; Qin; Qi; … arXiv: 2609.13287

Figure from LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents
LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents

Diffusion large language models (dLLMs) achieve high decoding efficiency through block-parallel, arbitrary-order generation, making them attractive for latency-sensitive applications. GUI agents represent a natural testbed for this paradigm, as they must repeatedly perceive screen states and emit structured, spatially grounded actions in real time.

Key insight: LLaDA-UI brings block-wise diffusion to vision-language GUI agents.


Is Bash All You Need? An Empirical Study of Tool Interfaces for Enterprise Digital Worker Agents

Mak; Hazel; Suresh; Susheel; Bhatnagar; Sahil; … arXiv: 2609.11999

In this study, we examine whether a general shell can outperform specialized tools on enterprise tasks. Shell-based agents have shown strong results in coding, but enterprise work also involves moving between applications and services, coordinating with coworkers, and performing professional analysis.

Key insight: Empirical study: is Bash enough as the tool interface for enterprise digital agents?


OrchSLM: Probing the Dynamics of Small Language Model Orchestration

Zhang; Chengxi; Yao; Yu arXiv: 2609.13470

Although large language models (LLMs) have demonstrated remarkable capabilities, their reliance on cloud-scale infrastructure poses fundamental challenges for deployment in agentic pipelines, including latency, privacy, connectivity, and substantial computational cost. Small language models (SLMs) offer a compelling alternative: recent studies suggest that many repetitive and narrowly scoped subtasks in agentic workloads may be better served by…

Key insight: OrchSLM probes the dynamics of small language model orchestration.