Tuesday’s cs.AI announcement day lists 147 new and 261 cross-lists (replacements skipped; listing total 408). Stack filter for agent systems, memory/context, computer-use / GUI / tools / skills / harnesses, multi-agent, persistence/identity, and local/open models keeps 40 papers — a dense but readable day spanning tool boundaries, stateful handoffs, plan-injection attacks, and open models for agentic search.
Trofimov; Artem; Novikov; Boris arXiv: 2609.15397
AI agents increasingly execute long-running workflows that externalize effects through independently supplied tools. Under retries, speculative execution, concurrency, and partial failures, the resulting external state may be inconsistent with the workflow's intended resolution: required effects may be missing or duplicated, aborted effects may survive, and committed effects may depend on provisional state that is later withdrawn.
Key insight: Tool-call success is not workflow success — external state can diverge under retries and partial failure.
Zhu; Sibo; Fan; Shicheng; Wang; Xinyue; … arXiv: 2609.15364
Digital agents must often adapt to new environments whose interfaces, tools, and failure modes are not fully captured by pretrained models. We introduce RSIAgent, a training-free multi-agent framework for recursive self-improvement through autonomous memory construction.
Key insight: Training-free recursive self-improvement via autonomous memory construction for unfamiliar environments.
Peng; Yanwen; Zhang; Delvin Ce; Wang; Xi; … arXiv: 2609.15054
Multi-Agent (MA) systems are effective at solving complex tasks that demand planning, tool use, and the synthesis of evidence from multiple sources. Existing systems typically adopt Hierarchical Manager-Worker (HMW) or Router-based Message Passing (RMP) structures as their communication protocol.
Key insight: A shared bus substrate can replace hierarchical manager-worker or router message passing in multi-agent systems.
Li; Yu; Cai; Qikun; Huang; Tao; … arXiv: 2609.15011
Memory-augmented and tool-using agents expose exact private values when remote LLMs process retrieved memory, tool actions, and intermediate observations. One-way masking limits direct exposure but removes values needed for trusted execution and can leak them through later observations.
Key insight: Structure-preserving trustworthy virtual memory keeps private values usable without raw exposure to remote LLMs.
Agachan; Burak; Duijn; Max van; Zohrehvand; Amirhossein arXiv: 2609.14767
Hierarchical orchestration, in which a Manager agent reviews worker output and can send it back for revision, is the default coordination pattern in production multi-agent LLM frameworks. Classical organizational theory predicts that the authority link speeds convergence on decisive output; work on sycophancy and Degeneration-of-Thought predicts that authoritative critique makes LLM output worse.
Key insight: Loop-back authority in manager-worker teams can hurt LLM output despite classical org-theory predictions.
Zhu; Genliang arXiv: 2609.14744
By acquiring compute, credentials, accounts, services, and other agents, autonomous AI agents can introduce new authority into a task. Payment, budget, OAuth, mandate, and fulfillment checks can validate transaction conditions without deciding whether a returned resource may become usable authority.
Key insight: Acquired compute, credentials, and services need runtime authorization before they become usable agent authority.
Sharma; Siddharth; Pandey; Nilesh Prasad; Gungor; Onat; … arXiv: 2609.14138
As LLM agents become integrated into increasingly complex workflows, they must continually acquire new capabilities while retaining competence on previously learned tasks. Lifelong agents address this through experience replay, injecting past interactions into the prompt to leverage prior experience during inference.
Key insight: Lifelong agents need joint memory and budget optimization at inference time, not only experience replay.
Deng; Runzhi; Zhong; Yiming; Zhao; Fang; … arXiv: 2609.13800
Routing and cascades reduce tool-agent cost by transferring control across models, but stateful handoffs must preserve accepted choices, realized effects, and unfinished obligations. We formulate this as commitment-constrained residual completion and introduce Commitment-Frontier Residual Completion (CFRC).
Key insight: Stateful handoffs should complete residual obligations instead of restarting from a clean slate.
Zhang; Zhihui; Liu; Wei arXiv: 2609.13672
AI agents can be interrupted while editing files, calling tools, or carrying out multi-step tasks. Restarting repeats completed work, but continuing from unverified or outdated progress can carry earlier errors forward.
Key insight: Treat recoverability as a system primitive for long-horizon agents interrupted mid-tool or mid-edit.
Luo; Han; Xu; Xian; Liu; Yinhe; … arXiv: 2609.13667
Geospatial agents are increasingly expected to support recurring and evolving analytical tasks rather than execute isolated workflows. In such settings, effective agents must distill prior execution experience into reusable geospatial procedural knowledge to guide future planning and tool use.
Key insight: Hierarchical skill learning with collaborative revision for recurring geospatial agent workflows.
Cao; Xinyun; Szekeres; Adriana; Faisal; Fazle Elahi arXiv: 2609.13548
Web agents can utilize reusable tools to reduce the cost and latency of low-level browser interaction, but automatically discovered tool collections can be large, redundant, and poorly aligned with user demand. We present AutoTailor, a meta-agentic framework for constructing and maintaining a compact set of trajectory-derived Model Context Protocol (MCP) APIs.
Key insight: AutoTailor selects and adapts web-agent capabilities to stay aligned with user intent.
Ianta; Alexandru; Stroulia; Eleni arXiv: 2609.13491
The strong performance of AI Agents across an impressive variety of tasks is driving an unprecedented investment in agentic infrastructures, however the cost of processing tokens is fast increasing. Web agents automate the execution of web-application tasks described in natural language, by analyzing the web-application's user interface (UI) and interacting with it.
Key insight: Model application behavior to cut token cost on recurring web-agent task patterns.
Mahmud; Kazi Abrar; Dhurubo; Nilotpaul Kundu; Kirttonia; Tamal; … arXiv: 2609.13335
Large Language Models (LLMs) have enabled more natural human-robot interaction, but open-source models often exhibit unstable long-horizon reasoning and inefficient action execution when deployed in agentic robotic frameworks. This paper presents an enhanced ROS-Agent based architecture that improves task reliability and execution efficiency for agentic robotic systems using open-source LLMs.
Key insight: Bridge thought and action to tame long-horizon instability in open-source LLM agents.
Chidambaram; Keertana; Ilyas; Andrew; Syrgkanis; Vasilis arXiv: 2609.15989
Chain-of-thought (CoT) monitoring is a safety strategy where the reasoning of a large language model "actor" is inspected by a "monitor" (often another language model) for signs of unsafe planning, deception, or misalignment. We find that planting harmful but benign-sounding reasoning in the actor's context can steer it to perform adversarial actions while evading monitors, an attack we term "plan injection".
Key insight: Plan injection plants benign-sounding harmful plans that actors paraphrase while evading CoT monitors.
Lin; Honghao; Woodruff; David P.; Deng; Yuan; … arXiv: 2609.15983
Language models can produce plausible short proofs, but may still be unreliable on long-horizon research problems, where progress depends on a sequence of uncertain and interdependent decisions. We introduce Stellar Colosseum, a model-agnostic harness for allocating inference across research in mathematics and theoretical computer science.
Key insight: A many-agent harness (Stellar Colosseum) for long-horizon mathematical research collaboration.
Chen; Ruishuo; Wang; Xun; Chen; Yu; … arXiv: 2609.15982
Skills extend an LLM agent beyond its parametric knowledge, and the gain they promise rests on picking the right one. Deployed harnesses route by preloading every skill's metadata into the context, which disperses the agent's attention and caps the library size.
Key insight: Elicit native skill routing from a frozen LLM without training a separate router.
Qiu; Junhao; Hu; Qinglong; Tong; Xialiang; … arXiv: 2609.15820
Large language models have advanced automated algorithm discovery by synthesizing executable code, but existing frameworks trap them in rigid search pipelines with pre-defined control flows. This limitation restricts adaptive reasoning, blocks cross-paradigm transfer, and discards valuable execution feedback.
Key insight: AlgoEvo self-evolves agentic search for automated algorithm discovery.
Noyan; Halil Burak arXiv: 2609.15422
AI agents are provisioned the same as employee-owned hosts in many enterprise settings with a static credential set fixed at deployment which includes all permissions the employee role might ever need. Role-based access control made this compromise for human principals because scoping access per task was infeasible.
Key insight: Empirical evaluation of task-based permission scoping for AI agent architectures.
Kang; Haoxiang; Wen; Ming arXiv: 2609.15396
LLM-based agents increasingly rely on persistent skills, i.e., reusable procedural prompts, to adapt without weight updates. Existing skill self-evolution methods directly revise skill text based on execution feedback, but each oracle evaluation requires a full agent rollout, creating a supervision bottleneck that confines search to failure-patching updates.
Key insight: SkillLift learns dense rubrics from sparse oracles to evolve skills efficiently.
Wang; Yuhang arXiv: 2609.15293
When Emergence World placed frontier LLM agents in an unsupervised multi-agent simulation, the results were alarming: agents committed crimes, starved, and enforced unanimous conformity -- without any external attacker. This paper identifies the mechanism.
Key insight: Without oversight, LLM agents collapse via an enforcement gap between intent and execution.
Feng; Yunhao; Lin; Ruixiao; Wen; Ming; … arXiv: 2609.15134
Computer-use agents increasingly interact with browsers, terminals, file systems, and external services, introducing safety risks that emerge through runtime behavior rather than generated content alone. Existing guard models target static prompts and responses and are poorly suited to agent execution; existing executable safety platforms produce evaluation verdicts rather than the normalized supervision a guard model needs to learn across…
Key insight: HazardAuditor turns executable threats into safer computer-use agent behavior.
Chen; Zixiang; Niu; Sufeng; Liu; Yingchi; … arXiv: 2609.15066
We present Salesforce Koa, an enterprise language model built by post-training the open-weight Nemotron-3-Super-120B foundation model with reinforcement learning using Group Relative Policy Optimization (GRPO). Salesforce Koa is trained on public and synthetically generated data, with no customer data, to improve tool use and agentic capabilities while preserving strong general-purpose performance.
Key insight: Salesforce Koa post-trains Nemotron-3-Super-120B with GRPO for enterprise agentic tool use.
Yu; Chengxin; Fan; Zhaoxin; Wu; Faguo; … arXiv: 2609.15009
Designing effective memory mechanisms is crucial for advancing LLM-driven Multi-Agent Systems (MAS), helping agents learn together and perform better over time. While recent work has led to strong cooperation skills, most methods still use flat, unstructured memories, which easily get filled with noise and erase differences between agents.
Key insight: CoMem couples collective and individual memory for evolutionary multi-agent systems.
Wang; Bingzheng; Gu; Xiaoyan; Wang; Wentao; … arXiv: 2609.14987
Large language model (LLM) agents interact with external environments through tool invocation, but tool outputs can also expose them to indirect prompt injection (IPI) attacks. Existing defenses mainly rely on prompt hardening, content filtering, pre-generated plans, or permission constraints.
Key insight: ActGuard audits actions before execution to blunt indirect prompt injection in tool agents.
Jiang; Jianhua; Yuan; Dongbo; Li; Weihua arXiv: 2609.14976
Long-horizon LLM agents accumulate memory across sessions, creating sparse but high-impact risks: stale facts, conflicting updates, cross-user leakage, revoked-memory reuse, and constraint decay. Standard aggregate scores hide per-risk failure rates--a model achieving 78% average accuracy may still leak data in 4% of episodes--and benchmark compression preferentially discards the rare high-severity events that distinguish a mostly-working model…
Key insight: MemRiskBench evaluates long-horizon agents with trace-aware, risk-preserving metrics.
Baig; Mirza Samad Ahmed; Gillani; Syeda Anshrah; Ali; Asher; … arXiv: 2609.14780
Multi-tenant tools commonly accept a tenant identifier and validate it against the caller's entitlement. For a large language model (LLM) agent, that pattern delegates resource selection to a process whose context may contain attacker controlled instructions.
Key insight: Structural tenant isolation (Stochastic Deputy) for tool-using multi-tenant LLM agents.
Zhang; Zhenyu; Yang; Jiudong arXiv: 2609.14399
Natural language prompts and skills serve as the strategic backbone of LLM-based agents. Recent advances in prompt and skill optimization have achieved notable gains, yet all existing methods optimize a \emph{single text template---missing the synergy among multiple complementary strategies.
Key insight: MOSCOPT mixes skills for collective optimization across LLM agents.
Wang; Zhongyu arXiv: 2609.14066
Although existing multi-agent Retrieval-Augmented Generation (RAG) systems have demonstrated promise on complex multimodal reasoning tasks, they remain fundamentally limited in reasoning depth and memory structure, suffering from inadequate retrieval and state blindness when answering knowledge-intensive questions. To address these limitations, we propose GraMRAG, a graph memory-guided multi-agent RAG framework that integrates a dynamic…
Key insight: GraMRAG orchestrates multi-agent multi-step reasoning over graph memory with retrieval.
Liu; Xiaoqun; Yan; Qiben arXiv: 2609.14060
Quantization is one of the default deployment paths for open-weight LLM agents, but it is not behavior-preserving: an adversary can release a full-precision checkpoint that passes audits yet misbehaves once quantized, termed as quantization-conditioned attack (QCA). Prior QCA work targets free-text generation, where harm is mediated by a human reader.
Key insight: AGENTQ shows quantization-conditioned backdoors can compromise LLM agents.
Shim; Minsun; Karim; Ramisha Raida; Jakkula; Ruthwik; … arXiv: 2609.14003
Personal AI agents built on large language models (LLMs) are increasingly given access to a user's private data and communications in order to provide personalized assistance. This access creates a persistent privacy risk: the agent must decide whether a given sensitive information should be disclosed to a particular party.
Key insight: Confuse the model to control flow: privacy leakage paths in tool-using agents and mitigations.
Huang; Shuhuai; Zhang; Jingfeng; Jia; Hong arXiv: 2609.13889
Harness design has transformed the development of LLM-based agents by integrating memory, tool use, and runtime control. However, this design also introduces security and privacy risks because malicious instructions from external sources may be written into persistent memory and persist across sessions.
Key insight: Malicious instructions can persist via harness memory poisoning across sessions.
Wei; Hongliang; Tu; Xiaobing; Wang; Yinggui; … arXiv: 2609.13739
Language-model agents are increasingly deployed through diverse harnesses that differ in system prompts, tool schemas, control loops, and trajectory formats. The same model can perform unevenly across these interfaces, making robustness to harness variation an important objective.
Key insight: HarnessBandit jointly schedules learnability and transferability across multi-harness agents.
Mostafavi; Seyedakbar arXiv: 2609.13731
The transition from passive foundation models to autonomous, goal-directed agentic AI systems has introduced unprecedented capabilities by coupling recursive cognitive reasoning loops, persistent memory architectures, live tool execution planes, and multi-agent collaboration topologies. However, granting probabilistic neural cores execution authority across filesystems, networks, and cloud infrastructure dissolves classical security perimeters:…
Key insight: Survey of threats, defenses, and systems issues for trustworthy agentic AI.
Zhao; Zhenyu; Zhao; Roy arXiv: 2609.13637
Persistent agents need evaluations that distinguish identity facts they can recall from those they express and enact. We introduce PAI-Bench, a provider-neutral benchmark for fidelity to a versioned, update-governed identity contract.
Key insight: Persistent identity in deployed AI is more than recall — a dedicated benchmark.
He; Jiyan; Liang; Guang; Liu; Hao; … arXiv: 2609.13356
In this work, we present ZGCM-1, a fully open 7B dense foundation model trained from scratch with extreme data, system, and algorithmic efficiency. ZGCM-1 is founded on a core premise: compact models cannot passively memorize the open web, but can overcome parametric capacity limits by coupling deliberate internal thinking with active external tool use.
Key insight: ZGCM-1: fully open 7B dense foundation model trained for math and agentic search efficiency.
Tian; Yuxin; Duan; Zenghao; Pang; Liang; … arXiv: 2609.13353
Agent skills are reusable units for language-model agents, but their risks emerge through model decisions, user context, tool calls, and execution feedback rather than through stable signatures or a single sandbox run. Existing static, dynamic, and benchmark-style evaluations rarely preserve public evidence that can be inspected, searched, and reused.
Key insight: SkillAtlas builds an attack-trace library targeting agent skills.
Larsen; Oliver Aleksander; Moghaddam; Mahyar T. arXiv: 2609.13334
Enterprise AI agents often succeed in a demonstration and then stall once they must operate day after day. An industry report estimates that most pilots never reach production and that deployed systems rarely retain feedback or improve over time, while agent benchmarks show single-run successes masking unreliable repetition.
Key insight: Agentic Company OS proposes substrate inversion for sustained enterprise agent deployment.
Gu; Zhangxuan; Chen; Haoxing; Qin; Qi; … arXiv: 2609.13287
Diffusion large language models (dLLMs) achieve high decoding efficiency through block-parallel, arbitrary-order generation, making them attractive for latency-sensitive applications. GUI agents represent a natural testbed for this paradigm, as they must repeatedly perceive screen states and emit structured, spatially grounded actions in real time.
Key insight: LLaDA-UI brings block-wise diffusion to vision-language GUI agents.
Mak; Hazel; Suresh; Susheel; Bhatnagar; Sahil; … arXiv: 2609.11999
In this study, we examine whether a general shell can outperform specialized tools on enterprise tasks. Shell-based agents have shown strong results in coding, but enterprise work also involves moving between applications and services, coordinating with coworkers, and performing professional analysis.
Key insight: Empirical study: is Bash enough as the tool interface for enterprise digital agents?
Zhang; Chengxi; Yao; Yu arXiv: 2609.13470
Although large language models (LLMs) have demonstrated remarkable capabilities, their reliance on cloud-scale infrastructure poses fundamental challenges for deployment in agentic pipelines, including latency, privacy, connectivity, and substantial computational cost. Small language models (SLMs) offer a compelling alternative: recent studies suggest that many repetitive and narrowly scoped subtasks in agentic workloads may be better served by…
Key insight: OrchSLM probes the dynamics of small language model orchestration.