Monday's cs.AI announcement day (2026-09-28) lists 88 new and 113 cross-lists (replacements skipped; listing total 201). Stack filter for agent systems, memory/context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent, persistence/identity, and local/open serving keeps 32 papers — MCP↔data-space mediation, mobile GUI deeplinks, HasMem / mutable transcripts / ActKV context stacks, MoMHa harness optimization, code-based skills, executive-control failures (LLM Parkinsonism), and related stack work.
Jaime Alonso Ruiz; Carlos Aparicio; Gabriel Huecas et al. arXiv: 2609.30341
Data Spaces enable sovereign and governed data sharing across organizational boundaries, but their integration with AI agents remains challenging due to mismatches between probabilistic language model interactions and policy-driven data infrastructures. This article presents an…
Key insight: MCP can mediate between LLM agents and data spaces as an architectural boundary, not just a tool API.
Yuchen Sun; Chenglin Cai; Gongjie Zhang et al. arXiv: 2609.30887
Mobile GUI agents complete tasks using GUI actions like taps and swipes. These actions are broadly applicable across applications, but reaching a navigation interface. A single deeplink call can replace a sequence of screen-by-screen GUI actions. We therefore introduce hybrid…
Key insight: App-native deeplinks let mobile GUI agents hop to targets instead of long tap sequences.
Zihong He; Junxiao Shen; Chen Liang et al. arXiv: 2609.30797
Text-based memory and context compression support reuse of past interactions. Resizing continuous memory changes the input to a frozen LLM, coupling capacity allocation with readout. We propose Hard-Origin Adaptively Softened Memory (HasMem). Frozen hard-prompt embeddings…
Key insight: Long-term agent memory benefits from hard-origin signals with adaptive softening of writes.
Dan Barry; Andrew Hines arXiv: 2609.31354
Contemporary large language model (LLM) chat systems treat conversation history as an immutable sequence of turns that defines the model's working context. However, user intent in real interactions is not static: it evolves through correction, refinement, and shifting…
Key insight: Editable conversation state (mutable transcripts) reduces context pollution in long agent dialogues.
Zihan Wang; Cheng Tang; Lei Gong et al. arXiv: 2609.31395
Agentic LLM inference accumulates long KV caches across iterative observation-reasoning-action loops, imposing substantial memory overhead and limiting serving throughput. Existing compression methods emphasize overall output quality, overlooking the asymmetric importance of…
Key insight: Action-guided KV cache management cuts agent-loop memory cost without blind eviction.
Subhojyoti Mukherjee; Md Mehrab Tanjim arXiv: 2609.30967
Most work on improving large language models treats accuracy as the sole objective. We argue that the harness, the Python code surrounding the model that constructs prompts, routes calls, and parses outputs, is a first-class design surface whose quality is inherently…
Key insight: Harnesses can be multi-objective optimized for accuracy, safety, and token spend together.
Bartłomiej Cupiał; Jens Tuyls; Maciej Wołczyk et al. arXiv: 2609.31076
Language agents struggle to act and learn in environments that require long sequences of low-level actions. Code-based abstractions can make these agents more productive by letting them invoke reusable skills instead of repeatedly selecting individual actions. The code handles…
Key insight: Code-based skills let language agents move up and down an abstraction ladder with executable structure.
Guanyu Nie; Fangzhou Zhu; Shixiong Kai et al. arXiv: 2609.30861
Language-model agents increasingly improve by converting execution experience into reusable external skills. Yet repeated skill updates form a learning process of their own: locally useful edits can accumulate into redundant or task-specific instructions, while new updates can…
Key insight: Skill evolution needs regularization so self-improving agents do not overfit brittle skills.
Zihao Zhu; Siwei Lyu; Adel Bibi et al. arXiv: 2609.30383
A skill is a modular package of natural-language instructions, executable scripts, and reference resources that an agent can load at runtime to extend its capabilities for a specific task. Skill-based agent systems therefore enable flexible reuse of third-party capabilities,…
Key insight: Skill composition creates cascading attack surfaces even when each skill looks benign alone.
Shane Caldwell; Max Harley; Ads Dawson et al. arXiv: 2609.30325
Agents are increasingly deployed with real autonomy in web application and network penetration testing, where a single out-of-scope action can breach a client's engagement boundary. Existing offensive-security benchmarks measure raw hacking capability; as those benchmarks…
Key insight: Engagement-boundary benchmarks test whether agents stay in scope under goal pressure.
Dongsheng Xiao; Zeyuan Wang; Xuzhe Xia et al. arXiv: 2609.30662
Large language models (LLMs) can plan, use tools, write code, and execute long-horizon workflows, yet strong local competence does not guarantee project-level executive control. Agents may continue acting after the original objective is satisfied, producing low-value…
Key insight: Autonomous LM agents can fail at executive control—continuing after the objective is done—so they need a global stop layer.
Xiaoyang Li; Yiqi Wang; Chencheng Zhu et al. arXiv: 2609.30813
Evaluating claim admission in shared agent memory is challenging because repeated claims may be mistaken for independent evidence. An agent may copy or paraphrase a retrieved belief, while admitting a false claim exposes subsequent agents to it. To study this problem, we…
Key insight: Shared agent memory needs an epistemic admission gate for which beliefs may enter the store.
Jiaqi Ding; Guorong Wu arXiv: 2609.30558
Agent memory systems are increasingly used to maintain long-term user preferences, task states and evolving facts, but current evaluations often collapse memory behavior into final-answer accuracy. We introduce MemProbe, a cognitive-science-inspired framework for diagnosing…
Key insight: Stability–plasticity tradeoffs in agent memory can be probed with cognitive experimental paradigms.
Zhensheng Zou (Peking University); Guoqing Wang (Peking University); Dan Hao (Peking University) arXiv: 2609.31430
Tool observations dominate the context of software-engineering agents, making long interaction histories costly to maintain. Existing context compression methods can discard information needed by later actions, while adapting agents to soft-token representations can compromise…
Key insight: For software-engineering agents, compress latent observations (what you see), not only verbalized thoughts.
Guangzhe Zhang arXiv: 2609.31381
Context projection replaces older tool observations with compact, addressable excerpts, reducing repeated input while potentially adding evidence-retrieval turns. We study this trade-off in ReVerPi, a Pi extension with archived observations and matched full/projected…
Key insight: Selective context projection can hide capped failures when completed pairs dominate the view.
Justice Owusu Agyemang; Michael Agyare; Kwame Opuni-Boachie Obour Agyekum et al. arXiv: 2609.30293
The Model Context Protocol (MCP) enables AI agents to discover and call tools, but loading every definition becomes expensive as connected catalogs grow. We present Cartograph, a federated MCP proxy that changes agent-visible tool discovery from $O(n)$ catalog traversal to…
Key insight: Federated tool discovery with operator-attested retrieval scales safer tool catalogs for agents.
Alexander Gill; Md Farhan Ishmam; Xuyen Nguyen et al. arXiv: 2609.30604
Existing computer-use agent benchmarks do not fully evaluate agents acting as assistants. A useful assistant retrieves information across complex, multi-step workflows, synthesizes it into artifacts (documents, presentations, spreadsheets), and navigates program interfaces to…
Key insight: Web-agent hard work starts after search: synthesize, organize, and display knowledge.
Haoran Zhang; Hengtong Zhang; Zhiyu Liang et al. arXiv: 2609.31301
Large language model agents increasingly act on software systems, no longer merely generating text but also changing databases and online services. However, an approved database update may succeed yet leave an unapproved notification because execution can produce persistent…
Key insight: Runtime validation should check persistent outcomes, not only approved tool actions.
Yiran Hu; Nan Jiang; Shanchao Liang et al. arXiv: 2609.30725
Although effective, coding agents often incur substantial monetary costs. Their recurring cost-inefficient behaviors remain underexplored. We conduct the first study of behavioral cost inefficiencies in coding agents, analyzing 1,200 trajectories from Claude Code and…
Key insight: Coding agents exhibit measurable cost-inefficient behaviors that harnesses can mitigate.
Md Shohel Arman; Igor Molybog arXiv: 2609.31587
We investigate whether natural-language documentation helps coding agents resolve software issues, and we build the tools to construct and evaluate it. We introduce a roundtrip benchmark that scores code descriptions by whether code regenerated from them passes the original…
Key insight: Compact documentation for coding agents can be optimized, but gains often fail to transfer across stacks.
Leon Goldberg; Gal Engelberg; Eden Yavin et al. arXiv: 2609.30345
Cloud-security investigation is dominated by population tasks: which identities can read a data store, how many resources fail a control, which assets are reachable from another account. These resolve against a complete inventory, not a named object. A partial answer to one is…
Key insight: Enterprise security investigations need more than coding agents—an agentic security brain over cloud evidence.
Salma Roshdy Aly; Hussein Assaf; Ziad Kobti arXiv: 2609.30328
When one language model judges whether another's code is correct, it does not report the absence of evidence. It returns a confident verdict with reasoning attached, indistinguishable from a verdict it had grounds for. Multi-agent verification, which decomposes a judgment into…
Key insight: Multi-agent code judges need grounding checks and the option to decline when evidence is thin.
Jinfeng Xu; Zheyu Chen; Ziyue Peng et al. arXiv: 2609.30734
Multi-agent LLM workflows use planning, execution, verification, and summarization to improve task performance, yet the value of each component depends on the state already produced. Executing every component can waste computation or overwrite a correct intermediate answer. We…
Key insight: Counterfactual credit assignment teaches multi-agent LLM workflows what steps to skip.
Jakub Masłowski; Jarosław A. Chudziak arXiv: 2609.31422
Large language model-based multi-agent debate (MAD) systems are being increasingly used as complex decision pipelines in distributed processes. However, their final synthesis phase still remains inadequately controlled. Even with detailed debate logs, summarizing models are…
Key insight: An active provenance gate can blunt fabricated consensus in multi-agent debate synthesis.
Weida Liang; Shi Qiu; Zhun Wang et al. arXiv: 2609.31318
AI agents combine language models with external data and tools that can modify files, call APIs, or execute code. Security failures can arise when adversarial content changes an agent's tool use or when the surrounding software contains vulnerabilities such as path traversal or…
Key insight: Repository-to-runtime red-teaming automates finding agent stack vulnerabilities end-to-end.
Minghui Yu; Ke Mu; Gang Wu arXiv: 2609.30836
Deploying small language models (SLMs) on offline, resource-constrained edge devices such as remote sensing satellites presents a fundamental challenge: their limited reasoning capacity hinders reliable execution of multi-step agent tasks requiring complex tool orchestration.…
Key insight: Offline edge SLMs need architecture-level decoding help for multi-step tool orchestration.
Junyi Shen; Noppanat Wadlom; Zhengyuan Su et al. arXiv: 2609.31047
Agentic LLM workflows decide their execution paths at runtime. Downstream computation may be predictable, or may have run before, yet it cannot begin until the model or the user resolves the branch. We call this serialization the branch-resolution barrier. Caching alone does…
Key insight: Speculative subgraph reuse speeds dynamic agentic LLM serving under branching tool loops.
Mukul Chhabra; Shail Patel; Luigi Medrano arXiv: 2609.30471
Reference-based LLM-as-a-judge evaluation assumes the reference answer is the target. In deployed agentic systems that operate over dynamic entities (support cases, assets, accounts), the closest available reference typically applies the correct procedure to a different entity,…
Key insight: Production agentic eval can gate retrieval by context awareness (CARGO).
Seungho Lee; Changbin Lee arXiv: 2609.30614
Dataspace connectors decide whether a transfer may occur, not what the transferred value contains, tolerable for contracted applications, not for LLM agents that compose tool calls and spawn sub-agents. Research on agents that generate governance artifacts evaluates output…
Key insight: Agentic dataspaces create authorship hazards when agents write as if they were human authors.
Chang Gong; Jingping Bi; Di Yao et al. arXiv: 2609.31186
Artificial intelligence is advancing rapidly, with increasingly capable systems taking larger roles in reasoning, decision-making, scientific discovery, and autonomous development. As AI begins to participate in its own improvement, from model training and experience…
Key insight: Recursive self-improvement needs an evolutionary safety taxonomy before harnesses self-modify unchecked.
Maokai Qin; Chuan Qin; Qi Zhang et al. arXiv: 2609.30971
Embodied agents offer a promising route to automating scientific experimentation, yet their progress is constrained by the lack of reliable and systematic evaluation environments. Existing simulation-based laboratory benchmarks rely heavily on manual task engineering, making it…
Key insight: Protocol-to-task compilers can scale benchmarks for scientific embodied agents.
Joseph Sifakis arXiv: 2609.30291
The purpose of this article is to highlight the central role of autonomous systems as the ultimate stage in the development of AI, to explain the underlying technical challenges that require a combination of connectionist AI and symbolic AI, and to integrate AI and systems…
Key insight: Autonomous-system design still centers long-term memory and collective coordination around agent cognition.