Thursday’s cs.AI listing had 150 new submissions and cross-lists. This page keeps the 32 that are about agent systems: memory and context, harnesses and skills, tool use and computer-use, multi-agent coordination, persistent identity, and local open-weight models aimed at that stack. Out-of-stack bio, medical, climate, quantum, generic eval, and weak keyword hits are omitted.
Subagents vs Agent Skills compares reusable skill packages against subagent execution for long-horizon tasks. Kernel-managed shared memory proposes system-wide personalization across agents. JarvisGUI targets cross-device GUI workflows with dynamic composition. AgentHijack and multimodal prompt-injection studies stress computer-use trust boundaries. What Should an Agent Forget? separates storage from use; Fortunate Recall adds ontology-driven lifecycle management. RobustSGPO evolves harnesses under search-space control; AgentAudit frames full-lifecycle trust evaluation. Eight papers have a figure extracted from the HTML/PDF.
Ryan Lum; Yongfeng Zhang arXiv: 2609.10144
AI systems become more useful when they can adapt to the people using them, but in multi-agent systems, useful context learned by one agent often remains unavailable to others. We present kernel-managed shared memory, a system-level abstraction in which specialized agents write structured, tagged memories while the agent-system kernel, not individual agents, governs retrieval, privacy enforcement, and prompt injection.
Key insight: Share personalization context across agents via kernel-managed memory.
Wasu Top Piriyakulkij; Rachel Lawrence; Alicia Curth et al. arXiv: 2609.09233
How can language model agents effectively leverage libraries of reusable knowledge to solve long-horizon tasks? Recent work has increasingly focused on agent skills: reusable capabilities represented as skill packages, i.e., multi-file bundles containing instructions, scripts, and other resources that help agents perform specific tasks.
Key insight: Choose subagents vs skill packages for reusable long-horizon knowledge.
Liang Qu; Jianxin Li; Hua Wang arXiv: 2609.09565
Agentic graph learning (AGL) has recently achieved promising results on graph reasoning tasks, where an agent powered by a large language model (LLM) sequentially samples the graph as evidence to support its final prediction. Existing methods either employ a single agent or orchestrate multiple role-based agents to reason and learn over the entire graph, but both essentially rely on a shared reasoning policy across different graph regions, which can be suboptimal for graph...
Key insight: Learn multi-agent structure via graph signatures, not flat prompts.
Divyanshu Kumar; Nitin Aravind Birur; Tanay Baswa et al. arXiv: 2609.09647
Agentic systems are rapidly moving to production, where they read untrusted inputs, call tools with real permissions, and act autonomously, expanding the security surface beyond chat-only models. Yet standard evaluations remain single-turn and fail to capture multi-step agent vulnerabilities.
Key insight: Red-team agentic systems with a taxonomy-driven risk discovery loop.
Yanze Cao arXiv: 2609.09774
Procedural memory lets language agents reuse successful routines, but reuse presumes that a stored routine remains applicable. We study what happens when that presumption is deliberately violated.
Key insight: Measure reuse and interference when procedural memory changes under web tasks.
Shrey Nag; Sachita; Abhishek Kumar Singh et al. arXiv: 2609.09875
Existing evaluation frameworks mostly assess only one part of AI agents, such as task completion (AgentBench) or security robustness (AgentDojo, ASB), rather than the complete pipeline of planning, tool selection, tool execution, memory and reasoning. Failures can occur at any stage, yet existing benchmarks rarely identify their precise source.
Key insight: Evaluate agent trust across the full lifecycle, not a single score.
Priyanka Mary Mammen; Emil Joswin; Srujananjali Medicherla arXiv: 2609.09448
As agentic systems getting adopted rapidly in safety critical applications, it is vital to measure the confidence associated with the agentic actions. In comparison to the traditional machine learning systems, agentic workflows have complex failure modes with planning, tool invocation and dynamic environment interactions.
Key insight: Calibrate agent confidence from internal representations of success.
Zibo Zhao; Jijun Shi; Mo Zhou et al. arXiv: 2609.09646
Semantic-gradient-based prompt optimization (SGPO) improves agent harnesses using execution feedback, but its local update rule leaves the choice of edit scope and operation unresolved. We introduce RobustSGPO, which specifies the requested edit, constructs and checks the patch, and continues search from either the incumbent or retained snapshots.
Key insight: Evolve agent harnesses with controlled search, not unbounded mutation.
Xing Zhang; Guanghui Wang; Yanwei Cui et al. arXiv: 2609.09815
Compound LLM systems often solve a coordination problem by adding a higher-level LLM. The resulting meta-agent reads workers' outputs, writes the final answer, allocates later calls, and decides when to stop.
Key insight: Manage compound LLM systems with merge operators, not another model.
Yuhang Li; Yuchen Li arXiv: 2609.10263
Persistent language agents need stored experience to remain available across time, while each answer requires evidence suited to a particular question. A superseded fact can mislead a current-state answer and still be essential for a historical query.
Key insight: Separate stored memory from what the agent should use next.
Ansuman Mullick; Eray Tüzün arXiv: 2609.10413
Current LLM memory systems treat all personal facts identically, so stores grow without bound while retrieval precision degrades. The core challenge is lifecycle management: which memories should persist, which should be replaced, and at what rate, conditioned on the behavioral type of each fact.
Key insight: Manage memory lifecycle with explicit retention and coherence rules.
Zixiang Chen; Yuheng Lu; Zihao Cheng et al. arXiv: 2609.10451
Real-world GUI usage frequently involves workflows that span multiple devices and platforms, requiring the transfer of intermediate results, maintenance of shared state, and coordination across heterogeneous environments. However, existing GUI benchmarks overwhelmingly evaluate agents on single-device, statically defined tasks, thus leaving such cross-device capabilities largely unexamined, resulting in an overly optimistic assessment of agents' readiness for real-world us...
Key insight: Compose GUI tasks across devices with shared intermediate state.
Bo Yan; Weikai Lin; Song Wang arXiv: 2609.09395
Language models act through tools, yet practical agents face libraries containing thousands of interfaces. We introduce the tool menu as the short, ordered subset of available tools shown to an agent before execution.
Key insight: Treat tool menus as execution priors for online agents.
Praphul Singh; Shanu Kumar; Akshat Agarwal; Ganesh Kumar arXiv: 2609.09458
As LLM agents move from answering questions to carrying out procedures, failures can be unwarranted rather than visibly wrong: the final response looks acceptable even though the system skipped the check, branch, dependency, or invariant that made the answer justified. Output-only evaluation sees the answer, and trace-aware judging sees activity, but neither identifies which obligations were active for the query.
Key insight: Match procedural instruction conformance with query-conditioned execution.
Haoran Gao; An Li; Zhen Li; Jun Cai arXiv: 2609.09625
As Digital Twin (DT) systems evolve beyond state synchronization toward task-oriented and knowledge-driven operation, Cognitive Digital Twins (CDTs) have emerged as an extension that incorporates cognitive capabilities into twin operation. Existing CDT studies often focus on specific enabling techniques, such as learning modules, knowledge graphs, and large language models, while providing limited insight into how cognition can be systematically integrated into DT architec...
Key insight: Move from state sync to cognitive self-evolution for coherent agents.
Hyojeong Yu; Hyukhun Koh; Minsung Kim et al. arXiv: 2609.09664
Large language models (LLMs) are increasingly deployed as personalized assistants that interact with users over extended periods of time. As conversations grow longer, relying on full interaction histories becomes increasingly inefficient and unreliable: long contexts introduce substantial computational overhead, making it difficult for models to consistently identify and utilize the most relevant information for the current request.
Key insight: Compose GUI tasks across devices with shared intermediate state.
Benjamin Gruenbaum; Doron Porat; Assaf Natanzon et al. arXiv: 2609.09853
LLM agents for enterprise systems of record cannot be evaluated on customer production data, and no existing substitute provides ground truth. We present the Era by Eon Benchmark for evaluating LLM agents that use enterprise tools.
Key insight: Benchmark agents on generated enterprise estates with exact ground truth.
Arnab Chattopadhayay; Debdipta Halder arXiv: 2609.10036
Large language model agents produce fluent action sequences across a wide range of tasks, yet they fail in characteristic ways once the environment becomes partially observable. Ambiguous feedback pushes them into premature commitments.
Key insight: Augment LLMs with belief-state engines for planning under partial observability.
Rui Sun; Zhan Shi; Bing He arXiv: 2609.10315
Reinforcement learning with verifiable rewards (RLVR) has advanced language-model reasoning in domains such as mathematics and code, where objective answers are inexpensive to check. Diagnostic reasoning over complex data lacks this advantage: establishing the true cause of an anomaly often requires costly expert investigation and may remain ambiguous after the fact.
Key insight: Train reasoning agents for causal exploration with synthesized rewards.
Hongming Zhang; Zhaozhen Gu; Fengshuo Bai et al. arXiv: 2609.10441
While Large Language Models (LLMs) have demonstrated impressive capabilities, they often struggle with extremely long contexts due to fixed context limits. To address this, sequential approaches like MemAgent extend the effective context by reading text in segments and iteratively updating a fixed-size memory.
Key insight: Use convolutional memory for long-context reasoning without full attention blowups.
Zhihao Liu; Hongyu Sun; Zhiyuan Fu et al. arXiv: 2609.09212
This paper presents an end-to-end evaluation framework for image-triggered command injection against computer-use agents (CUAs). The goal is to test whether a local visual patch can induce verifiable environmental consequences along the full chain of screenshot input, VLM generation, action parsing, and environment execution.
Key insight: Harden computer-use agents against visual and multimodal input attacks.
Hamed Jafarzadeh Asl; Yuanhao Yu; Vahid Partovi Nia arXiv: 2609.09476
In-vehicle assistants must translate natural-language requests into accurate vehicle function calls under strict memory and latency constraints, making small language models (SLMs) attractive for on-device deployment. For such models, a key design choice is how the available function surface is presented.
Key insight: Move small models from fixed keys to readable schemas for agent function calls.
Yanzhe Chen; Zechen Bai; Zhijun Cao et al. arXiv: 2609.10522
Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to "play" robots through a compact semantic interface linking intent to action.
Key insight: Evolve agent harnesses with controlled search, not unbounded mutation.
Viet K. Nguyen; Mohammad I. Husain arXiv: 2609.09404
Agentic AI frameworks let a language model plan, keep memory, and call tools that reach real files, mail, and services. Most of these agents also read images, which gives an attacker a way to put text into the agent's context without going through the user.
Key insight: Harden computer-use agents against visual and multimodal input attacks.
Mark Marron; Earl T. Barr arXiv: 2609.10248
Traditional software delivery assumes a static paradigm: code is constructed prior to execution and deployed as a fixed artifact. We present Agentic Just-In-Time Software Construction (A-JIT), a paradigm that replaces static binaries with dynamic, software systems that can perpetually evolve to meet changing demands.
Key insight: Construct software agentically just-in-time instead of static scaffolds alone.
Syed Ghazanfar Abbas; Dongyan Xu arXiv: 2609.03247
Large language model (LLM) security has largely focused on role-playing jailbreaks, with less attention to what happens when a user asks an LLM to verify an identity claim through a test designed by the model itself. We study this behavior through a staged developer-identity experiment with ChatGPT, Claude, Qwen, Mistral, and Llama.
Key insight: Watch self-issued auth patterns when agents act as developers.
Iliano Fasolino arXiv: 2609.09243
Retrieval-augmented generation (RAG) grounds a language model in retrieved documents, which reduces hallucination but creates a new attack surface: if retrieved text is tampered with, the model may repeat the falsehood. We study how much a small quantized model, Llama 3.1 8B, degrades when a fraction of its retrieved context is poisoned.
Key insight: Stress-test RAG under poisoned or adversarial documents.
Jingjie Ning; Shanshan Zhong; Xiaochuan Li; Ji Zeng arXiv: 2609.09219
AI research agents combine prior knowledge, public sources, and experimental feedback to produce useful results. The Discovery Certification Protocol (DCP) turns claims about these results into executable recovery and feedback tests.
Key insight: Evaluate agent trust across the full lifecycle, not a single score.
Kevin Hartman arXiv: 2609.09671
When an agent writes code, the development framework becomes the control system for a non-deterministic worker. Spec-first, agent-driven frameworks have gained rapid traction since 2025; the installable ones, GitHub Spec Kit, obra/superpowers, BMAD, and GSD, and our own, all capture intent through a specification or durable planning artifacts.
Key insight: Enforce spec-first, test-driven agent development on coding tasks.
Fumihiko Tachibana; Daisuke Miyashita; Jun Deguchi arXiv: 2609.09768
In Retrieval-Augmented Generation (RAG) systems, a large number of retrieved chunks are concatenated to form the input context so that users can receive high-quality responses based on external knowledge. As a result, the input context length increases substantially, leading to a larger prefill workload and, in turn, a longer time to first token (TTFT).
Key insight: Combine KV-cache-aware fine-tuning with selective recomputation.
Harang Ju; Sinan Aral arXiv: 2609.09789
Organizational design in the era of artificial intelligence requires experimental methods that can test how human-AI groups coordinate, delegate, and make decisions. Programmable platforms coordinate live human-to-human sessions or real-time human-AI chat, but researchers cannot easily declare experiment protocols in which AI participants both communicate and act on shared work within one auditable configuration.
Key insight: Run live human-AI collaboration experiments on a shared platform.
Daniel Alejandro Coll Tejeda; Pedro García López; Daniel Barcelona-Pons arXiv: 2609.10239
Graph-based retrieval can improve multi-hop question answering, but existing approaches often incur high query-time costs and produce diffuse, oversized contexts that reduce generation efficiency. We present LiteRAG, a graph-based retrieval method that replaces expensive retrieval-time LLM control with query-conditioned algorithmic exploration and reasoning-chain context construction.
Key insight: Stress-test RAG under poisoned or adversarial documents.