Tuesday's cs.AI announcement day (2026-09-29) lists 456 new and 516 cross-lists (replacements skipped; listing total 972). Stack filter for agent systems, memory/context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent, persistence/identity, and local/open serving keeps 43 papers — on-device GUI experience reuse, CUA×SWE, skill and harness evolution, collaborative/agent memory, multi-server MCP orchestration, PhoneCLI mobile commands, and related stack work.
Taehwan Park; Changmin Lee; Hayeon Lee et al. arXiv: 2609.32166
Running GUI agents on edge devices can keep sensitive screens and interaction histories local, but the computational cost of inference at every action step makes deployment challenging. Existing GUI agent systems either perform full vision-language model (VLM) inference at each action step or reuse coarse-grained knowledge matched to prior tasks.
Key insight: PastForward advances GUI/mobile computer-use agents with practical interaction shortcuts or evals.
Prince Zizhuang Wang; Chenhao Liang; Zelong Xu et al. arXiv: 2609.32600
Software development requires more than editing code: developers repeatedly run software, interact with its interfaces, visually inspect its behavior, and use these observations to decide what to change next and whether a change works. Existing coding agents and computer-use agents are largely studied in isolation, leaving this integrated development process underexplored.
Key insight: CUA-SWE advances GUI/mobile computer-use agents with practical interaction shortcuts or evals.
Yifan Wang; Hao Cheng; Xiaomin Li et al. arXiv: 2609.32990
Long horizon agent self-evolution without model weight updates is essential for enabling deployed agents to accumulate reusable skills and improve over time. Prior self-evolution work has focused primarily on short-horizon tasks, while repository-level software engineering remains unexplored despite being an ideal testbed for long-horizon adaptation.
Key insight: Certified Long-Horizon Code Agent Evolution via Validation-Gated Skill Optimization treats skills or harnesses as the evolvable control surface around the model.
Weiyuan Li; Jinghan Xu; Aili Chen et al. arXiv: 2609.34649
Harness evolution improves LLM agents by learning from execution trajectories, but existing experience- and skill-based methods are less effective on long-horizon tasks. As interactions grow, useful evidence can be buried by redundant or outdated context, making context management itself a key bottleneck.
Key insight: Beyond Skill Evolution treats skills or harnesses as the evolvable control surface around the model.
K. R. Jayaram; Vatche Isahagian; Vinod Muthusamy et al. arXiv: 2609.32091
AI agents are stateless across sessions by default and therefore operationally amnesic: each session begins with little durable knowledge of prior failures, repairs, preferences, or successful strategies. As a result, agents repeat the same mistakes and discard hard-won experience.
Key insight: Memory as Middleware for Self-Improving AI Agents proposes a concrete memory or context mechanism for long-horizon agents.
Sen Zhao; Ruiqi Kong; Zuyu Zhang et al. arXiv: 2609.32192
Multi-agent workflows require task-relevant information to be shared across agents, while irrelevant, stale, unverified, or incompatible information must remain isolated. We call this task-conditioned scope of information a collaborative memory boundary.
Key insight: CoMemBench proposes a concrete memory or context mechanism for long-horizon agents.
Kaiwei Liu; Jiqian Dong; Liran Dong et al. arXiv: 2609.32731
Agent skills encapsulate reusable procedural knowledge that enables LLM agents to perform tasks, and they can be improved automatically using trajectories from interactions with the environment. This is the classic problem of skill evolution.
Key insight: SkillVine treats skills or harnesses as the evolvable control surface around the model.
Xin Yan; Zhengbo Jiao; Jiaqi Liu et al. arXiv: 2609.32750
Reinforcement learning enables computer-use agents to improve through interaction with real software environments, including websites and desktop applications.
Key insight: CUA-Sandbox advances GUI/mobile computer-use agents with practical interaction shortcuts or evals.
Zhixiang Zhang; Zesen Liu; Wai Ip Lai et al. arXiv: 2609.33123
Self-evolving agent harnesses continually update persistent components such as memory, prompts, skills, and tools. We call this process harness evolution. However, such evolution could introduce unexpected safety risks.
Key insight: Compositional Safety Failures in Harness Evolution treats skills or harnesses as the evolvable control surface around the model.
Eliott Jacopin; Éric Jacopin; Koichi Takahashi arXiv: 2609.33731
The Model Context Protocol (MCP) isolates servers by design: only the host can orchestrate cross-server workflows. When the host is a large language model, the resulting orchestrations are non-deterministic, non-reproducible, and pay one inference round-trip per tool call.
Key insight: HTN Planning as a Coordination Layer for Multi-Server MCP Tool Orchestration improves how agents orchestrate or recover from MCP/tool interfaces.
Mingda Zhang; Qiang Huang; Yanjin Li et al. arXiv: 2609.33867
LLM-based agents can improve themselves across tasks by reusing and revising the skills they orchestrate into executable procedures. Flow-based training fits this loop: it samples procedures in proportion to reward, and the flow through each skill credits it for the next library revision.
Key insight: R Flow treats skills or harnesses as the evolvable control surface around the model.
Ning Wang; Zhiren Gong; Bingdong Li et al. arXiv: 2609.34397
Agent skill evolution seeks to improve reusable procedural guidance for large language model (LLM) agents through iterative revision. Existing methods base each revision mainly on execution trajectories or feedback, leaving recurring behavioral requirements across tasks implicit and tying revision to the behavior of the current skill.
Key insight: SkillFocus treats skills or harnesses as the evolvable control surface around the model.
Xiaonan Xu; Wenjing Wu arXiv: 2609.35381
Many Model Context Protocol (MCP) servers wrap web APIs built for human developers, and their error messages tell the reader to run a command, edit a configuration, open a web page or wait. Many agents that read them can only call the server's tools.
Key insight: MCP Error Messages Written for Developers Hurt the Most Capable Agents Most improves how agents orchestrate or recover from MCP/tool interfaces.
Anjie Xu; Zhiyu Zhang; Ruiqing Ding et al. arXiv: 2609.32274
Agent skills are expected to improve task performance. Yet we find that they often provide no benefit, and can even hurt performance while incurring additional token costs. Can we predict whether a skill will help before the agent acts?
Key insight: When Does a Skill Add Value? Task-Conditional Gain Prediction for Selective Skill Use treats skills or harnesses as the evolvable control surface around the model.
Feng Liang; Yupeng Li; Runhao Zeng et al. arXiv: 2609.32339
Reusable skills help LLM-based agents solve complex tasks, but the agent must receive guidance before it commits to an ineffective approach. Existing skill mechanisms often expose only metadata and load full content on demand, leaving useful guidance unavailable until the agent decides to retrieve it.
Key insight: Enabling Timely Guidance before Skill Retrieval treats skills or harnesses as the evolvable control surface around the model.
Xutao Mao; Rui Qian; Linghan Chen et al. arXiv: 2609.32635
LLM agents now execute tasks end to end with permission to change real systems and increasingly orchestrate subagents that differ in capability and cost. Prior work treats the choice of subagent as an optimization problem. Yet the orchestrator makes this choice from the identities that subagents display, and an attacker can spoof them.
Key insight: Trust the Brand, Lose Control treats skills or harnesses as the evolvable control surface around the model.
Hongyi Du; Tianyi Zhang; Weijia Zhang et al. arXiv: 2609.32965
Multiple agents may often conflict in an organization: for example, one coding agent changes an interface in a repository, but another continues to develop on the old version where existing tests become stale. A conversation can resolve the episode, but when the participants change, what makes the lesson continue to govern the team?
Key insight: Relic studies coding-agent loops, evals, or memory for software tasks.
Shuyang Zhang arXiv: 2609.33153
Large language model (LLM) agents increasingly draw on external tools and reusable skills selected at run time from libraries that hold thousands of entries.
Key insight: What Does a Skill Actually Do? Estimands and Evaluation Validity for Tool and Skill Use in LLM Agents studies coding-agent loops, evals, or memory for software tasks.
Lirui Luo; Kelong Mao; Heming Xia et al. arXiv: 2609.34422
Language-model agents increasingly tackle long-horizon tasks whose interaction histories exceed the model's active context. Recent work has begun to use reinforcement learning to make memory control part of the policy, often relying on predefined memory tools within domain-specific training environments of relatively short horizons.
Key insight: Coding Agent Memory Post-training proposes a concrete memory or context mechanism for long-horizon agents.
Shobhan Roy arXiv: 2609.35557
The compiler must read modules a physics-based solver cannot build without; the coding agent must not read that intellectual property. The harness does not ship that rule. We classified fifteen read routes against a container, permission rules and a sandbox. None of the three can tell which program is reading.
Key insight: The Compiler May Read It, the Agent May Not treats skills or harnesses as the evolvable control surface around the model.
Vincent-Daniel Yun; Woosang Lim; Haneul Yoo et al. arXiv: 2609.32259
Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender.
Key insight: Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs proposes a concrete memory or context mechanism for long-horizon agents.
Yaorui Shi; Yuchun Miao; Yuxin Chen et al. arXiv: 2609.32423
The harness surrounding a language model is a central determinant of agent performance. Recent methods optimize harnesses by searching over complete programs, where individual mechanisms are difficult to isolate and reuse. We introduce PluginRSI, which represents a harness as a composition of atomized plugins and organizes harness evolution around these plugins.
Key insight: PluginRSI treats skills or harnesses as the evolvable control surface around the model.
Jiahong Dai; Zhuochen Yang; Pengyang Shao et al. arXiv: 2609.32495
An agent harness, the code that turns a model into an agent, writes its own record of each run, and that record is all a later reader gets when a run is disputed, investigated or audited. We call a record evidentiary when a reader who was not there can check it without trusting the writer. Across sixteen deployed frameworks, none writes one in full.
Key insight: Hearsay treats skills or harnesses as the evolvable control surface around the model.
Wenxu Jia; Xize Cheng; Zihan Zhang et al. arXiv: 2609.32522
Long-term memory enables agents to accumulate information and reason across sessions, yet existing research primarily focuses on dyadic text or image-text conversations, leaving long-term memory for multi-party spoken conversations underexplored. This setting requires preserving conversational content, identifying participants across sessions, and retaining who speaks to whom.
Key insight: Beyond Dyadic Memory proposes a concrete memory or context mechanism for long-horizon agents.
Yulin Hu; Yanyan Zhao; Zimo Long et al. arXiv: 2609.32574
Long-term memory is essential for multimodal agents that interact with users across sustained conversations. However, user memories are not always explicitly stated: they may also be implied by recurring background objects in images, ambient sounds in audio, or other peripheral multimodal cues.
Key insight: CUE-Mem proposes a concrete memory or context mechanism for long-horizon agents.
Jinlan Liu; Hongliang Sun; Yong Wang et al. arXiv: 2609.32584
Long-term memory enables large language model (LLM) agents to leverage historical interactions for future tasks.
Key insight: EMIR proposes a concrete memory or context mechanism for long-horizon agents.
Euntae Choi; Sumin Song; Sungjoo Yoo arXiv: 2609.33146
An LLM agent is defined by two things: the weights inside its model and the harness of components assembled around it. Harnesses are still handcrafted, and HarnessX, which evolves them automatically, starts each benchmark from a handcrafted harness, reports gains on the tasks it evolved on, and budgets 100 to 175 million meta-agent tokens per benchmark.
Key insight: LiteEvo treats skills or harnesses as the evolvable control surface around the model.
Xianglong Shi; Ruijie Yang; Sirui Zhao et al. arXiv: 2609.33268
Large language models increasingly serve as long-horizon assistants and agents, where they must both accumulate information across interactions and make the relevant parts available when later requests depend on them.
Key insight: LSTMem proposes a concrete memory or context mechanism for long-horizon agents.
Jayant Parashar; Eugene F. Douglass; William C. Bastian et al. arXiv: 2609.33822
An agent harness controls how a language model accesses information, uses tools, preserves memory, and checks its work. Improving this software is costly when each evaluation requires a long interaction with an environment. We introduce Vestrum, a framework that turns failures in execution traces into scoped harness changes without training the task model.
Key insight: Vestrum treats skills or harnesses as the evolvable control surface around the model.
Mingxi Zou; Langzhang Liang; Zhuo Wang et al. arXiv: 2609.34132
As LLM agents increasingly rely on persistent memory for long-horizon and personalized behavior, they can retain and reuse information across interactions, but this also creates a lasting channel through which malicious memory writes can influence future behavior.
Key insight: From Attack Success to Attack Severity proposes a concrete memory or context mechanism for long-horizon agents.
Chidera Biringa; Lucas Yannul; Xiaowen Wang et al. arXiv: 2609.34242
AI agents require memory that preserves information across user-agent exchanges, user-to-user conversations, and group conversations with or without agent participation, while supporting updates as evidence changes or is removed. We present Stashbird, an agent memory system that links source episodes to derived memory state through explicit provenance.
Key insight: Stashbird proposes a concrete memory or context mechanism for long-horizon agents.
Wanqi Zhou; Jiawei Lu; Yang Wang et al. arXiv: 2609.34438
Long-term memory is essential for language agents to maintain coherent and effective behavior over extended, multi-session interactions. Existing memory systems mainly use retrieval at read time, while write-time memory formation still relies on direct extraction or compression.
Key insight: Remember by Asking proposes a concrete memory or context mechanism for long-horizon agents.
Cong Li; Hao Sun; Zenan Li et al. arXiv: 2609.34661
Existing program-reasoning benchmarks ask large language models to predict a program's behavior on a given input.
Key insight: Codoku studies coding-agent loops, evals, or memory for software tasks.
Jinnan Guo; Hao Mark Chen; Kapil Vaswani et al. arXiv: 2609.35366
LLM agents solve complex tasks by iteratively changing files, invoking local tools, and interacting with remote services, which modifies state across their local environment and remote services. Today, agents and users must manage these changes explicitly, whether reverting exploratory actions or recovering from erroneous ones.
Key insight: Planarian targets on-device or local agent serving constraints.
Zhihao Zhang; Chao Wang; Rujia Li et al. arXiv: 2609.33910
LLM agents increasingly rely on user approval to authorize security-sensitive actions at runtime. Such approvals are granted within a specific task and execution context. In long-lived agents, authorization decisions may need to persist across tasks or sessions.
Key insight: When Consent Outlives Context studies coding-agent loops, evals, or memory for software tasks.
Shaojin Chen; Huihao Jing; Wun Yu Chan et al. arXiv: 2609.32378
LLM-based agents are increasingly deployed with authority over consequential resources and decisions in real systems. These agents often operate alongside human and LLM-based participants who hold different forms of authority. Yet workflow roles, permission settings, and review mechanisms do not necessarily reflect the authority realized in practice.
Key insight: AuthorityLens contributes a system-level idea for agent stacks.
Shengyao Wang; Jiang Liu arXiv: 2609.32395
The KV cache of a transformer entangles every token's representation with its entire prefix: a passage encoded once cannot be reused under a different prefix or removed without recomputing everything after it, so exact cache reuse is limited to shared prefixes. We present SMem, an architecture whose context representation is a cache by construction.
Key insight: Memory as a cache proposes a concrete memory or context mechanism for long-horizon agents.
Yongxian Wei; Yilin Zhao; Runxi Cheng et al. arXiv: 2609.32521
Current agents remain largely stateless across tasks, limiting their ability to continually improve from prior interactions and making memory essential for long-horizon agentic behavior.
Key insight: MemAgent proposes a concrete memory or context mechanism for long-horizon agents.
Song-Li Wu; Jingyi Wang; Zhaocheng Du et al. arXiv: 2609.33244
Large Language Model (LLM) agents increasingly rely on external memory to support long-horizon reasoning and decision making. Existing memory systems typically retrieve historical trajectories or summaries as independent context fragments, overlooking the procedural dependencies underlying multi-step execution.
Key insight: ActiveMem proposes a concrete memory or context mechanism for long-horizon agents.
Minghao Li; Bangyan Li; Zifan Wang et al. arXiv: 2609.34565
Long-horizon tasks require LLM agents to continually draw on information from earlier interactions. However, retaining the full history increases context costs, while compressing it risks losing details needed later, and the relevance of historical information often becomes apparent as the task progresses.
Key insight: FlowState proposes a concrete memory or context mechanism for long-horizon agents.
Guilin Zhang; Kai Zhao; Priyanka Mudgal et al. arXiv: 2609.35692
Users of Workday's deployed LLM-based agents often request features which can be addressed by defining named procedures, also known as skills, in the LLM context, effectively augmenting agents' capabilities. However, as an agent's skills library grows in size, so does the agent's operational cost.
Key insight: Report treats skills or harnesses as the evolvable control surface around the model.
Keliang Li; Heng Wang; Chen Hu et al. arXiv: 2609.33646
Browser-use agents require seamless alignment between structured web metadata and visual information, while preserving relevant context across long interactions. Existing interfaces often rely on either screenshot-level action prediction or static Set-of-Marks overlays, leaving the model to resolve dense DOM-pixel alignment before every operation.
Key insight: Probe to Act advances GUI/mobile computer-use agents with practical interaction shortcuts or evals.
Yangqin Jiang; Lingrui Xu; Chao Huang arXiv: 2609.35671
Mobile GUI agents operate through a perception--action loop: at each step they screenshot the device, invoke a vision--language model (VLM), and emit an action. It is slow, costly, and brittle, yet most of what it does is navigation---and everyday navigation is static, ordered, and endlessly repeated.
Key insight: PhoneCLI advances GUI/mobile computer-use agents with practical interaction shortcuts or evals.