Tuesday's cs.AI announcement day (2026-09-29) lists 456 new and 516 cross-lists (replacements skipped; listing total 972). Stack filter for agent systems, memory/context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent, persistence/identity, and local/open serving keeps 43 papers — on-device GUI experience reuse, CUA×SWE, skill and harness evolution, collaborative/agent memory, multi-server MCP orchestration, PhoneCLI mobile commands, and related stack work.


Research Papers

PastForward: Faster On-Device GUI Agents via Computational Experience Reuse

Taehwan Park; Changmin Lee; Hayeon Lee et al. arXiv: 2609.32166

Figure from PastForward: Faster On-Device GUI Agents via Computational Experience Reuse
PastForward: Faster On-Device GUI Agents via Computational Experience Reuse

Running GUI agents on edge devices can keep sensitive screens and interaction histories local, but the computational cost of inference at every action step makes deployment challenging. Existing GUI agent systems either perform full vision-language model (VLM) inference at each action step or reuse coarse-grained knowledge matched to prior tasks.

Key insight: PastForward advances GUI/mobile computer-use agents with practical interaction shortcuts or evals.

CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering

Prince Zizhuang Wang; Chenhao Liang; Zelong Xu et al. arXiv: 2609.32600

Figure from CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering
CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering

Software development requires more than editing code: developers repeatedly run software, interact with its interfaces, visually inspect its behavior, and use these observations to decide what to change next and whether a change works. Existing coding agents and computer-use agents are largely studied in isolation, leaving this integrated development process underexplored.

Key insight: CUA-SWE advances GUI/mobile computer-use agents with practical interaction shortcuts or evals.

Certified Long-Horizon Code Agent Evolution via Validation-Gated Skill Optimization

Yifan Wang; Hao Cheng; Xiaomin Li et al. arXiv: 2609.32990

Long horizon agent self-evolution without model weight updates is essential for enabling deployed agents to accumulate reusable skills and improve over time. Prior self-evolution work has focused primarily on short-horizon tasks, while repository-level software engineering remains unexplored despite being an ideal testbed for long-horizon adaptation.

Key insight: Certified Long-Horizon Code Agent Evolution via Validation-Gated Skill Optimization treats skills or harnesses as the evolvable control surface around the model.

Beyond Skill Evolution: Self-Evolving Context Management Policies for Long-Horizon Agent Harnesses

Weiyuan Li; Jinghan Xu; Aili Chen et al. arXiv: 2609.34649

Figure from Beyond Skill Evolution: Self-Evolving Context Management Policies for Long-Horizon Agent Harnesses
Beyond Skill Evolution: Self-Evolving Context Management Policies for Long-Horizon Agent Harnesses

Harness evolution improves LLM agents by learning from execution trajectories, but existing experience- and skill-based methods are less effective on long-horizon tasks. As interactions grow, useful evidence can be buried by redundant or outdated context, making context management itself a key bottleneck.

Key insight: Beyond Skill Evolution treats skills or harnesses as the evolvable control surface around the model.

Memory as Middleware for Self-Improving AI Agents

K. R. Jayaram; Vatche Isahagian; Vinod Muthusamy et al. arXiv: 2609.32091

AI agents are stateless across sessions by default and therefore operationally amnesic: each session begins with little durable knowledge of prior failures, repairs, preferences, or successful strategies. As a result, agents repeat the same mistakes and discard hard-won experience.

Key insight: Memory as Middleware for Self-Improving AI Agents proposes a concrete memory or context mechanism for long-horizon agents.

CoMemBench: Benchmarking Collaborative Memory Boundaries across Multi-Agent Workflow Topologies

Sen Zhao; Ruiqi Kong; Zuyu Zhang et al. arXiv: 2609.32192

Figure from CoMemBench: Benchmarking Collaborative Memory Boundaries across Multi-Agent Workflow Topologies
CoMemBench: Benchmarking Collaborative Memory Boundaries across Multi-Agent Workflow Topologies

Multi-agent workflows require task-relevant information to be shared across agents, while irrelevant, stale, unverified, or incompatible information must remain isolated. We call this task-conditioned scope of information a collaborative memory boundary.

Key insight: CoMemBench proposes a concrete memory or context mechanism for long-horizon agents.

SkillVine: Agent Skill Evolution via Branching Exploration

Kaiwei Liu; Jiqian Dong; Liran Dong et al. arXiv: 2609.32731

Agent skills encapsulate reusable procedural knowledge that enables LLM agents to perform tasks, and they can be improved automatically using trajectories from interactions with the environment. This is the classic problem of skill evolution.

Key insight: SkillVine treats skills or harnesses as the evolvable control surface around the model.

CUA-Sandbox: Efficient Environments for Computer-Use Agent Reinforcement Learning

Xin Yan; Zhengbo Jiao; Jiaqi Liu et al. arXiv: 2609.32750

Figure from CUA-Sandbox: Efficient Environments for Computer-Use Agent Reinforcement Learning
CUA-Sandbox: Efficient Environments for Computer-Use Agent Reinforcement Learning

Reinforcement learning enables computer-use agents to improve through interaction with real software environments, including websites and desktop applications.

Key insight: CUA-Sandbox advances GUI/mobile computer-use agents with practical interaction shortcuts or evals.

Compositional Safety Failures in Harness Evolution: Identification and Runtime Monitoring

Zhixiang Zhang; Zesen Liu; Wai Ip Lai et al. arXiv: 2609.33123

Self-evolving agent harnesses continually update persistent components such as memory, prompts, skills, and tools. We call this process harness evolution. However, such evolution could introduce unexpected safety risks.

Key insight: Compositional Safety Failures in Harness Evolution treats skills or harnesses as the evolvable control surface around the model.

HTN Planning as a Coordination Layer for Multi-Server MCP Tool Orchestration

Eliott Jacopin; Éric Jacopin; Koichi Takahashi arXiv: 2609.33731

The Model Context Protocol (MCP) isolates servers by design: only the host can orchestrate cross-server workflows. When the host is a large language model, the resulting orchestrations are non-deterministic, non-reproducible, and pay one inference round-trip per tool call.

Key insight: HTN Planning as a Coordination Layer for Multi-Server MCP Tool Orchestration improves how agents orchestrate or recover from MCP/tool interfaces.

R$^2$ Flow: Recursive Self-Improvement via Recursive Skill Evolution

Mingda Zhang; Qiang Huang; Yanjin Li et al. arXiv: 2609.33867

LLM-based agents can improve themselves across tasks by reusing and revising the skills they orchestrate into executable procedures. Flow-based training fits this loop: it samples procedures in proportion to reward, and the flow through each skill credits it for the next library revision.

Key insight: R Flow treats skills or harnesses as the evolvable control surface around the model.

SkillFocus: Evolving Agent Skills via Capability Decomposition

Ning Wang; Zhiren Gong; Bingdong Li et al. arXiv: 2609.34397

Agent skill evolution seeks to improve reusable procedural guidance for large language model (LLM) agents through iterative revision. Existing methods base each revision mainly on execution trajectories or feedback, leaving recurring behavioral requirements across tasks implicit and tying revision to the behavior of the current skill.

Key insight: SkillFocus treats skills or harnesses as the evolvable control surface around the model.

MCP Error Messages Written for Developers Hurt the Most Capable Agents Most

Xiaonan Xu; Wenjing Wu arXiv: 2609.35381

Many Model Context Protocol (MCP) servers wrap web APIs built for human developers, and their error messages tell the reader to run a command, edit a configuration, open a web page or wait. Many agents that read them can only call the server's tools.

Key insight: MCP Error Messages Written for Developers Hurt the Most Capable Agents Most improves how agents orchestrate or recover from MCP/tool interfaces.

When Does a Skill Add Value? Task-Conditional Gain Prediction for Selective Skill Use

Anjie Xu; Zhiyu Zhang; Ruiqing Ding et al. arXiv: 2609.32274

Agent skills are expected to improve task performance. Yet we find that they often provide no benefit, and can even hurt performance while incurring additional token costs. Can we predict whether a skill will help before the agent acts?

Key insight: When Does a Skill Add Value? Task-Conditional Gain Prediction for Selective Skill Use treats skills or harnesses as the evolvable control surface around the model.

Enabling Timely Guidance before Skill Retrieval: Retaining Helpful Warm Tips in Agent Context

Feng Liang; Yupeng Li; Runhao Zeng et al. arXiv: 2609.32339

Reusable skills help LLM-based agents solve complex tasks, but the agent must receive guidance before it commits to an ineffective approach. Existing skill mechanisms often expose only metadata and load full content on demand, leaving useful guidance unavailable until the agent decides to retrieve it.

Key insight: Enabling Timely Guidance before Skill Retrieval treats skills or harnesses as the evolvable control surface around the model.

Trust the Brand, Lose Control: How Identity Hijacks LLM Agent Orchestration

Xutao Mao; Rui Qian; Linghan Chen et al. arXiv: 2609.32635

LLM agents now execute tasks end to end with permission to change real systems and increasingly orchestrate subagents that differ in capability and cost. Prior work treats the choice of subagent as an optimization problem. Yet the orchestrator makes this choice from the identities that subagents display, and an attacker can spoof them.

Key insight: Trust the Brand, Lose Control treats skills or harnesses as the evolvable control surface around the model.

Relic: From Multi-Agent Collaboration to Persistent Organizational Capability

Hongyi Du; Tianyi Zhang; Weijia Zhang et al. arXiv: 2609.32965

Multiple agents may often conflict in an organization: for example, one coding agent changes an interface in a repository, but another continues to develop on the old version where existing tests become stale. A conversation can resolve the episode, but when the participants change, what makes the lesson continue to govern the team?

Key insight: Relic studies coding-agent loops, evals, or memory for software tasks.

What Does a Skill Actually Do? Estimands and Evaluation Validity for Tool and Skill Use in LLM Agents: A Critical Review

Shuyang Zhang arXiv: 2609.33153

Large language model (LLM) agents increasingly draw on external tools and reusable skills selected at run time from libraries that hold thousands of entries.

Key insight: What Does a Skill Actually Do? Estimands and Evaluation Validity for Tool and Skill Use in LLM Agents studies coding-agent loops, evals, or memory for software tasks.

Coding Agent Memory Post-training: Unlocking the Memory Potential of Pre-trained File Operations for Long-Horizon Tasks via Reinforcement Learning

Lirui Luo; Kelong Mao; Heming Xia et al. arXiv: 2609.34422

Language-model agents increasingly tackle long-horizon tasks whose interaction histories exceed the model's active context. Recent work has begun to use reinforcement learning to make memory control part of the policy, often relying on predefined memory tools within domain-specific training environments of relatively short horizons.

Key insight: Coding Agent Memory Post-training proposes a concrete memory or context mechanism for long-horizon agents.

The Compiler May Read It, the Agent May Not: Keeping Part of a Research Code Away from a Coding Agent

Shobhan Roy arXiv: 2609.35557

The compiler must read modules a physics-based solver cannot build without; the coding agent must not read that intellectual property. The harness does not ship that rule. We classified fifteen read routes against a container, permission rules and a sandbox. None of the three can tell which program is reading.

Key insight: The Compiler May Read It, the Agent May Not treats skills or harnesses as the evolvable control surface around the model.

Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs

Vincent-Daniel Yun; Woosang Lim; Haneul Yoo et al. arXiv: 2609.32259

Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender.

Key insight: Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs proposes a concrete memory or context mechanism for long-horizon agents.

PluginRSI: Recursive Improvement of Agent Harnesses with Reusable Plugins

Yaorui Shi; Yuchun Miao; Yuxin Chen et al. arXiv: 2609.32423

Figure from PluginRSI: Recursive Improvement of Agent Harnesses with Reusable Plugins
PluginRSI: Recursive Improvement of Agent Harnesses with Reusable Plugins

The harness surrounding a language model is a central determinant of agent performance. Recent methods optimize harnesses by searching over complete programs, where individual mechanisms are difficult to isolate and reuse. We introduce PluginRSI, which represents a harness as a composition of atomized plugins and organizes harness evolution around these plugins.

Key insight: PluginRSI treats skills or harnesses as the evolvable control surface around the model.

Hearsay: Can an Auditor Trust the Record a Deployed Agent Harness Writes?

Jiahong Dai; Zhuochen Yang; Pengyang Shao et al. arXiv: 2609.32495

An agent harness, the code that turns a model into an agent, writes its own record of each run, and that record is all a later reader gets when a run is disputed, investigated or audited. We call a record evidentiary when a reader who was not there can check it without trusting the writer. Across sixteen deployed frameworks, none writes one in full.

Key insight: Hearsay treats skills or harnesses as the evolvable control surface around the model.

Beyond Dyadic Memory: Interaction-Aware Multimodal Memory with Adaptive Agentic Retrieval for Multi-Party Spoken Conversations

Wenxu Jia; Xize Cheng; Zihan Zhang et al. arXiv: 2609.32522

Long-term memory enables agents to accumulate information and reason across sessions, yet existing research primarily focuses on dyadic text or image-text conversations, leaving long-term memory for multi-party spoken conversations underexplored. This setting requires preserving conversational content, identifying participants across sessions, and retaining who speaks to whom.

Key insight: Beyond Dyadic Memory proposes a concrete memory or context mechanism for long-horizon agents.

CUE-Mem: Benchmarking Long-Term User Memory via Implicit Cues in Multimodal Conversations

Yulin Hu; Yanyan Zhao; Zimo Long et al. arXiv: 2609.32574

Long-term memory is essential for multimodal agents that interact with users across sustained conversations. However, user memories are not always explicitly stated: they may also be implied by recurring background objects in images, ambient sounds in audio, or other peripheral multimodal cues.

Key insight: CUE-Mem proposes a concrete memory or context mechanism for long-horizon agents.

EMIR$^2$: Evolution-Aware Memory with Intent-Guided Multi-Round Retrieval

Jinlan Liu; Hongliang Sun; Yong Wang et al. arXiv: 2609.32584

Long-term memory enables large language model (LLM) agents to leverage historical interactions for future tasks.

Key insight: EMIR proposes a concrete memory or context mechanism for long-horizon agents.

LiteEvo: Automated, Cost-Efficient Harness Evolution for Generalization to Unseen Tasks

Euntae Choi; Sumin Song; Sungjoo Yoo arXiv: 2609.33146

An LLM agent is defined by two things: the weights inside its model and the harness of components assembled around it. Harnesses are still handcrafted, and HarnessX, which evolves them automatically, starts each benchmark from a handcrafted harness, reports gains on the tasks it evolved on, and budgets 100 to 175 million meta-agent tokens per benchmark.

Key insight: LiteEvo treats skills or harnesses as the evolvable control surface around the model.

LSTMem: Hierarchical Long Short-Term Online Memory for Large Language Models

Xianglong Shi; Ruijie Yang; Sirui Zhao et al. arXiv: 2609.33268

Large language models increasingly serve as long-horizon assistants and agents, where they must both accumulate information across interactions and make the relevant parts available when later requests depend on them.

Key insight: LSTMem proposes a concrete memory or context mechanism for long-horizon agents.

Vestrum: Improving Agent Harnesses by Adapting Their Verification, Structure and Memory

Jayant Parashar; Eugene F. Douglass; William C. Bastian et al. arXiv: 2609.33822

An agent harness controls how a language model accesses information, uses tools, preserves memory, and checks its work. Improving this software is costly when each evaluation requires a long interaction with an environment. We introduce Vestrum, a framework that turns failures in execution traces into scoped harness changes without training the task model.

Key insight: Vestrum treats skills or harnesses as the evolvable control surface around the model.

From Attack Success to Attack Severity: Counterfactual Memory Attacks on LLM Agents

Mingxi Zou; Langzhang Liang; Zhuo Wang et al. arXiv: 2609.34132

As LLM agents increasingly rely on persistent memory for long-horizon and personalized behavior, they can retain and reuse information across interactions, but this also creates a lasting channel through which malicious memory writes can influence future behavior.

Key insight: From Attack Success to Attack Severity proposes a concrete memory or context mechanism for long-horizon agents.

Stashbird: Efficient Speaker-Indexed Memory for Conversational Agents

Chidera Biringa; Lucas Yannul; Xiaowen Wang et al. arXiv: 2609.34242

AI agents require memory that preserves information across user-agent exchanges, user-to-user conversations, and group conversations with or without agent participation, while supporting updates as evidence changes or is removed. We present Stashbird, an agent memory system that links source episodes to derived memory state through explicit provenance.

Key insight: Stashbird proposes a concrete memory or context mechanism for long-horizon agents.

Remember by Asking: Retrieval-Induced Memory Evolution for LLM Agents

Wanqi Zhou; Jiawei Lu; Yang Wang et al. arXiv: 2609.34438

Long-term memory is essential for language agents to maintain coherent and effective behavior over extended, multi-session interactions. Existing memory systems mainly use retrieval at read time, while write-time memory formation still relies on direct extraction or compression.

Key insight: Remember by Asking proposes a concrete memory or context mechanism for long-horizon agents.

Codoku: Renewable Program-Reasoning Challenges for Frontier Coding Agents

Cong Li; Hao Sun; Zenan Li et al. arXiv: 2609.34661

Existing program-reasoning benchmarks ask large language models to predict a program's behavior on a given input.

Key insight: Codoku studies coding-agent loops, evals, or memory for software tasks.

Planarian: Managing Agent State with Statepoints

Jinnan Guo; Hao Mark Chen; Kapil Vaswani et al. arXiv: 2609.35366

LLM agents solve complex tasks by iteratively changing files, invoking local tools, and interacting with remote services, which modifies state across their local environment and remote services. Today, agents and users must manage these changes explicitly, whether reverting exploratory actions or recovering from erroneous ones.

Key insight: Planarian targets on-device or local agent serving constraints.

When Consent Outlives Context: Residual Authority Replay in Long-Lived Agents

Zhihao Zhang; Chao Wang; Rujia Li et al. arXiv: 2609.33910

LLM agents increasingly rely on user approval to authorize security-sensitive actions at runtime. Such approvals are granted within a specific task and execution context. In long-lived agents, authorization decisions may need to persist across tasks or sessions.

Key insight: When Consent Outlives Context studies coding-agent loops, evals, or memory for software tasks.

AuthorityLens: Rethinking LLM-Based Agent Systems Through the Lens of Authority

Shaojin Chen; Huihao Jing; Wun Yu Chan et al. arXiv: 2609.32378

LLM-based agents are increasingly deployed with authority over consequential resources and decisions in real systems. These agents often operate alongside human and LLM-based participants who hold different forms of authority. Yet workflow roles, permission settings, and review mechanisms do not necessarily reflect the authority realized in practice.

Key insight: AuthorityLens contributes a system-level idea for agent stacks.

Memory as a cache: Exact context reuse and deletion by construction

Shengyao Wang; Jiang Liu arXiv: 2609.32395

The KV cache of a transformer entangles every token's representation with its entire prefix: a passage encoded once cannot be reused under a different prefix or removed without recomputing everything after it, so exact cache reuse is limited to shared prefixes. We present SMem, an architecture whose context representation is a cache by construction.

Key insight: Memory as a cache proposes a concrete memory or context mechanism for long-horizon agents.

MemAgent: Learning to Manage Heterogeneous Memory Providers for LLM Agents

Yongxian Wei; Yilin Zhao; Runxi Cheng et al. arXiv: 2609.32521

Figure from MemAgent: Learning to Manage Heterogeneous Memory Providers for LLM Agents
MemAgent: Learning to Manage Heterogeneous Memory Providers for LLM Agents

Current agents remain largely stateless across tasks, limiting their ability to continually improve from prior interactions and making memory essential for long-horizon agentic behavior.

Key insight: MemAgent proposes a concrete memory or context mechanism for long-horizon agents.

ActiveMem: Dynamic Latent Memory Trees for Long-Horizon Agents

Song-Li Wu; Jingyi Wang; Zhaocheng Du et al. arXiv: 2609.33244

Large Language Model (LLM) agents increasingly rely on external memory to support long-horizon reasoning and decision making. Existing memory systems typically retrieve historical trajectories or summaries as independent context fragments, overlooking the procedural dependencies underlying multi-step execution.

Key insight: ActiveMem proposes a concrete memory or context mechanism for long-horizon agents.

FlowState: Execution State as Memory for Long-Horizon LLM Agents

Minghao Li; Bangyan Li; Zifan Wang et al. arXiv: 2609.34565

Long-horizon tasks require LLM agents to continually draw on information from earlier interactions. However, retaining the full history increases context costs, while compressing it risks losing details needed later, and the relevance of historical information often becomes apparent as the task progresses.

Key insight: FlowState proposes a concrete memory or context mechanism for long-horizon agents.

Report: Progressive Disclosure of Agent Skills

Guilin Zhang; Kai Zhao; Priyanka Mudgal et al. arXiv: 2609.35692

Users of Workday's deployed LLM-based agents often request features which can be addressed by defining named procedures, also known as skills, in the LLM context, effectively augmenting agents' capabilities. However, as an agent's skills library grows in size, so does the agent's operational cost.

Key insight: Report treats skills or harnesses as the evolvable control surface around the model.

Probe to Act: Elevating Browser-Use Agent via Active Visual Probing

Keliang Li; Heng Wang; Chen Hu et al. arXiv: 2609.33646

Browser-use agents require seamless alignment between structured web metadata and visual information, while preserving relevant context across long interactions. Existing interfaces often rely on either screenshot-level action prediction or static Set-of-Marks overlays, leaving the model to resolve dense DOM-pixel alignment before every operation.

Key insight: Probe to Act advances GUI/mobile computer-use agents with practical interaction shortcuts or evals.

PhoneCLI: From App Interfaces to Callable Commands for Mobile Agents

Yangqin Jiang; Lingrui Xu; Chao Huang arXiv: 2609.35671

Figure from PhoneCLI: From App Interfaces to Callable Commands for Mobile Agents
PhoneCLI: From App Interfaces to Callable Commands for Mobile Agents

Mobile GUI agents operate through a perception--action loop: at each step they screenshot the device, invoke a vision--language model (VLM), and emit an action. It is slow, costly, and brittle, yet most of what it does is navigation---and everyday navigation is static, ordered, and endlessly repeated.

Key insight: PhoneCLI advances GUI/mobile computer-use agents with practical interaction shortcuts or evals.