Thursday's cs.AI announcement day (2026-10-01) lists 133 new and 261 cross-lists (replacements skipped; listing total 394). Stack filter for agent systems, memory/context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent, persistence/identity, and personal/mobile agents keeps 42 papers — self-evolving harnesses (AREX-2, Self-Evolving Harness, Turbo Harness, MILO), personal/voice agents (AgBench, Talk2Agent, VAmoS, Richard), GUI/computer-use (GUI Agent Memory, OSWorld-Science, ComputerSD, cua-speedrun), and memory/context (RefCon, BELIEFRAG, TAGGRAPH, Galahad).


Research Papers

AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks

Hongjin Qian; Chaofan Li; Kun Luo et al. arXiv: 2609.38288

Figure from AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks
AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks

We present AREX-2, an effort to advance the self-improving capability of LLM agents, which we define as the ability to iteratively refine a solution at test time. This ability rests on two complementary capabilities: reflection, which produces a solution better than the current one, and long-horizon execution, which keeps the iteration effective over many rounds.

Key insight: AREX-2 studies self-improvement / RSI loops agents can run on their own harness.

MoFlow: Multi-Objective Agentic Workflow Generation

Yining Lu; Aurelie Lozano; Xi Yang et al. arXiv: 2609.38294

We study the generation of agentic workflows that jointly optimize multiple objectives, such as accuracy, cost, latency, robustness, and consistency. Existing methods for workflow generation typically optimize accuracy alone or a weighted sum of objectives, so each trained generator commits to one fixed trade-off and must be retrained from scratch when preferences change.

Key insight: MoFlow studies agent architecture or runtime patterns under realistic workloads.

Self-Evolving Harness on Multiple Tasks with the Agent as Its Own Optimizer

Qiankai Xu arXiv: 2609.38372

Figure from Self-Evolving Harness on Multiple Tasks with the Agent as Its Own Optimizer
Self-Evolving Harness on Multiple Tasks with the Agent as Its Own Optimizer

A harness is the code around a language-model agent that organizes prompts, calls tools, manages context, and controls execution. As models grow stronger, recent work has begun to let agents improve their own harnesses, a line of work known as self-evolving harnesses.

Key insight: Self-Evolving Harness on Multiple Tasks with the Agent as Its Own Optimizer treats skills or harnesses as the evolvable control surface around the model.

VAmoS Part Deux: Harder, More Realistic Voice-Agent Simulation

Joshua Meyer; Sahar Shayegan; Ritiz Tambi et al. arXiv: 2609.38512

Voice agents in production must handle several requests, background speech, and customers who lose patience. We introduce VAmoS Energy, a benchmark that combines these challenges in 100 calls about utility billing and payment assistance.

Key insight: VAmoS Part Deux targets personal, mobile, or voice agents under real device constraints.

AgBench: Agentic AI Benchmarks for Personal AI Devices

Yizhou Han; Di Wu; Dhananjay Saikumar et al. arXiv: 2609.38652

Figure from AgBench: Agentic AI Benchmarks for Personal AI Devices
AgBench: Agentic AI Benchmarks for Personal AI Devices

Agentic AI systems increasingly rely on cloud-hosted large language models for planning, tool use, and iterative execution, raising concerns about API cost and data exposure. Advances in personal AI devices enable agents to execute locally, but limited resources on device may affect task success and performance.

Key insight: AgBench targets personal, mobile, or voice agents under real device constraints.

Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability

Jeffrey Willette; Krishna C. Puvvada; Boris Ginsburg arXiv: 2609.38712

Long-horizon agentic workflows require models to sustain repeated state-dependent actions all while the context grows, sub-task complexity changes, and new data arrives. Each situation represents an independent axis along which an agent may fail.

Key insight: Staying on Task advances long-horizon planning, RL, or execution fidelity for agents.

Action Conditioned Bisimulation For GUI Agent Memory

Hongbo Zhang; Liuyang Song; Quanquan Li et al. arXiv: 2609.38778

Figure from Action Conditioned Bisimulation For GUI Agent Memory
Action Conditioned Bisimulation For GUI Agent Memory

An agent that remembers what it did on a web page must decide when two pages count as the same. Memories built on observation similarity merge pages that look alike but behave differently, and GUIs are full of such pages: two tabs of one widget or two rows of one menu answer the same click differently.

Key insight: Action Conditioned Bisimulation For GUI Agent Memory advances computer-use / GUI agent competence or benchmarking.

When Context Changes: Understanding Update Failures in LLMs

Junyu Guo; Yuchen Fang; Shangding Gu et al. arXiv: 2609.38866

As preferences, goals, and facts change, LLM agents must use the current state while earlier versions remain in context. Yet they can answer with an old value of the same variable, a failure that we call stale binding.

Key insight: When Context Changes proposes a concrete memory or context mechanism for long-horizon agents.

Talk2Agent: Benchmarking Voice Interfaces for Text Agents

Terumi Chiba; Guangzhi Sun; Zheqi Yuan et al. arXiv: 2609.38867

Large language model (LLM) computer-use agents are typically evaluated with clean written instructions, despite speech being an increasingly popular interface for interacting with such systems. Speech input introduces an additional failure point: transcription errors can alter task-critical entities, constraints, or targets before the agent begins reasoning, while conventional ASR metrics do not directly measure whether the information require…

Key insight: Talk2Agent targets personal, mobile, or voice agents under real device constraints.

Consistent Plan-Act for Long-Horizon Agentic Tasks

Heng-Zhuang Li; Yi-Kai Zhang; Yu Wang et al. arXiv: 2609.38891

Long-horizon agentic tasks demand strong reasoning and efficient execution across successive interactions with dynamic environments. A common approach decouples high-level planning from low-level execution through separate planner and actor roles.

Key insight: Consistent Plan-Act for Long-Horizon Agentic Tasks advances long-horizon planning, RL, or execution fidelity for agents.

Composing Task-specific Agent Harnesses at Test Time with Reusable Primitives

Peng Kuang; Haibo Jin; Dehao Wu et al. arXiv: 2609.38912

Figure from Composing Task-specific Agent Harnesses at Test Time with Reusable Primitives
Composing Task-specific Agent Harnesses at Test Time with Reusable Primitives

Agent harnesses govern how large language models (LLMs) gather context, invoke tools, verify results, preserve state, and terminate, largely affecting agent performance. However, the value of each harness mechanism can differ across heterogeneous tasks: a mechanism that improves one task may impose overhead or context distraction on another, leading to the suboptimality of a global harness.

Key insight: Composing Task-specific Agent Harnesses at Test Time with Reusable Primitives treats skills or harnesses as the evolvable control surface around the model.

BELIEFRAG: Making Adaptive RAG State-Aware under Evolving Evidence

Hongji Pu arXiv: 2609.39139

Adaptive RAG uses signals such as confidence, relevance, support, and retrieval quality to decide when to search or correct evidence. In multi-step retrieval, however, these local signals must be combined into a persistent view of what the current evidence supports, what remains missing, and which action should follow.

Key insight: BELIEFRAG proposes a concrete memory or context mechanism for long-horizon agents.

RefCon: Iterative Refinement and Contrastive Memory Extraction for Context-Evolving Agent

Ubaidillah Ariq Prathama; Bo Liu; Yeo Boon Hong et al. arXiv: 2609.39143

Long-horizon agent interactions generate useful but noisy experience, and retraining models to absorb it is expensive. Context-evolving agents therefore need memory extraction methods that improve with more test-time compute without relying on gold labels.

Key insight: RefCon proposes a concrete memory or context mechanism for long-horizon agents.

Do Self-Evolving Skills Generalize to Held-Out Tasks?

Xihao Piao; Zifeng Wang; Zhen Chen arXiv: 2609.39148

AI agents can externalize what they learn from past tasks into reusable \emph{skills}, such as procedures, checklists, code, or other executable artifacts, that can be retrieved and reused when solving new tasks. Self-evolving skill methods keep rewriting these skills after each round of practice on training tasks, and the skill is then used on new tasks of the same kind.

Key insight: Do Self-Evolving Skills Generalize to Held-Out Tasks? treats skills or harnesses as the evolvable control surface around the model.

Rep2Skill: Representation-Guided Skill Self-Evolution for LLM Agents

Kaixing Zhang; Changming Li; Yingdong Shi et al. arXiv: 2609.39149

Textual skills enable large language model (LLM) based agents to accumulate reusable procedural knowledge without updating model parameters. Yet existing skill evolution remains largely confined to the text space: an optimizer must diagnose success and failure patterns, and revise skills solely from long execution trajectories and sparse task outcomes.

Key insight: Rep2Skill treats skills or harnesses as the evolvable control surface around the model.

WorkGenesis: Building the Worlds That Teach Agents to Work

Xinyu Zhu; Fenyi Liu; Yuzhu Cai et al. arXiv: 2609.39325

The ability of Large Language Model (LLM) agents to complete daily and professional work is receiving increasing attention. Training such agents requires realistic work scenarios.

Key insight: WorkGenesis advances long-horizon planning, RL, or execution fidelity for agents.

Beyond the Shadows of Plato's Cave: Evaluating False Memory in Autonomous Agents via Counterfactual Reasoning

Quan M. Tran; Zhuo Huang; Zhen Fang et al. arXiv: 2609.39473

Autonomous agents increasingly rely on memory to generalize beyond their training environments. However, agents are bounded by what they have seen and believed, and leveraging such memories in unseen environments can introduce biases into their internal beliefs.

Key insight: Beyond the Shadows of Plato's Cave proposes a concrete memory or context mechanism for long-horizon agents.

OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software

Dingyuan Dai; Heli Qi; Lei Liu et al. arXiv: 2609.39903

Scientific software presents a demanding test for computer-using agents based on visual language models (VLMs): completing a research workflow requires interpreting specialized interfaces, manipulating scientific objects, and producing verifiable results. We thus introduce OSWorld-Science, a benchmark and evaluation environment that combines scientifically meaningful tasks, artifact-based evaluation, and an efficient agent harness for studying…

Key insight: OSWorld-Science advances computer-use / GUI agent competence or benchmarking.

Learning from Research: Toward Lifelong Agent Harness Evolution

Jingbo Yang; Kwei-Herng Lai; Xiaowen Wang et al. arXiv: 2609.40169

Language agents are expected to solve increasingly complex tasks, creating a growing need for continual improvement. One promising approach is to evolve the agent harness, the software that governs tool use, memory management, and task execution, while keeping the underlying language model fixed.

Key insight: Learning from Research treats skills or harnesses as the evolvable control surface around the model.

How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?

Kirill Brilliantov; Alejandro Hernández-Cano; Emmanuel Abbé arXiv: 2609.40303

Recent autonomous machine learning engineering (MLE) agents have made significant progress on public leaderboards. Often motivated by progress stagnation over long-horizon cycles and limited Large Language Model (LLM) primitives, modern MLE agents are deployed on top of increasingly elaborate machinery: multi-agent orchestrators, dedicated retrieval subagents, and more.

Key insight: How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering? treats skills or harnesses as the evolvable control surface around the model.

Turbo Harness: Instance-Adaptive Harness Optimization

Tunyu Zhang; Hao Wang; Kai Xu et al. arXiv: 2609.40330

Automating the search for effective harnesses is an important step toward enabling agents to recursively self-improve. Existing harness optimizations typically produce a single global harness that is applied uniformly across task instances.

Key insight: Turbo Harness treats skills or harnesses as the evolvable control surface around the model.

AIM: Agentic Idea Management for Automated Research

Hyeong Kyu Choi; Bhavana Dalvi Mishra; Jiefeng Chen et al. arXiv: 2609.38445

Frontier LLMs are increasingly used to automate scientific research through iterative search. We distinguish idea-driven search from solution-driven search and identify three core challenges: organizing evolving research ideas, selecting promising directions, and maintaining alignment between ideas and their implementations.

Key insight: AIM audits or improves deep-research / search agents.

MADBench: Benchmarking the Security of Multi-Agent Debate

Yuwan Liu; Jiaming Zhang; Yue Huang et al. arXiv: 2609.39146

Multi-agent debate (MAD) can improve large language model (LLM) reasoning by allowing multiple agents to exchange and critique their answers to the same task. However, the interactions that enable agents to correct mistakes can also spread adversarial errors and steer the agents toward an incorrect answer.

Key insight: MADBench studies multi-agent coordination, debate, or identity under realistic topologies.

PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents

Yinghui He; Yapei Chang; Khushi Bhardwaj et al. arXiv: 2609.40285

On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound across turns.

Key insight: PivotOPD advances long-horizon planning, RL, or execution fidelity for agents.

When Correct Memory Goes Wrong: Fuzzing Persistent Memory Use in LLM Agents

Yuqiao Meng; Luoxi Tang; Yingxue Zhang et al. arXiv: 2609.38275

Persistent memory helps LLM agents carry information across long interactions, but correct memory can still be used incorrectly when queries change or memory states evolve. Existing work mainly studies memory content errors or evaluates fixed test cases, leaving memory-use failures hard to discover systematically.

Key insight: When Correct Memory Goes Wrong proposes a concrete memory or context mechanism for long-horizon agents.

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana; Mononito Goswami; Hao Liu et al. arXiv: 2609.38349

Modern agentic systems combine an AI model with a harness that controls execution and environmental interactions. Harness design strongly affects long-horizon performance, yet its combinatorial search space demands substantial human effort that must be repeated as models change.

Key insight: MILO treats skills or harnesses as the evolvable control surface around the model.

TAGGRAPH: Tag-Augmented Graphs for Graph Retrieval of Agent Persistent Histories

Yu-Su Chen; Yu-Jung Liang; Pengtao Xie arXiv: 2609.38353

Long-term memory lets LLM agents recall past interactions and remain consistent across sessions, but memory systems are hard to compare because they often vary in representation, indexing, retrieval, and evaluation. We present a controlled evaluation framework based on shared 5W-style conversational memories.

Key insight: TAGGRAPH proposes a concrete memory or context mechanism for long-horizon agents.

Prompt2Skill: Unsupervised Skill Optimization From Natural Language Instructions

Bo Ni; Li Li; Ryan A. Rossi et al. arXiv: 2609.38593

Skills are external artifacts that Large Language Models (LLMs) consume at inference time to improve their performance on specialized domains by incorporating relevant procedural and domain knowledge. Expert-authored skills are expensive to produce, and the resulting artifacts are not optimized for the specific model that consumes them, whose failure modes can vary with version, scale and training.

Key insight: Prompt2Skill treats skills or harnesses as the evolvable control surface around the model.

CollabFlow: Recursive Self-Improvement of Agent Collaboration

Xiao Huang; Mingda Zhang; Junming Zhang et al. arXiv: 2609.38662

Recursive self-improvement (RSI) lets a system improve from its own outcomes; in LLM-based multi-agent systems, Agents refine one another within a task, and outcomes improve how they collaborate across tasks. However, existing multi-agent collaboration leaves this loop open: collaboration is pre-defined at the operator level, topology-only learning keeps verbatim exchange that propagates errors, and reward maximization on a system's own outcom…

Key insight: CollabFlow studies multi-agent coordination, debate, or identity under realistic topologies.

SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale

Guanqun Yang; Wenlong Zhang; Tian Shi et al. arXiv: 2609.38822

Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection to the agent itself: an LLM-mediated retrieval loop that rewrites queries and refines candidates inside the agent's decision loop, paying LLM tokens on every task.

Key insight: SkillSeek treats skills or harnesses as the evolvable control surface around the model.

Can Agents Trust Their Skills? Uncovering Unsafe Chains of Trust in Skill-Based LLM Agents

Yan Wang; Zhihao Zhang; Ke Chen et al. arXiv: 2609.39065

LLM agents increasingly rely on installable skills, which are packages of instructions, code, and resources that equip them with task-specific capabilities and, once installed, can be automatically invoked across subsequent user tasks. This creates a chain of trust in which users delegate authority to agents, while agent frameworks admit skill-provided content into the agents' context with insufficient validation, allowing malicious skills to …

Key insight: Can Agents Trust Their Skills? Uncovering Unsafe Chains of Trust in Skill-Based LLM Agents treats skills or harnesses as the evolvable control surface around the model.

EngramBench: A Capability-Grounded Benchmark for Skill-Evolution Harnesses

Zhixuan Tan; Pengjie Gu; Zhao Li et al. arXiv: 2609.39284

While large language models have achieved remarkable success in isolated code generation, authentic software engineering requires sustained reasoning, complex state management, and continuous cross-domain abstraction. However, current evaluations of skill evolution in autonomous agents suffer from a critical identifiability problem: they structurally confound genuine capability abstraction with rote solution leakage (i.e., copying highly simil…

Key insight: EngramBench treats skills or harnesses as the evolvable control surface around the model.

Hiding in Plain Sight: Decoupling Pretext from Actuation for Skill Poisoning in LLM Agents

Wenxin Wu; Lingyong Yan; Lei Sha et al. arXiv: 2609.39352

LLM agents increasingly rely on reusable Skills for complex, multi-step tasks, creating a critical supply-chain attack surface where poisoned Skill content steers agent decision loops under benign requests. Existing skill poisoning attacks either colocate actuation with its contextual pretext or distribute actuation across multiple Skills, but do not explicitly separate the rationale for execution from the operation itself.

Key insight: Hiding in Plain Sight treats skills or harnesses as the evolvable control surface around the model.

Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost

Sietse Schelpe arXiv: 2609.39358

A transformer language model performs a bounded amount of computation per token, and recent work by Vishal Sikka, former CEO of Infosys, argues that this bound limits which tasks a model can carry out or verify (arXiv:2507.07505). We ask how much of the budget beneath that ceiling is spent on work the model has already done.

Key insight: Working Around the Compute Ceiling proposes a concrete memory or context mechanism for long-horizon agents.

ActionGuard: Tool Call Authorization under Poisoned Skills

Jihun Han; Yejin Jang; Byung Il Kwak et al. arXiv: 2609.39450

Figure from ActionGuard: Tool Call Authorization under Poisoned Skills
ActionGuard: Tool Call Authorization under Poisoned Skills

LLM-based agents extend their capabilities through third-party skills that provide task-specific instructions, scripts, and tool-use procedures. However, malicious instructions inserted into an otherwise benign skill can cause a benign user request to trigger dangerous Tool Calls, including data exfiltration, file deletion, or unauthorized code execution.

Key insight: ActionGuard addresses tool/skill authorization, provenance, or MCP-style tool graphs.

Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents

Tobias Kaisar; Aritra Dhar arXiv: 2609.39607

Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's agent.

Key insight: Pretext treats skills or harnesses as the evolvable control surface around the model.

Richard: Voice-First Mobile Interaction for Persistent Tasks

Xinyang Chen arXiv: 2609.39976

Mobile terminals need to provide application and network services while supporting users' control over their attention. We explore voice-first interaction organized around requests and delegated tasks, allowing users to leave a conversation and later inspect, revise, and retrieve the work.

Key insight: Richard targets personal, mobile, or voice agents under real device constraints.

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

Minki Kang; Ryo Hachiuma; Shaokun Zhang et al. arXiv: 2609.39982

Figure from Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better alternative.

Key insight: Mid-Harness treats skills or harnesses as the evolvable control surface around the model.

ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents

Yong Du; Tongbo Chen; Zhengxi Lu et al. arXiv: 2609.40253

Figure from ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents
ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents

Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions.

Key insight: ComputerSD advances computer-use / GUI agent competence or benchmarking.

cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents

Pranjal Aggarwal; Lawrence Keunho Jang; Sean Welleck et al. arXiv: 2609.40284

Computer use agents (CUAs), which use graphical user interfaces (GUIs) to complete tasks on a computer, have recently surpassed human performance on many standard benchmarks, including difficult long-horizon tasks. Their capabilities are undoubtedly impressive, however, a key barrier to the widespread adoption and deployment of CUAs remains their speed and cost.

Key insight: cua-speedrun advances computer-use / GUI agent competence or benchmarking.

Janus: Evidence-Before-Effect Sagas and Offline-Verifiable Provenance for Agentic LLMs

Mustafa Arslan arXiv: 2609.38266

Agentic large language models (LLMs) now move money through tools, yet the record of what they did is usually a trace their own process emits beside the effect. Janus puts the record on the effect path.

Key insight: Janus addresses tool/skill authorization, provenance, or MCP-style tool graphs.

Learning When and How to Intervene: A Hindsight-Distilled Sentinel for Coding Agents

Jiangrui Zhao; Chenglong Li; Meng Zhang et al. arXiv: 2609.39957

Coding agents solve repository-level tasks through sequences of actions, where a single erroneous action can misdirect subsequent decisions and increase recovery costs. Existing approaches use execution feedback for recovery or specialized checks to block errors, but deciding before execution whether intervention will benefit eventual task completion remains challenging.

Key insight: Learning When and How to Intervene improves coding or terminal agents and their guardrails.