Wednesday's cs.AI announcement day (2026-09-30) lists 207 new and 299 cross-lists (replacements skipped; listing total 506). Stack filter for agent systems, memory/context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent, persistence/identity, and local/open serving keeps 42 papers — meta-reasoning for agentic inference, meta-skills for harness design, local LM self-checks (HARISSA), context compression (FOCUS, CLMs), skill/harness evolution (RASO, MoSI Branches, SafeCoEvo), and agent memory retrieval (UpliftMem).


Research Papers

Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language Models

Ferreira, Leonardo; Liu, Gardenia; Zheng, Kaden arXiv: 2609.35875

Multi-agent debate (MAD) reportedly improves reasoning and factuality over single-model inference, but prior work treats agents as symmetric peers, leaving open what drives the gains. We test the hypothesis that cognitive diversity among agents is the driver, in the setting where the question is still measurable: small open-weight models with benchmark headroom.

Key insight: Beyond Symmetric Agents studies multi-agent coordination, debate, or identity under realistic topologies.

Harness Evolution as Learning: Approximation, Generalization, and Optimization Limits of Self-Improving Personal Agents

Gan, Zeyu; Gong, Zixuan; Liu, Yong arXiv: 2609.36892

Figure from Harness Evolution as Learning: Approximation, Generalization, and Optimization Limits of Self-Improving Personal Agents
Harness Evolution as Learning: Approximation, Generalization, and Optimization Limits of Self-Improving Personal Agents

As the capabilities of large language models (LLMs) continue to advance, increasing attention is turning to how to translate their abilities into useful behavior. Personal agents bring this question into everyday settings, where models are expected to serve individual users and continually adapt to their preferences.

Key insight: Harness Evolution as Learning treats skills or harnesses as the evolvable control surface around the model.

How Can Recommendation Feedback Evolve Agent Memory?

Mao, Shanwen; Li, Mingming; Zhang, Hao et al. arXiv: 2609.37544

Content-generation agents continuously receive impressions, clicks, conversions, and negative feedback from recommendation systems, providing real-world outcome signals for memory evolution. However, these signals are delayed and noisy, confounded by audience composition, placement, and recommendation policies, and may result from the combined influence of multiple memories, making accurate attribution difficult.

Key insight: How Can Recommendation Feedback Evolve proposes a concrete memory or context mechanism for long-horizon agents.

Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL

Huang, Youling; Xu, Tiankuo; Liu, Jiaji et al. arXiv: 2609.37898

Reinforcement learning for long-horizon agents typically relies on sparse outcome-based rewards. This leads to a severe cold-start problem, as early-stage policies often fail to solve sampled tasks, leaving little useful reward signal for learning.

Key insight: Guide, Then Let Go advances long-horizon planning, RL, or execution fidelity for agents.

Video-RSI: Recursive Self-Improvement of Video Understanding Agents via Harness Evolution

Luo, Bingjun; Guo, Jialin; Li, Siqi arXiv: 2609.37950

Video understanding agents acquire evidence through an executable harness that controls what they observe and how they use those observations. However, execution traces contain only the evidence acquired by the current harness, leaving competing explanations for failure unresolved and limiting the basis for self-improvement.

Key insight: Video-RSI treats skills or harnesses as the evolvable control surface around the model.

Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning

Dahal, Paras; Bakhtin, Anton; Cohen, Taco et al. arXiv: 2609.38147

As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to stop.

Key insight: Thinking Before Thinking treats skills or harnesses as the evolvable control surface around the model.

BRIDGE: Bilevel Retrieval-Credit-Aware Agentic Reinforcement Learning

Xiao, Quan; Liu, Mingda; Liu, Gaowen et al. arXiv: 2609.36505

Agentic reinforcement learning (ARL) with verifiable rewards improves the ability of large language models (LLMs) to tackle knowledge-intensive tasks by learning to interleave search and reasoning. However, most existing ARL methods optimize only LLM-generated tokens and treat retrieved evidence as environment observations.

Key insight: BRIDGE proposes a concrete memory or context mechanism for long-horizon agents.

UpliftMem: Learning Set-Level Uplift for Agent Memory Retrieval

Liang, Mengkun; Qiang, Haoran; Liu, Guannan et al. arXiv: 2609.36805

Figure from UpliftMem: Learning Set-Level Uplift for Agent Memory Retrieval
UpliftMem: Learning Set-Level Uplift for Agent Memory Retrieval

Large language model (LLM) agents reuse external memory to guide new tasks, but effective retrieval requires learning which memory sets improve execution. Such learning relies on costly outcome feedback: ordinary retrieval observes only executed sets, while evaluating alternatives requires additional rollouts.

Key insight: UpliftMem proposes a concrete memory or context mechanism for long-horizon agents.

ARC-KV: Amortizing Anchor Search for Reconstruction-Based KV Cache Compaction

Shen, Zheyu; Wang, Guanhua; Tu, Dezhan et al. arXiv: 2609.36835

Long-context large language model inference is bottlenecked by KV caches that grow linearly with sequence length. This burden is especially severe for long, reusable context prefixes, whose cache must serve many downstream queries.

Key insight: ARC-KV proposes a concrete memory or context mechanism for long-horizon agents.

SkillCome: Group Contrast Skill Optimization with Dual Memory

Li, Haolin; Hong, Feng; Li, Ang et al. arXiv: 2609.37128

Skill evolution improves the capabilities of large language models by analyzing trajectories generated under a given skill and modifying the skill accordingly. Existing approaches typically generate a single trajectory per question.

Key insight: SkillCome treats skills or harnesses as the evolvable control surface around the model.

Mixture of Self-Improving Branches For Agent Harness Optimization

Dong, Haoyu; Zhou, Yuhang; Lin, Zihao et al. arXiv: 2609.37834

Figure from Mixture of Self-Improving Branches For Agent Harness Optimization
Mixture of Self-Improving Branches For Agent Harness Optimization

Harness optimization provides a practical setting for recursive self-improvement (RSI), where agent-generated modifications inform subsequent changes through execution feedback. Recent work such as Meta-Harness implements this process through iterative code generation and evaluation, but retains a fixed development set and proposal policy.

Key insight: Mixture of Self-Improving Branches For treats skills or harnesses as the evolvable control surface around the model.

You Cannot Pick a Provider From the Price List: Market-Aware Routing for Open-Weight LLM Inference

He, Liang; Wen, Jingbo; Chen, Yixiong et al. arXiv: 2609.37902

Existing LLM routers choose among models using static per-model costs. We show that open-weight inference markets introduce a second, largely ignored decision axis: after choosing a model, a client must still choose which provider serves it.

Key insight: You Cannot Pick a Provider targets safer or cheaper local/open-weight deployment for agent stacks.

HARISSA: Inference-Time Self-Checks for Efficient and Safe Local Language Model Deployment

Alkiek, Kenan; Lee, Moontae; Jurgens, David et al. arXiv: 2609.38006

Running a language model locally offers advantages in privacy, latency, and cost, but local hardware fits only small models, which are less capable than frontier models. The usual remedy for a hard query, escalating it to a cloud model, gives up the privacy and cost advantages of running locally.

Key insight: HARISSA targets safer or cheaper local/open-weight deployment for agent stacks.

Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation

Chu, Jaewon; Lee, Ji Soo; Park, Jihwan et al. arXiv: 2609.38024

Figure from Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation
Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation

An agent skill is a reusable, actionable natural-language artifact that guides an agent to perform a task effectively under a given harness. Recent studies have explored the optimization of agent skills, contributing to a growing collection of publicly available skills spanning diverse tasks, domains, and harnesses.

Key insight: Retrieval-Augmented Skill Optimization via Cross-Harness treats skills or harnesses as the evolvable control surface around the model.

UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training

Jain, Ashish; Sandhu, Armaan arXiv: 2609.38043

Interactive agent benchmarks and multi-turn reinforcement learning increasingly place a second language model in the role of the user. This simulated user controls what information the agent receives and when, yet current benchmarks score only the agent and do not directly measure whether the user correctly executed its assigned role.

Key insight: UserProxyBench adds a practical building block for LLM agent systems.

Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

Qian, Cheng; Zhu, Kunlun; Li, Beibin et al. arXiv: 2609.38143

Figure from Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI
Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

Agent performance depends on both reasoning ability and the environment in which it acts. We study test-time AI-for-AI, asking how a Builder can learn to construct better execution environments for a Target while both models' weights remain fixed.

Key insight: Learning Meta-Skills for Agent Harness treats skills or harnesses as the evolvable control surface around the model.

Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning

Jung, Dongwon; Ramesh, Hemanth Neelgund; Wang, Yifan et al. arXiv: 2609.36178

Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed to success.

Key insight: Targeting Pivotal Decisions for Credit advances long-horizon planning, RL, or execution fidelity for agents.

Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression

Zhu, Xingyu; Pu; Yi et al. arXiv: 2609.36322

Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new positional coordinate: a token's phase, or its position relative to compression-window boundaries.

Key insight: Periodic Weak Spots proposes a concrete memory or context mechanism for long-horizon agents.

SafeCoEvo: Co-Evolving Safety Harnesses and Guards for LLM Agents at Test-Time

Cheng, Yu; Hu, Yongkang; Ma, Shuaijie et al. arXiv: 2609.36580

Figure from SafeCoEvo: Co-Evolving Safety Harnesses and Guards for LLM Agents at Test-Time
SafeCoEvo: Co-Evolving Safety Harnesses and Guards for LLM Agents at Test-Time

LLM agents deployed in real-world environments continually encounter new tasks and safety risks, while execution feedback typically becomes available only after each task is completed. However, existing self-evolving approaches commonly rely on multiple rounds of optimization over fixed and repeatedly accessible task distributions, fundamentally differing from test-time adaptation in real-world deployment, where only experience accumulated from past tasks can be used to improve safety decisions on future unseen ...

Key insight: SafeCoEvo treats skills or harnesses as the evolvable control surface around the model.

Topological Coherence for Self-evolving Multi-agent Systems

Zhao, Sen; Kong, Ruiqi; Zhang, Zuyu et al. arXiv: 2609.37953

Complex tasks inherently couple workflow structure, agent responsibility, collaboration, and memory access: task regions delimit responsibility and tool scope, cross-region dependencies give rise to handoffs, and ownership boundaries delimit private and selectively shared memory. Existing methods can jointly optimize agent and communication structures, yet such optimization does not by itself require responsibility, handoff, and memory boundaries to remain consistent with task dependencies.

Key insight: Topological Coherence for Self-evolving Multi-agent proposes a concrete memory or context mechanism for long-horizon agents.

MAADBench: The Refreshable Paradigm for Anomaly Detection in Multi-Agent Systems

Ma, Lei; Hofmann, Dennis; Xu, Haowen et al. arXiv: 2609.36556

Recent studies report that LLM-based multi-agent systems (MAS) fail at rates of 41%-87%, yet to our knowledge, no benchmark to date supports systematic anomaly detection (AD) for them. Building MAS AD benchmarks is hard because they must remain fresh as LLM systems evolve: tasks may leak into training data and thus be memorized by LLMs, traces and anomaly patterns expire as backbones evolve, and labels must be provided reliably for each refresh.

Key insight: MAADBench studies multi-agent coordination, debate, or identity under realistic topologies.

When Upstream Messages Override Correct Answers: A Controlled Study of Multi-Agent LLM Collaboration

Gong, Yaxin; Zhang, Gangyi; Gao, Chongming et al. arXiv: 2609.36855

Multi-agent LLM systems rely on message passing among specialized agents to accomplish complex tasks. However, an upstream agent may provide useful information or an incorrect answer that causes a downstream agent to override a correct answer supported by its own evidence.

Key insight: When Upstream Messages Override Correct studies multi-agent coordination, debate, or identity under realistic topologies.

WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents

Mao, Bo; He, Hang; Wang, Linting et al. arXiv: 2609.36887

Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent harness, and evaluator. Scaling environments in isolation, however, does not guarantee commensurate gains in model performance, because reliable learning signals depend on coherent interactions among all components of the agentic interaction system.

Key insight: WEFT improves tool-use or MCP-style orchestration for general-purpose agents.

CADOC: Cache-Aware Dynamic Object Context for Long-Horizon Agents

Yao, Junjie; Zhou, Zhangchen; Xu, Zhi-Qin John arXiv: 2609.37012

For a long-horizon agent, context is the bottleneck: the history is resent with every request, the window caps task length, and reasoning degrades as the history grows. Replacing structured objects with compact retrieval Cards shortens the prompt and keeps the exact originals retrievable, but editing the history can break prefix-cache reuse, and prior recoverable methods time their edits by forecasts of future reuse or by preset intervals.

Key insight: CADOC proposes a concrete memory or context mechanism for long-horizon agents.

Rational Clarification by Assistive Agents via Value-of-Information Reasoning

Nguyen-Hien, T. Duy; Teh, Yee Whye; Lee, Wee Sun et al. arXiv: 2609.37588

Users of language-based assistive agents often make ambiguous requests. In response, an assistant can either directly act on its interpretation of the request --- risking misalignment with the user --- or ask a clarifying question.

Key insight: Rational Clarification by Assistive Agents advances long-horizon planning, RL, or execution fidelity for agents.

FOCUS: Training-Free Decision-Preserving Context Compression for LLM Agents

Dixit, Shantanu; Bastos, Anson; Zhang, Xuchao et al. arXiv: 2609.37590

Figure from FOCUS: Training-Free Decision-Preserving Context Compression for LLM Agents
FOCUS: Training-Free Decision-Preserving Context Compression for LLM Agents

LLM agents accumulate interaction histories that grow linearly with task length, causing quadratic inference cost scaling and performance degradation from attention dilution. Existing context-compression methods learn what to discard offline: by contrastively optimizing guidelines, distilling compressors, or training compression policies.

Key insight: FOCUS proposes a concrete memory or context mechanism for long-horizon agents.

Beyond a single latent space: a dual-latent world model for long-horizon planning

Zhao, Delin; Yue, Zhengrong; Zhuang, Shaobin et al. arXiv: 2609.37644

Latent world models often struggle with long-horizon planning despite accurate short-term predictions. Recursive rollouts accumulate errors, while distance concentration in high-dimensional latent spaces can weaken goal discrimination.

Key insight: Beyond a single latent space advances long-horizon planning, RL, or execution fidelity for agents.

ContextRender: From Execution Dependencies to Agent Context

Kashmira, Savini; Dantanarayana, Jayanaka L.; Tang, Lingjia et al. arXiv: 2609.37743

LLM agents performing long-horizon tasks accumulate tool results that later steps may need. Passing the full history to every invocation is costly even when it fits within the context window, while reducing it risks omitting needed information.

Key insight: ContextRender proposes a concrete memory or context mechanism for long-horizon agents.

Character Training for Risk-Averse Agents

Dhoot, Arav; Pandey, Punya Syon; Johnson, Jamie et al. arXiv: 2609.38093

Risk aversion in resources could prevent misaligned AI agents from causing catastrophic harm. Misaligned but risk-averse agents would tend to favor safer strategies like making deals with humans over riskier strategies like rebelling.

Key insight: Character Training for Risk-Averse Agents adds a practical building block for LLM agent systems.

Do LLM Agents Execute the Plans They Declare? From Planning-Mode Declaration to Pattern-Specific Execution

Oota, Subba Reddy; Herrera, Francisco; Sagrera, Jordi Cabot et al. arXiv: 2609.38108

Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment. However, successful planning requires two distinct capabilities: selecting an appropriate plan for the task and executing it faithfully.

Key insight: Do LLM Agents Execute the advances long-horizon planning, RL, or execution fidelity for agents.

Prompted Identity Degrades Cooperation in Multi-Agent LLM Systems

Del Giudice, Xavier; Palma, Alessio; Migliarini, Matteo et al. arXiv: 2609.35928

Multi-agent LLM systems increasingly mix models from several providers, yet exposing each agent's underlying model identity to its peers significantly impairs cooperation. We show that when agents are aware of each other's model family, the group splits into clusters, where agents prefer interacting with others carrying their same label, although nothing in the task rewards or asks for such a split.

Key insight: Prompted Identity Degrades Cooperation in studies multi-agent coordination, debate, or identity under realistic topologies.

Reward-rate Policy Gradient for Efficient Machine Learning Engineering Agents

Tian, Muhang; Yang, Sherry arXiv: 2609.36393

Traditional reinforcement learning (RL) techniques focus on maximizing expected cumulative reward, where each action assumes to take a constant unit of time. However, this assumption does not hold for agentic RL tasks such as machine learning engineering (MLE) agents, where actions involve data loading, feature engineering, and model training that take variable durations.

Key insight: Reward-rate Policy Gradient for Efficient advances long-horizon planning, RL, or execution fidelity for agents.

Frontier Autolab: Organizational Memory, Adversarial Dissent and Temporal Leakage in Multi-Agent LLM Firms Across Fifty Years of Technological Change

Ghosh, Bravish arXiv: 2609.36739

Multi-agent LLM systems are increasingly structured like organizations, with roles, critics and shared memory, yet they are evaluated on tasks that last minutes. We ask how such an organization behaves when the ground it stands on keeps moving.

Key insight: Frontier Autolab proposes a concrete memory or context mechanism for long-horizon agents.

SKILLLITE: Evidence-Guided Malicious Skill Auditing with Compact LLMs

Ou, Haoran; Deng, Gelei; Zhang, Xuanye et al. arXiv: 2609.36879

As LLM-based agents perform increasingly complex tasks, Agent Skills have emerged as a flexible mechanism for extending their capabilities. An Agent Skill packages task-specific instructions with executable components and auxiliary resources to provide specialized functionalities.

Key insight: SKILLLITE treats skills or harnesses as the evolvable control surface around the model.

VACE: Validation-Gated Alternating Co-Evolution of Agent Models and Harnesses

Qi, Jiexing; He, Yu; Liu, Jun et al. arXiv: 2609.37105

Language model agents can be improved by updating their model weights or refining the harness that guides task execution. These components are coupled: weight updates change how the model uses the harness, while harness updates change the trajectories used for training.

Key insight: VACE treats skills or harnesses as the evolvable control surface around the model.

Concealing LLM-Based Multi-Agent Topology via Phantom Structure Injection

He, Longzhu; Wen, Zelang; Li, Xinfeng et al. arXiv: 2609.37567

Driven by the rapid advancement of large language models (LLMs), LLM-based multi-agent systems (MAS) have emerged as a powerful paradigm for collaborative reasoning over complex tasks. A key design element of MAS is the communication topology, which governs information flow among agents and often encodes proprietary knowledge about the system architecture.

Key insight: Concealing LLM-Based Multi-Agent Topology via studies multi-agent coordination, debate, or identity under realistic topologies.

AgentBug-Smith: Automatically Reproducing Real-World Harness Bugs in Agentic Systems

Cheng, Yiming; Rahardja, Alfin Wijaya; Zhang, Mengshi et al. arXiv: 2609.37864

Agent harness bugs exhibit unique characteristics and remain challenging for state-of-the-art software agents to repair. Progress in this area is further hindered by existing benchmarks, which contain only a small and fixed number of executable harness bugs while requiring hundreds of human hours to construct.

Key insight: AgentBug-Smith treats skills or harnesses as the evolvable control surface around the model.

Auditable Long-Term Memory: A Deterministic Retrieval Chain Measured at 479/475 of 500 on LongMemEval-S

Chanhnourack, Christopher J. arXiv: 2609.38021

We evaluate an auditable long-term memory system on LongMemEval-S. Its retrieval chain uses hybrid candidate retrieval, cross-encoder reranking, coverage-first packet compilation, and deterministic reasoning scaffolds; an LLM is used only as a replaceable final reader.

Key insight: Auditable Long-Term Memory proposes a concrete memory or context mechanism for long-horizon agents.

StateTape: Action-Conditioned Evidence Lifecycle Modeling for Long-Horizon Coding Agents

Yu, Ziyang; Zhao, Liang; Zhu, Bowen et al. arXiv: 2609.36319

Despite the recent success of coding agents built on large language models, it remains challenging to run them over long horizons, since every observation is appended to the context and the context grows with each one. History-based maintenance is a common remedy, which masks or summarizes old observations, or prunes what a model reads as useless, and bounds the context at little cost.

Key insight: StateTape advances long-horizon planning, RL, or execution fidelity for agents.

Engineering Simplicity: Simple Mechanism Interfaces Steer LLM Agents

Zhu, Kehang; Shah, Anand; Parkes, David arXiv: 2609.36365

Can interaction formats and textual scaffolds help large language model (LLM) agents make better decisions, and do better decisions come with better explanations? We study these questions in auctions and matching, multi-agent environments with explicit rules and known optimal strategies.

Key insight: Engineering Simplicity studies multi-agent coordination, debate, or identity under realistic topologies.

Context Language Models

Shao, Rulin; Shen, Shannon Zejiang; Yin, Junjie Oscar et al. arXiv: 2609.37725

Figure from Context Language Models
Context Language Models

We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file.

Key insight: Context Language Models proposes a concrete memory or context mechanism for long-horizon agents.

WitnessGym: Benchmarking Coding Agents on the Construction of Bug Witnesses

Qi, Haomin; Xu, Xiangzhe; Huang, Yiming et al. arXiv: 2609.36635

Bug validation asks a coding agent to produce an executable witness for a reported bug. The witness combines a concrete input with a testing harness and exposes faulty behavior during execution.

Key insight: WitnessGym treats skills or harnesses as the evolvable control surface around the model.