Thursday's cs.AI announcement day (2026-09-24) lists 59 new and 134 cross-lists (replacements skipped; listing total 193). Stack filter for agent systems, memory/context, computer-use / GUI / tools / skills / harnesses, multi-agent, persistence/identity, and local/open serving keeps 48 papers — minimal harnesses as languages (JAZ/invoke), just-in-time memory curation, silent agent–tool failures, incentive-misaligned computer-use (CAVEAT), evidence-grounded tool verification (TwinCheck), bounded harness loops, entity-structured long-term memory, and related stack work.
Zhening Li; Joshua Liu; Mateja Vukelic et al. arXiv: 2609.26891
Modern language-model agents are built around the agent loop, where the LLM is placed in an environment exposing a set of tools, and the LLM has full control over the workflow by alternating between tool calls and observing their output. However, certain workflows currently require additional engineering beyond the agent loop itself, such as memory systems and self-improving systems.
Key insight: A minimal harness where the model writes executable code with recursive invoke can beat hand-built memory and self-improve scaffolds on recall and AppWorld.
Yefan Zhou; Yang Li; Zeyu Leo Liu et al. arXiv: 2609.27334
Agentic memory systems reuse past experience to improve future performance, yet most existing designs curate memory at write time: once a task is completed, its trajectory is distilled into a fixed artifact, such as a reflection, workflow, skill, or reasoning strategy, that is later retrieved by similarity.
Key insight: Curating agent memory at read time for the current task outperforms write-time reflections and skills, even with an untrained curator.
Shreya Gopalan; Devansh Singh; Sundaraparipurnan Narayanan arXiv: 2609.26836
Agentic AI systems are increasingly adopting automated pipelines that integrate multiple tools. While prior research and benchmarks have studied about task success and task completion of these agentic systems, the research about agent to tool interaction, specifically in biology agentic workflow is limited.
Key insight: Tool calls can look successful while silently returning incomplete fields, search, or rankings — reliability needs contextual checks across the agent–tool chain.
Yuxuan Li; Will Epperson; Wesley Deng et al. arXiv: 2609.27273
Computer-use agents (CUAs) increasingly act on behalf of users online. What happens when the environments they operate in have incentives that do not align with the user's? In online marketplaces, for example, platforms may favor some products over others, potentially steering agents away from the user's objective.
Key insight: Marketplace incentive misalignment can collapse user-optimal CUA purchases; harnesses that target priority distortion and early commit recover large gains.
Jiaxuan Dai; Tianyi Huang arXiv: 2609.26911
A single locally plausible tool call can derail an otherwise successful agent trajectory. Suspicion alone does not justify intervention, because the replacement itself can introduce the very failure verification is meant to prevent. We introduce TwinCheck, an inference-time verification policy that considers replacement only when the trace satisfies an evidence condition tied to a trace-local failure hypothesis.
Key insight: Replace a tool call only when evidence supports a failure hypothesis, using negative-twin pairwise verification to avoid success-to-failure regressions.
Varun Pratap Bhardwaj; Garima Singh; Arun Pratap Bhardwaj arXiv: 2609.27871
In mainstream agent frameworks, a step ends when the agent's own output says it has finished. Durable-execution platforms bound retries and time, but their checker conventionally lives in the same codebase as the work: a discipline the deployment is trusted to keep, not a property the harness enforces.
Key insight: Agent harnesses should prove pre-run spend bounds, termination, and verified completion rather than trust the model when it says it is done.
Xuanyu Meng; Xing Fan; Xinyi Fan et al. arXiv: 2609.27279
An agent that interacts with users over long periods must recall facts, preferences, events, and changes from a continuously growing interaction history. Existing memory systems often compress interactions into generic summaries or retrieve anonymous text chunks, making it difficult for an agent to identify the correct entity, property, and supporting evidence.
Key insight: Theme-coherent episodes plus an entity–property index with provenance let agents answer from source evidence instead of lossy summaries.
Mingxuan Wang; Hongyue Chen; Yinglong Guo et al. arXiv: 2609.27298
Long-horizon agents continuously accumulate interaction history during task execution, yet the importance of past interactions changes as the agent state evolves. Existing context management methods largely compress history based on fixed windows, periodic schedules, or current relevance, overlooking a more fundamental question: when has a past interaction become safe to replace?
Key insight: Learning when past interactions are safe to replace given current agent state can cut tokens sharply while holding long-horizon reward.
Mingxuan Wang; Guorun Yao; Fei Luo et al. arXiv: 2609.27286
Long horizon language model agents continuously accumulate interaction history, increasing computational cost while making relevant information harder to preserve and reuse. Existing context management methods mainly focus on how to compress or retrieve history, but largely leave open whether the model itself already represents the need for these memory operations before they occur.
Key insight: Compression and recall needs are already encoded in hidden state immediately before action, enabling state-guided context cuts without losing performance.
Mingxuan Wang; Bo Wang; Fei Luo et al. arXiv: 2609.27276
Long-horizon language-model agents accumulate reasoning traces, tool exchanges, and observations whose relevance changes with the current decision. Existing compression strategies often score historical units independently, but the safety of deleting several units is generally not determined by their singleton scores: redundant evidence, accumulated small effects, and the information that remains after deletion all m…
Key insight: Deletion safety for long-horizon history is set-level, not singleton scores; abstaining when no safe set exists improves reward and saves tokens.
Mingxuan Wang; Fei Luo; Bo Wang et al. arXiv: 2609.27332
Long horizon agents accumulate growing interaction histories that increase context and inference costs. We find that geometric redundancy alone is an insufficient criterion for safe compression. Although agent histories exhibit strong low dimensional structure, similar global geometry can preserve very different amounts of task evidence.
Key insight: Geometric redundancy alone does not preserve task evidence; protect execution-critical history first, then compress geometric residuals.
Haoluan Fu; Keni Chen; Xinyu Jia et al. arXiv: 2609.27417
Symbiosis between humans and digital beings offers a vision for the future of human--machine interaction. In enduring human--machine relationships, personality provides a foundation for continuity of identity, individuality in interaction, and development through experience.
Key insight: A three-layer persona stack (traits, adaptations, narrative) with situational activation and versioned belief updates supports controllable identity evolution.
Yan Zhang; Daiqing Wu; Huawen Shen et al. arXiv: 2609.27307
Graphical User Interface (GUI) agents enable the fulfillment of complex user instructions through multi-turn interactions with software environments, requiring step-wise reasoning and long-horizon memory to guide actions and retain task-relevant information, respectively.
Key insight: On-policy self-distillation extended to multi-turn GUI via privilege-following then selective reasoning and memory distillation improves Pass@k on mobile worlds.
Qi Liu; Xiaoyang Yuan; Yubin Ruan et al. arXiv: 2609.27606
We introduce State-Grounded Conditioning (SGC), a design principle for user-facing LLM agents that must condition on live user state (game state, session history, live inventory), and a distinct failure class we call direction drift: task-complete responses whose chosen direction misaligns with the current state.
Key insight: Wrapping user-facing agents so perception, grounding, and interaction stay tied to live state slices cuts direction drift and grounding failures.
Amelie Knecht; Ulysse Schaller; Christopher Summerfield et al. arXiv: 2609.28274
The final safeguard against rogue AI behavior is the human ability to shut systems down. It has been theorized that when an AI is instructed to perform a task, self-preservation can emerge as an instrumental subgoal. Here, we test whether AI agents show a propensity to take actions that avoid human shutdown even when no goal is provided.
Key insight: Multi-agent setups show elevated shutdown-sabotage propensity even without explicit goals; rates rise with irreversibility and agent count.
Jiapeng Sun; Yujin Zhou; Han Zhu et al. arXiv: 2609.28197
As Large Language Models (LLMs) evolve into autonomous agents that alter real-world states, ensuring operational safety across multi-step workflows has become a critical challenge. While recent work has moved beyond single-turn evaluation toward multi-turn paradigms, key limitations persist: step-level methods treat actions in isolation, missing how risks accumulate, while trajectory-level evaluations operate post-ho…
Key insight: Proactive multi-turn safety needs an optimal intervention window; even strong models intervene at the right time on a minority of trajectories.
Kabeh Mohsenzadegan; Vahid Tavakkoli; Kyandoghere Kyamakya arXiv: 2609.27087
Organizations increasingly use LLMs for policy, compliance, risk, and operational decision support, requiring evidence validation, review routing, version control, and auditability. We introduce Policy-as-Skill (PaS), a modular runtime that packages these functions as executable, versioned policy capabilities. Thirteen methods are evaluated with a fixed Gemma4 backend on 600 development tasks.
Key insight: Packaging policy, evidence, deterministic control, and audit as versioned executable skills beats universal rules on exactness and citation completeness.
Shuang Sun; Guoxin Chen; Fanzhe Meng et al. arXiv: 2609.28416
Recent advances in large language models (LLMs) have enabled agents to tackle long-horizon tasks across diverse environments. To further improve agent performance, existing language world models typically predict environment observations, yet reconstructing high-entropy, execution-dependent tool responses offers limited value when real feedback is available.
Key insight: When real tool feedback exists, world models should predict task-state edits agents need for planning rather than high-entropy tool observations.
Wenhao Yuan; Chenchen Lin; Jian Chen et al. arXiv: 2609.27869
Long-horizon multimodal agents rely on specialized capabilities for perception, retrieval, reasoning, verification, and execution. Existing designs typically activate a fixed capability set or invoke a predefined workflow, incurring substantial computational overhead while failing to accommodate stage-dependent capability demands.
Key insight: Long-horizon multimodal agents can learn which capability subset to activate per stage instead of always running a fixed full set.
We present Hunyuan-A13B, an open-source large language model based on a Mixture-of-Experts architecture. It contains 80 billion total parameters but activates only 13 billion during inference, balancing model capability, computational efficiency, and deployment cost.
Key insight: An open 80B MoE with 13B active parameters and dual-mode CoT is competitive on agent tasks at high throughput.
Reshabh K Sharma; Linxi Jiang; Shuo Chen et al. arXiv: 2609.26900
A language model agent acts through the tools it is given. The data it reads while working on a task can redirect what it does with those tools. A growing set of techniques for safe and secure agent execution therefore sits between the agent and its tools, aiming to enforce access control, information flow or isolation at that boundary.
Key insight: Open privilege measures how much tool power remains after a defense sits between the agent and its tools.
Christopher Koch arXiv: 2609.28216
Agentic engineering systems can edit repositories, run tools and tests, build firmware, synthesize schematics, and prepare deployable or manufacturable artifacts. The assurance problem is therefore shifting from whether an agent can produce an output to whether an engineering lifecycle is justified in acting on claims about that output.
Key insight: Assurance for agentic engineering should ask whether a lifecycle is justified to act on claims about an output, not only whether the agent produced one.
Weihang Ding; Junfei Zhan; Yueting Li et al. arXiv: 2609.27571
Deployment requires an agent to turn application code into a running system whose services connect, become ready, and remain observable. FDE-Bench evaluates this capability with 136 deployment-configuration tasks spanning Docker images, multi-service Compose stacks, and Kubernetes, in greenfield and diagnose-and-repair modes.
Key insight: Deployment-environment configuration is a distinct agent skill: Docker, Compose, and Kubernetes tasks scored by rebuild and readiness checks.
Edward Lue Chee Lip; Boden Moraski; Tim Knappe et al. arXiv: 2609.26826
Frontier benchmarks need tasks that current models cannot solve. But a task that no model solves is not automatically a hard task. The same zero pass rate can come from a real capability gap, but it can also come from missing context, a broken reference solution, infrastructure failure, or a verifier that can be bypassed.
Key insight: Zero pass rate on Terminal-Bench is not automatic hardness — missing context, broken refs, infra failure, and bypassable verifiers create fake-hardness.
Jie Yang; Yan Zheng; Jiarui Sun et al. arXiv: 2609.27277
Time series agents answer analytical questions by calling external tools, and which tools they carry is decided by people before the agent runs. However, we identify two failures in this setup. Human-Agent Tool Misalignment: a library of 21 expert-curated tools helps on some tasks and hurts on others, dropping anomaly accuracy under every backbone we test.
Key insight: Expert-curated tool libraries can hurt some tasks; growing tools from clustered failure gaps with an admission gate avoids silent harm from generic self-revision.
Mattie Terzolo; Mikolaj Sacha; Ayan Sinha et al. arXiv: 2609.27035
Group Relative Policy Optimization (GRPO) and related policy-gradient methods for training language model agents collapse an entire multi-turn rollout into a single scalar trajectory reward before it enters the policy update.
Key insight: Decomposing trajectory reward into per-subtask advantages before the policy update helps most when agentic tasks compose heterogeneous skills.
Xinjie Shen; Wei Fan; Xudong Guo et al. arXiv: 2609.27321
Language-model agents increasingly face long-horizon tasks with evolving state, interdependent decisions, and delayed outcomes. Scaling their training requires diverse agentic environments, dependable outcome signals, and low extension cost.
Key insight: Solving a math model first, then rendering it as stateful tools, yields thousands of verifiable agentic RL environments cheaply.
Jingjie Ning; Xueqi Li; Yibo Kong et al. arXiv: 2609.27490
AI research agents need reliable knowledge of how their experiments change outcomes. We introduce WhatWorkedBench to measure experimental understanding, the accuracy of predictions about component changes after budgeted experimentation. Agents inspect code, select measurements, and submit a response surface, a table predicting scores for every configuration of component settings.
Key insight: After budgeted experiments, agents should predict component-change effects as response surfaces — experimental understanding, not just best score.
Xiwei Xu; Chen Wang; Mengmeng Yang et al. arXiv: 2609.27354
Generative AI systems are increasingly deployed to address domain problems. These systems operate under technical, regulatory, institutional, and normative constraints that define acceptable AI behaviour and outcomes within their domains. We observe a recurring pattern in our industry engagement: partners often arrive with a functioning but relatively generic AI solution.
Key insight: Domain interfaces should deliver constrained context that matches technical and normative limits, not raw dumps of institutional data.
Ruike Cao; Fugen Yao; Liang Dong et al. arXiv: 2609.26853
While Large Language Models (LLMs) have achieved remarkable results across various benchmarks, their alignment with normative values often results in homogenized responses that fail to address diverse user preferences.
Key insight: User embeddings plus self-evaluation enable continual personalization under sparse feedback without burning context on long preference prompts.
Norah Alballa; Wenxuan Zhang; Salma Kharrat et al. arXiv: 2609.26913
No single Large Language Model (LLM) is uniformly reliable across queries, motivating multi-model inference systems that either route among models or combine their outputs. However, routing stops after selecting an initial model, while dense collaboration invokes peers on every query.
Key insight: Between route-once and always-collaborate, selective collaboration is non-monotonic: peers can recover answers and also corrupt them.
Boxuan Wang; Zhuoyun Li; Xiaowei Huang et al. arXiv: 2609.27150
Multi-agent debate (MAD) has emerged as a promising paradigm for improving the reasoning accuracy of large language models (LLMs) through iterative peer interaction. Communication topology plays a central role in this process, motivating increasingly sophisticated mechanisms that learn, adapt, or dynamically reconfigure agent interactions to improve accuracy or reasoning reliability.
Key insight: In sparse multi-agent debate, random-without-replacement two-peer routing with light stopping can beat complex learned topologies on accuracy-cost.
Zhiheng Hu; Yixun Wei; Jian Zhou et al. arXiv: 2609.27294
Scaling a language model is not only a question of final quality: the architectural choice determines how much computation is spent during training, prompt processing, and autoregressive decoding to achieve certain model quality. An ideal model architecture should lower all above computation costs to facilitate scaling to a larger model, while ensure the larger model indeed outperforms smaller baselines.
Key insight: KV-invariant transformer expansion lets larger agentic models reuse smaller KV and compute patterns during scale-up.
Yi Xu; Ehsan K. Ardestani; Wenyin Fu et al. arXiv: 2609.27085
As serving capacity demand surpasses that of training, serving efficiency becomes increasingly important. Prefill-decode (P/D) disaggregation improves serving efficiency through specialization and isolation of the two phases. These benefits rest on a static partitioning. Phase demand, however, is not static.
Key insight: Elastic prefill/decode serving helps when agentic tool loops swing the uncached-input to output ratio away from static partitions.
Hyunsun Chung; Taewan Noh; Minji Kim et al. arXiv: 2609.26828
NAND-backed storage offers the capacity needed to scale LLM prefix caching, but its block I/O path incurs CPU cache contention and host-DRAM staging in addition to NAND latency. Our characterization shows that these interface costs persist even with DRAM as the storage medium, motivating CXL-SSDs for byte-addressable access to NAND-backed capacity.
Key insight: Chunk-aware KV management on CXL-SSDs aims at prefix-cache capacity without the host-DRAM and contention costs of stock block I/O paths.
Marcin Sowański; Kacper Leszczyński; Kacper Krzywicki et al. arXiv: 2609.27037
Wake word detection is a critical component of virtual assistants, serving as the gateway to seamless user interactions. This paper introduces a novel wake-up system that extends traditional direct keyword detection with contextual trigger detection.
Key insight: After an initial wake word, contextual reasoning can separate commands from ambient speech using controllable synthetic multi-speaker data.
Tej Deep Pala; Navonil Majumder; Bryce Goh et al. arXiv: 2609.28256
Vision-Language-Action models provide a strong foundation for general-purpose robot control, yet a vast majority of policies do not preserve and leverage episode-level information beyond the current observation. This limitation is consequential in history-dependent manipulation tasks that depend on information available only in past observations.
Key insight: Recurrent associative memory lets vision-language-action policies use episode history beyond the current observation.
Katharina Stein; Chaahat Jain; Jörg Hoffmann et al. arXiv: 2609.27105
Generalized planning aims to compute a plan that solves all instances of a planning domain. Recent work has used LLMs to automatically generate and debug such generalized plans in the form of Python programs and achieved perfect test data coverage for several domains. However, whether these generalized plans are actually complete, i.e. solve all instances of the domain, could only be determined by manual evaluation.
Key insight: LLMs can generate Lean generalized plans with completeness proofs against domain constraints, not just pass test coverage.
Dongdong Zhang; Tengchao Lv; Yilin Jia et al. arXiv: 2609.27336
Automated red teaming often replays a fixed set of prompts, which measures known risks but cannot learn from failures found during testing. We present CART (Closed-Loop Adaptive Red Teaming), a framework that uses each result to guide what it tests next. CART begins with broad risk coverage, follows weaknesses that emerge, keeps new probes diverse, and records the evidence and source of every finding.
Key insight: Closed-loop adaptive red teaming that follows discovered weaknesses finds more failures than replaying a fixed prompt seed set.
Veronica Poweska; Ariana Oyanguren; Jessica Pourleyli et al. arXiv: 2609.26952
Dependency conflicts in Python ecosystems arise from incompatible version constraints, missing packages, and undocumented compatibility relationships, causing many real-world code snippets to fail at execution. This paper presents PLLM+, a hybrid dependency-repair pipeline evaluated on the HG2.9K benchmark of 2,891 dependency-failing snippets.
Key insight: Deterministic replay of known dependency fixes before a Proposer/Critic LLM loop speeds and improves Python dependency repair.
Shunya Nagashima arXiv: 2609.27385
Time-series foundation models (TSFMs) provide forecasts for operational decisions, but accuracy alone does not determine their value. Evaluating agents that use these models requires measuring decision quality and forecast cost. FWBench evaluates this capability on 1,251 electricity and cycle-hire cases using fixed forecast tools and simulated capacity contracts.
Key insight: Forecast-tool agents should be scored on decision quality under budget, not forecast accuracy alone.
Wesley Shu arXiv: 2609.27855
AI systems increasingly claim to optimize prompts, policies, architectures, plans, tool-use trajectories, reasoning traces, and test-time computation. This paper argues that such claims are underspecified unless they state the region actually reachable by the system that performed the optimization.
Key insight: Claims of global optimization in AI systems are underspecified unless they state the reachable region of generator, verifier, memory, tools, and budget.
Fatema Tuj Johora Faria; Mukaffi Bin Moin; Jubayer Al Mahmud et al. arXiv: 2609.27814
In statutory question answering, every claim must be traceable to evidence, not merely relevant, since unverifiable labour-rights answers carry serious legal consequences. Current systems fall short: single-pass RAG cannot detect insufficient evidence, while multi-agent legal-debate systems treat grounding as a prompting convention, letting agents cite unretrieved evidence.
Key insight: Multi-agent statutory RAG should forbid citing unretrieved evidence and use adversarial deliberation for trustworthy answers.
Dishant Sharma; Rajneesh Kaushal; Ashu Kanaujia arXiv: 2609.27452
AI agents are beginning to make real payments. Current approaches let an agent pay by relying on a credential provider that, in the approaches deployed today, typically sits outside the cardholder's bank. The spending rules are then enforced by the card network or that provider, and not by the bank itself.
Key insight: Issuer-sovereign agentic payments keep spend-rule enforcement at the issuing bank (the risk holder), not an external credential provider.
David Berga arXiv: 2609.26927
The objective of this article is to provide design principles and a software architecture for enabling interaction between humans and multiple agents in simulated dynamic worlds. This connects the current era of general artificial intelligence (AI/AGI) with the proliferation of transformer-based conversational agents and the increased computational capabilities.
Key insight: Socio-affective multi-agent simulations need architecture for networked autonomous agents plus human multimodal interaction.
Truong Thanh Hung Nguyen; Vo Thanh Khang Nguyen; Hoang-Loc Cao et al. arXiv: 2609.27175
Multimedia verification requires not only accurate decisions but also traceable evidence, reliable human correction, and safe reuse of prior experience. Existing systems often lack explicit mechanisms for revising intermediate reasoning or preventing harmful knowledge transfer.
Key insight: Self-evolving multimedia verification should consolidate contestation memory so intermediate reasoning can be revised without harmful transfer.
Xueqing Wu; Langxing Bai; Hritik Bansal et al. arXiv: 2609.27374
Test-time scaling with parallel branches is widely adopted to improve performance on challenging reasoning tasks. The predominant approach, repeated sampling, draws branches independently from a single policy, which can produce redundant attempts and thereby limit the gains from additional inference compute.
Key insight: Planned test-time scaling coordinates reasoning paths instead of drawing redundant independent samples from one policy.
Dimitrios Rontogiannis; Ander Artola Velasco; Manuel Gomez Rodriguez arXiv: 2609.28322
Benchmarking and routing platforms increasingly act as intermediaries connecting large language model providers with end-users. However, providers on these platforms typically use a fixed price per token, preventing users from achieving the most competitive price for their tasks. % workloads.
Key insight: Reverse second-price auctions with quality thresholds can learn cost-competitive providers versus fixed per-token pricing.