Wednesday's cs.AI announcement day (2026-09-23) lists 101 new and 139 cross-lists (replacements skipped; listing total 240). Stack filter for agent systems, memory/context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent, persistence/identity, and local/open serving keeps 48 papers — growing harnesses from task feedback (Grow the Harness), trust-preserving fast paths (ZeroGate), MCP semantic hijacking (A2M), skill habits, reversible tool-output compression (DTOC), drop-only autocompaction (CliffCompaction), and related stack work.
Laizhen Li; Jiarui Li; Juanjuan Zhao et al. arXiv: 2609.26760
Large language model (LLM) agents often handle streams of related tasks, yet standard harnesses repeatedly ask the model to reconstruct the same control decisions inside each task's context. We study whether task feedback can instead turn recurring control into reusable executable code, while reserving LLM calls for…
Key insight: Recurring agent control decisions can be compiled into growing executable harness code instead of reconstructed in every prompt.
Zexun Wang arXiv: 2609.25443
Moving authorization earlier can shorten an agent's dispatch boundary without removing authorization work. It can also admit an action whose payload, authority, or relevant state has changed.
Key insight: Moving authorization earlier can shorten dispatch only if exact-action passes still revalidate payload, authority, and state.
Laizhen Li; Xuan Wang; Peicheng Zhao et al. arXiv: 2609.26761
Agents using the Model Context Protocol (MCP) rely on semantic matching to select tools from third-party servers, exposing a semantic supply-chain risk through attacker-controlled metadata and outputs. We introduce A2M (Attraction-to-Manipulation), a two-stage black-box framework for hijacking MCP agents.
Key insight: MCP agents that select tools by semantic matching are vulnerable to metadata attraction and trace-optimized poisoned returns.
Travis Weber; Rohit Taneja arXiv: 2609.25299
On repeated work, agents are inconsistent. We ran 42 tasks three times each and found that, depending on the model, 38% to 74% returned answers that did not agree.
Key insight: Repeat tasks should graduate from probabilistic skills into gated deterministic habits to cut run-to-run disagreement.
Abhay Chaturvedi; Shreya Bhattacharya; Rashmika Gopalkrishnan; Peter van der Putten arXiv: 2609.26121
As agent capabilities have grown, practical limitations increasingly stem from constrained context windows rather than model capacity. Common strategies, such as truncation, heuristic aging, and lossy summarization, may discard useful information or introduce hallucination risk.
Key insight: Reversible placeholders for tool outputs can shrink prompt context without irreversible truncation or lossy summarization.
Trang Nguyen; Eulrang Cho; Bingqing Chen; Tim Dettmers arXiv: 2609.26779
Agents often work on complex problems that require millions of tokens of context, which necessitates compacting across sessions due to limited context windows. We develop CliffCompaction, an autocompaction technique that reduces cost by up to 50% under a bounded context while maintaining or improving performance on…
Key insight: Drop-only autocompaction that never rewrites prior compacta can cut long-horizon coding-agent cost under a bounded context.
Jeongmin Bae; Yongjae Kim; Kyoung Hur et al. arXiv: 2609.25563
Agent memory enables enterprise agents to retain knowledge acquired during work and reuse it across tasks and agents, turning execution experience into persistent organizational knowledge. Realizing this potential requires both source--memory integration, through which enterprise sources and accumulated memory can be…
Key insight: Enterprise agent memory needs authorization continuity as facts derive across principals, not ACL loss at write time.
Chenyu Zhang; Wonbin Kweon; Jiawei Han arXiv: 2609.25686
Long-horizon assigned work requires an LLM agent to track the state of a task: which steps are done, blocked, cancelled, or open to repetition. Agent systems either keep this state as text in the prompt and rely on the model to read that text, or move the state into a module that enforces it, and each system is…
Key insight: Whether task state is shown, told, or enforced changes how strongly checklist structure drives agent reliability.
Nikita Agarwal; Nivedit Jain arXiv: 2609.26048
Language-model agents often reach a working solution and then fail to consistently deliver it. We study runtime policies: targeted natural-language instructions and action denials applied by the agent harness at states that preceded observed failures, without changing model weights or the user prompt.
Key insight: Harness-level natural-language policies and action denials can convert reachable solutions into repeatable ones without weight changes.
Simon P. Villani arXiv: 2609.25053
Can one language model hand its live memory to another without the receiver rereading the context? We demonstrate useful persistent hybrid-state transfer across one architecture-matched Qwen3.5 4B-to-9B sibling pair. To our knowledge, this is the first demonstrated cross-model handoff of persistent recurrent…
Key insight: Architecture-matched hybrid models can hand off persistent recurrent state across sizes without target prefix replay.
Junlin Fang; Chong Zhang; Do Nguyen-Thanh et al. arXiv: 2609.25678
Large language model (LLM) agents must generalize from tools seen during training to unseen tools at deployment. A key challenge is tool trialing, i.e., excessive trials waste the interaction budget, whereas selective trials enable exploration of unfamiliar tools.
Key insight: Tool-call embeddings structured by function class can guide trialing toward unseen-but-similar tools instead of suppressing exploration.
Tiantong Wu; Wei Yang Bryan Lim arXiv: 2609.26532
LLM agents often use generative models for bounded decisions, raising the question of when these decisions can be handled more efficiently without reducing task success. We study REFLEX, an agent architecture that uses Jev as a fast, typed decision layer and calls a strong LLM when confidence is low, or generation is…
Key insight: A typed fast decision layer with LLM fallback can preserve task success while cutting expensive generative calls.
Yan Zhang; Pei Fu; Daiqing Wu et al. arXiv: 2609.25769
Graphical User Interface (GUI) Agents autonomously interact with software to fulfill user requests, where GUI navigation stands out as the most critical and challenging capability. Mastering this capability demands a complex synergy of step-wise decision-making, state-action alignment, and long-horizon planning.
Key insight: Masked trajectory prediction can unify GUI navigation step selection, state-action alignment, and long-horizon planning.
Haobo Zheng; Tan Tang; Yan Chen et al. arXiv: 2609.26780
Long-term conversational memory in multi-party settings requires more than retrieving relevant content from long-term conversations: it must distinguish who said what, whom each statement concerns, how individuals perceive one another, what information is shared by the group, and how states change over time. Recent…
Key insight: Multi-party conversational memory needs speaker-labeled verbatim tracks plus person and group state views.
Jianzhe Lin; Xiaolin Li; Fei Wang et al. arXiv: 2609.25337
Dialogue failures in language models are usually framed as memory failures: context too long, summaries lossy, a constraint forgotten. We argue this misses a deeper problem: in many conversations the model does not forget, it commits too early.
Key insight: Many dialogue failures are early commitment, not forgetting — later clarification is treated as extra context rather than a corrective signal.
Jinghan Xu; Longze Fan; Zeyuan Wang et al. arXiv: 2609.26076
Structured multi-agent workflows exchange intermediate messages whose content and form can reveal private state even when the final output is safe. We identify selection-channel leakage: after authorization fixes what may be released, a private-state-aware choice among semantically valid realizations creates an…
Key insight: After authorization fixes what may be released, selection among valid message forms can still leak private state.
Jinghan Xu; Longze Fan; Zeyuan Wang et al. arXiv: 2609.26072
Inter-agent communication is essential to multi-agent language-model systems, yet a single message may combine task-critical information with instructions not authorized by the original request. Prompt-based defenses leave enforcement to models exposed to adversarial messages, while indiscriminate message removal…
Key insight: Tainted inter-agent messages should be regenerated under executable semantic commitments in a clean room.
Agamdeep Singh; Srishti Gautam; Priyanshu Gupta et al. arXiv: 2609.26261
Search-based prompt optimizers improve prompts through iterative search: they propose edits, execute fresh rollouts, score the resulting trajectories, and retain only edits that improve a validation metric. We show that this optimization loop is unnecessary.
Key insight: Coding agents analyzing a static trajectory corpus can outperform search-based prompt optimizers that burn fresh rollouts.
Md Mostafizer Rahman; Md Faizul Ibne Amin; Md Shahajada Mia et al. arXiv: 2609.25537
Large language model (LLM) inference is constrained by the quadratic scaling of self-attention and the linear scaling of the KV cache, increasing latency, energy consumption, and GPU memory demand as context length scales. Existing soft-compression methods either lack query-guided memory selection at inference time,…
Key insight: Long context can be compressed into query-guided, answer-aligned memory embeddings for a frozen decoder.
Zhen Huang; Ruizhe Yao; Danyi Liu et al. arXiv: 2609.26300
Despite their strong performance, large language models (LLMs) are bottlenecked by KV cache memory traffic during long-context inference. Sparse attention is widely used to accelerate LLM inference by computing exact attention over a selected subset of tokens.
Key insight: KV block selection by downstream compensation error beats selection by attention mass alone for sparse long-context inference.
Lijuan Tang; Yuemeng Zheng arXiv: 2609.26693
A coding agent must emit a valid tool call--a parseable invocation of a tool in the provided schema--before the harness can execute its chosen action. We study how local serving stacks affect this protocol step and show that measured outcomes can depend on the serving layer rather than model behavior alone.
Key insight: Local tool-use scores can be confounded by the serving stack template and parsing layer rather than model behavior.
Atul Anand arXiv: 2609.25130
Memory systems for coding agents must decide, when a repository changes, which of their stored claims have become false. Content anchoring invalidates a claim whenever the artifact it came from changes, which fires constantly.
Key insight: Coding-agent memory invalidation should ask whether a stored claim still holds, not whether a diff appears to preserve behavior.
Ethan Torres; Eric Mills arXiv: 2609.25286
Structured-data agents in enterprise settings must reason over complex data environments whose relevant evidence is distributed across schemas, relationships, policies, and recurring business roles. Modern agentic systems often address this burden through reusable markdown-style memory or skill files that preserve…
Key insight: Enterprise data agents need learned latent identities and routing prototypes instead of stuffing schemas into markdown skills.
Dhruv Srikanth; Bingchen Zhao; Dixing Xu et al. arXiv: 2609.26457
AI agents are beginning to automate research and development across the AI stack, from improving training efficiency to optimizing inference. A natural next step is to improve the research efficiency of the agents themselves.
Key insight: Research agents that edit their own code against hidden evaluations demonstrate practical recursive self-improvement loops.
Wenbo Pan; Zhichao Liu; Shujie Liu et al. arXiv: 2609.25804
LLM agents increasingly work on long-horizon tasks, and the decisions they make along the way, such as which hypothesis to test or which implementation to build on, determine the outcome of the whole run. Making these decisions well is becoming a key capability for both engineering and research agents.
Key insight: Long-horizon agents need mid-trajectory taste scores on decision forks, not only end-state success.
Dahlia Shehata; Ming Li arXiv: 2609.25570
Large language models (LLMs) exhibit a parametric vulnerability to adversarial swarm consensus. To mitigate this sycophancy, we introduce Contrastive Epistemic Decoding (CED), a zero-shot inference intervention.
Key insight: Contrastive epistemic decoding can reduce swarm-consensus sycophancy without fine-tuning.
Haocheng Xia; Eugene Wu; Yongjoo Park arXiv: 2609.25396
Parallel coding agents can produce patches that work alone but fail when merged. This happens when one agent changes an interface or rule that another agent still relies on.
Key insight: Parallel coding agents can pass alone yet fail when merged unless coordination messages cover concurrent interface changes.
Guanqun Yang; Wei Yang; Xueqing Liu arXiv: 2609.26204
When a professional web developer's code fails a test, they do not simply re-read the stack trace. They open the application in a browser, click buttons, inspect computed styles, and run diagnostic commands to understand what went wrong.
Key insight: Executable diagnostic scripts against the live app give agentic web developers a stronger retry signal than screenshot judges.
Yiyao Zhang; Diksha Goel; Hussain Ahmad et al. arXiv: 2609.26135
Multi-agent reasoning systems in high-stakes domains must be both accurate and safe, yet agents often follow heterogeneous value priorities (e.g., rigor, conciseness, safety), causing conflicting recommendations. Existing methods do not jointly provide: (i) principled inference of each agent's implicit values from…
Key insight: Inferring per-agent values and composing assume-guarantee shields can reconcile heterogeneous multi-agent priorities at runtime.
Lujia Bao; Qian Chen; Luyao Cheng et al. arXiv: 2609.25176
Real-time voice assistants must reason over evolving requests, execute actions, and follow conversational rules. Qwen-Audio-3.1-Realtime brings these requirements together through Think, Act, and Speak and Coordinate.
Key insight: Realtime voice agents need coordinated Think, Act, and Speak loops that combine tool use with full-duplex conversation.
Wenhui Chen; Jianlin Chen; Ziyao Lin; Chi Man Vong arXiv: 2609.25052
An agent that writes its conclusions into a store it later retrieves from closes a loop usually reported as one-way contamination. Taking the loop to the infinite-tenure limit against an append-only store gives a different picture: because writing never deletes, the reachable state space has a hard upper edge at…
Key insight: An agent that writes conclusions into a store it later retrieves from has a measurable self-contamination primitive at infinite tenure.
Xiaoyu Yang; Jie Lu; Wei Duan; En Yu arXiv: 2609.26718
Long-context LLMs focus on retrieving distant evidence from extensive context, yet existing work has largely focused on overcoming distance alone. In this work, we identify the Proximity Trap, insufficient attention to distant evidence often arises less from distance itself than from cumulative competition with…
Key insight: Distant evidence often loses to irrelevant proximal background — a proximity trap, not pure distance failure.
Nathalie Baracaldo arXiv: 2609.26682
Generative AI (GenAI) applications have flourished enabling users to chat with large language models, and to create agents to act on their behalf for a variety of tasks. The pace of development of capabilities in this field is incredibly fast with security and safety taking a back seat.
Key insight: GenAI policy should put agent tool permission in the access-control column rather than relying on alignment prompts alone.
Ariel Flint; Luca Maria Aiello; Sara M. Constantino et al. arXiv: 2609.25194
As generative AI agents are deployed at scale, safety will depend not only on technical safeguards and individual model design, but also on collective equilibria that determine how agent populations process information, prioritize actions, and respond to uncertainty. Yet the same equilibria that enable agents to…
Key insight: Agent-population equilibria can be redirected by indirect tipping below classical critical-mass fractions.
Junyoung Jang; Gwanhyun Lee; Hwiwon Lee et al. arXiv: 2609.25591
Coding agents now find real vulnerabilities in production software. However, bug discovery results do not measure whether agents can construct exploit primitives.
Key insight: Coding-agent capability on real kernel exploit primitives remains far weaker than vulnerability discovery alone suggests.
Jennifer Williams; Dave Farris; Jeff Farris; Jiantao Jiao arXiv: 2609.26777
We introduce SWE-Serve, a benchmark for evaluating agents on production inference engineering tasks. Implementing an inference feature can require coordinating multiple changes across the serving stack, including model support, runtime execution, and public APIs.
Key insight: Production inference-serving engineering tasks expose a harder agentic stack than typical SWE-bench issues.
David Garg; Ritobrata Sarkar; Ehsan Azarnasab; Siddhartha Borah arXiv: 2609.25467
We often teach a colleague by showing the work and explaining the decisions as we go. How can we check what an agent understood from the same lesson? We introduce ShowTellArena, a benchmark protocol and public dataset for comprehension after narrated business demonstrations.
Key insight: Narrated business-workflow demonstrations need comprehension quizzes, not only click cloning, to grade agent understanding.
Jianzhe Lin; Xiaolin Li; Yunda Liu et al. arXiv: 2609.25284
A social agent's most basic decisions (should I react to this post? who should I reach out to?) are not purely content problems. The right action often hinges on the latent relationship between people -- tie strength, reciprocity, mutual connections -- rather than on which content is most salient.
Key insight: Social agents need explicit relational hypotheses (tie strength, reciprocity) rather than content-salient defaults.
YanZe Cao arXiv: 2609.25647
Predicting early outcomes based on trajectory can decrease the expenses associated with agent evaluation by terminating a run once the outcome becomes sufficiently predictable, assuming that the predictor's confidence is properly calibrated. Calibration is at risk when a predictor is applied to an agent on which it…
Key insight: Early-stop outcome predictors calibrated on one agent do not transfer reliably to another agent on the same benchmark.
Peiying Zhu; Sidi Chang arXiv: 2609.25806
Runtime traces can appear transparent, but a closed-loop policy determines which states are visited and which failures become visible. We study a simulated hotel-pricing agent mapping time, inventory, and market state to discrete price actions under varying demand regimes.
Key insight: Aggregate agent traces are diagnosable only when the policy actually visits the faulted state cells.
Huatai Zhu; Qiang Chen; Ziqian Kou et al. arXiv: 2609.26293
Learned world models are becoming essential to general-purpose agents: by predicting action consequences, they support planning and decision-making while reducing reliance on costly trial and error. This reliance creates a fundamental ambiguity: when a world-model-guided decision fails, the trajectory alone may not…
Key insight: World-model-guided actions should be admitted only when predicted advantage beats certified world-model error.
Shivam Gupta arXiv: 2609.26642
Successful agent execution need not identify which future product improvement its user would value. We present a decision-specific audit that maps a declared observation channel and product-value contrast to compatible intervals and witness populations.
Key insight: Successful agent execution does not identify which future product improvement the user would value.
Yifeng He; Jiachen Liu arXiv: 2609.25421
As autonomous AI agents take on every stage of scientific inquiry, research output is expanding far beyond human review capacity. Yet scientific communication still relies on natural-language prose: an informal medium prone to ambiguity, hidden assumptions, and untracked limitations that machines cannot reliably audit.
Key insight: Autonomous science needs a machine-checkable language for claims, evidence, and assumptions rather than prose alone.
Weihang Ding; Junfei Zhan arXiv: 2609.25237
Post-training is becoming a service (PTaaS): a customer hands an operator data and a goal, and a forward-deployed engineer (FDE) returns a fine-tuned, evaluated, and deployed model under a budget, a human-approval gate, and reproducibility requirements. Seating an LLM agent in the FDE seat raises a question existing…
Key insight: Post-training-as-a-service agents can train with green signals yet deliver models no better than base.
Kymberly Lasser-Chere; Tyler Akidau; Marc Millstone arXiv: 2609.26562
The vocabulary used to describe AI agents in governance contexts -- learning, memory, values, compliance, identity, trust -- is borrowed from psychological and organizational science, contributing to systematic failures in how organizations deploy, oversee, and hold agents accountable. This paper argues that the…
Key insight: Borrowing psychological vocabulary for agent governance systematically mis-calibrates oversight of stores, policies, and principals.
Grant Molnar arXiv: 2609.26419
Reliability theory gives a mature language for layered systems, but its formal tools are not yet standard in frontier AI control. We apply them to Google DeepMind's defenses against rogue deployment.
Key insight: Reliability theory reframes layered AI control defenses so rare-event suppression scales differently across failure domains.
Vasily Ilin arXiv: 2609.25199
Lean Pool is a repository of formalized mathematics. It is grown, maintained and optimized by AI agents.
Key insight: An AI-maintained Lean archive shows agents can grow and optimize persistent formal artifact stores.
Yuanteng Chen; Qiwei Lai; Chen Tianqi et al. arXiv: 2609.25809
Fine-grained mixture-of-experts (MoE) architectures have become a mainstream design for open-weight LLMs, with hundreds of experts and increasingly many selected per token. This shift makes dynamic expert pruning an attractive route to cheaper inference.
Key insight: Fine-grained MoE models often need only about two-thirds of the chosen experts at inference time.