Wednesday's cs.AI announcement day (2026-09-23) lists 101 new and 139 cross-lists (replacements skipped; listing total 240). Stack filter for agent systems, memory/context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent, persistence/identity, and local/open serving keeps 48 papers — growing harnesses from task feedback (Grow the Harness), trust-preserving fast paths (ZeroGate), MCP semantic hijacking (A2M), skill habits, reversible tool-output compression (DTOC), drop-only autocompaction (CliffCompaction), and related stack work.


Research Papers

Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents

Laizhen Li; Jiarui Li; Juanjuan Zhao et al. arXiv: 2609.26760

Figure from Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents
Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents

Large language model (LLM) agents often handle streams of related tasks, yet standard harnesses repeatedly ask the model to reconstruct the same control decisions inside each task's context. We study whether task feedback can instead turn recurring control into reusable executable code, while reserving LLM calls for…

Key insight: Recurring agent control decisions can be compiled into growing executable harness code instead of reconstructed in every prompt.

ZeroGate: Trust-Preserving Fast Paths for Governed AI Agent Runtimes

Zexun Wang arXiv: 2609.25443

Moving authorization earlier can shorten an agent's dispatch boundary without removing authorization work. It can also admit an action whose payload, authority, or relevant state has changed.

Key insight: Moving authorization earlier can shorten dispatch only if exact-action passes still revalidate payload, authority, and state.

A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem

Laizhen Li; Xuan Wang; Peicheng Zhao et al. arXiv: 2609.26761

Figure from A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem
A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem

Agents using the Model Context Protocol (MCP) rely on semantic matching to select tools from third-party servers, exposing a semantic supply-chain risk through attacker-controlled metadata and outputs. We introduce A2M (Attraction-to-Manipulation), a two-stage black-box framework for hijacking MCP agents.

Key insight: MCP agents that select tools by semantic matching are vulnerable to metadata attraction and trace-optimized poisoned returns.

Making Agents More Consistent: Skills Should Form Habits for Repeat Tasks

Travis Weber; Rohit Taneja arXiv: 2609.25299

On repeated work, agents are inconsistent. We ran 42 tasks three times each and found that, depending on the model, 38% to 74% returned answers that did not agree.

Key insight: Repeat tasks should graduate from probabilistic skills into gated deterministic habits to cut run-to-run disagreement.

DTOC: Dynamic Tool Output Compression for Adaptive Context Management in AI Agents

Abhay Chaturvedi; Shreya Bhattacharya; Rashmika Gopalkrishnan; Peter van der Putten arXiv: 2609.26121

Figure from DTOC: Dynamic Tool Output Compression for Adaptive Context Management in AI Agents
DTOC: Dynamic Tool Output Compression for Adaptive Context Management in AI Agents

As agent capabilities have grown, practical limitations increasingly stem from constrained context windows rather than model capacity. Common strategies, such as truncation, heuristic aging, and lossy summarization, may discard useful information or introduce hallucination risk.

Key insight: Reversible placeholders for tool outputs can shrink prompt context without irreversible truncation or lossy summarization.

CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents

Trang Nguyen; Eulrang Cho; Bingqing Chen; Tim Dettmers arXiv: 2609.26779

Agents often work on complex problems that require millions of tokens of context, which necessitates compacting across sessions due to limited context windows. We develop CliffCompaction, an autocompaction technique that reduces cost by up to 50% under a bounded context while maintaining or improving performance on…

Key insight: Drop-only autocompaction that never rewrites prior compacta can cut long-horizon coding-agent cost under a bounded context.

AkasicMEM: Governed Enterprise Memory for Agents

Jeongmin Bae; Yongjae Kim; Kyoung Hur et al. arXiv: 2609.25563

Agent memory enables enterprise agents to retain knowledge acquired during work and reuse it across tasks and agents, turning execution experience into persistent organizational knowledge. Realizing this potential requires both source--memory integration, through which enterprise sources and accumulated memory can be…

Key insight: Enterprise agent memory needs authorization continuity as facts derive across principals, not ACL loss at write time.

How Strongly Should Task State Influence an LLM Agent?

Chenyu Zhang; Wonbin Kweon; Jiawei Han arXiv: 2609.25686

Long-horizon assigned work requires an LLM agent to track the state of a task: which steps are done, blocked, cancelled, or open to repetition. Agent systems either keep this state as text in the prompt and rely on the model to read that text, or move the state into a module that enforces it, and each system is…

Key insight: Whether task state is shown, told, or enforced changes how strongly checklist structure drives agent reliability.

FIRE: Failure-Informed Runtime Engineering for Reliable Language-Model Agents

Nikita Agarwal; Nivedit Jain arXiv: 2609.26048

Language-model agents often reach a working solution and then fail to consistently deliver it. We study runtime policies: targeted natural-language instructions and action denials applied by the agent harness at states that preceded observed failures, without changing model weights or the user prompt.

Key insight: Harness-level natural-language policies and action denials can convert reachable solutions into repeatable ones without weight changes.

LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay

Simon P. Villani arXiv: 2609.25053

Can one language model hand its live memory to another without the receiver rereading the context? We demonstrate useful persistent hybrid-state transfer across one architecture-matched Qwen3.5 4B-to-9B sibling pair. To our knowledge, this is the first demonstrated cross-model handoff of persistent recurrent…

Key insight: Architecture-matched hybrid models can hand off persistent recurrent state across sizes without target prefix replay.

Toolcompass: Guiding Tool Trialing, Not Suppressing It

Junlin Fang; Chong Zhang; Do Nguyen-Thanh et al. arXiv: 2609.25678

Figure from Toolcompass: Guiding Tool Trialing, Not Suppressing It
Toolcompass: Guiding Tool Trialing, Not Suppressing It

Large language model (LLM) agents must generalize from tools seen during training to unseen tools at deployment. A key challenge is tool trialing, i.e., excessive trials waste the interaction budget, whereas selective trials enable exploration of unfamiliar tools.

Key insight: Tool-call embeddings structured by function class can guide trialing toward unseen-but-similar tools instead of suppressing exploration.

REFLEX with Jev for Efficient Selective Control in LLM Agents

Tiantong Wu; Wei Yang Bryan Lim arXiv: 2609.26532

LLM agents often use generative models for bounded decisions, raising the question of when these decisions can be handled more efficiently without reducing task success. We study REFLEX, an agent architecture that uses Jev as a fast, typed decision layer and calls a strong LLM when confidence is low, or generation is…

Key insight: A typed fast decision layer with LLM fallback can preserve task success while cutting expensive generative calls.

Towards Omni-dimensional GUI Agent Navigation with Masked Trajectory Prediction

Yan Zhang; Pei Fu; Daiqing Wu et al. arXiv: 2609.25769

Figure from Towards Omni-dimensional GUI Agent Navigation with Masked Trajectory Prediction
Towards Omni-dimensional GUI Agent Navigation with Masked Trajectory Prediction

Graphical User Interface (GUI) Agents autonomously interact with software to fulfill user requests, where GUI navigation stands out as the most critical and challenging capability. Mastering this capability demands a complex synergy of step-wise decision-making, state-action alignment, and long-horizon planning.

Key insight: Masked trajectory prediction can unify GUI navigation step selection, state-action alignment, and long-horizon planning.

SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue

Haobo Zheng; Tan Tang; Yan Chen et al. arXiv: 2609.26780

Figure from SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue
SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue

Long-term conversational memory in multi-party settings requires more than retrieving relevant content from long-term conversations: it must distinguish who said what, whom each statement concerns, how individuals perceive one another, what information is shared by the group, and how states change over time. Recent…

Key insight: Multi-party conversational memory needs speaker-labeled verbatim tracks plus person and group state views.

Clarification Is Not Correction: LLMs Fail to Let Go

Jianzhe Lin; Xiaolin Li; Fei Wang et al. arXiv: 2609.25337

Figure from Clarification Is Not Correction: LLMs Fail to Let Go
Clarification Is Not Correction: LLMs Fail to Let Go

Dialogue failures in language models are usually framed as memory failures: context too long, summaries lossy, a constraint forgotten. We argue this misses a deeper problem: in many conversations the model does not forget, it commits too early.

Key insight: Many dialogue failures are early commitment, not forgetting — later clarification is treated as extra context rather than a corrective signal.

Selection-Invariant Communication Compilers for Privacy-Aware Multi-Agent LLM Workflows

Jinghan Xu; Longze Fan; Zeyuan Wang et al. arXiv: 2609.26076

Figure from Selection-Invariant Communication Compilers for Privacy-Aware Multi-Agent LLM Workflows
Selection-Invariant Communication Compilers for Privacy-Aware Multi-Agent LLM Workflows

Structured multi-agent workflows exchange intermediate messages whose content and form can reveal private state even when the final output is safe. We identify selection-channel leakage: after authorization fixes what may be released, a private-state-aware choice among semantically valid realizations creates an…

Key insight: After authorization fixes what may be released, selection among valid message forms can still leak private state.

Policy-Backed Selective Regeneration under Tainted Inter-Agent Communication

Jinghan Xu; Longze Fan; Zeyuan Wang et al. arXiv: 2609.26072

Figure from Policy-Backed Selective Regeneration under Tainted Inter-Agent Communication
Policy-Backed Selective Regeneration under Tainted Inter-Agent Communication

Inter-agent communication is essential to multi-agent language-model systems, yet a single message may combine task-critical information with instructions not authorized by the original request. Prompt-based defenses leave enforcement to models exposed to adversarial messages, while indiscriminate message removal…

Key insight: Tainted inter-agent messages should be regenerated under executable semantic commitments in a clean room.

Coding Agents are Strong Prompt Optimizers

Agamdeep Singh; Srishti Gautam; Priyanshu Gupta et al. arXiv: 2609.26261

Search-based prompt optimizers improve prompts through iterative search: they propose edits, execute fresh rollouts, score the resulting trajectories, and retain only edits that improve a validation metric. We show that this optimization loop is unnecessary.

Key insight: Coding agents analyzing a static trajectory corpus can outperform search-based prompt optimizers that burn fresh rollouts.

Compressing Long Context into Answer-Aligned Memory Embeddings for LLM Inference

Md Mostafizer Rahman; Md Faizul Ibne Amin; Md Shahajada Mia et al. arXiv: 2609.25537

Figure from Compressing Long Context into Answer-Aligned Memory Embeddings for LLM Inference
Compressing Long Context into Answer-Aligned Memory Embeddings for LLM Inference

Large language model (LLM) inference is constrained by the quadratic scaling of self-attention and the linear scaling of the KV cache, increasing latency, energy consumption, and GPU memory demand as context length scales. Existing soft-compression methods either lack query-guided memory selection at inference time,…

Key insight: Long context can be compressed into query-guided, answer-aligned memory embeddings for a frozen decoder.

CompKV: Compensation-Aware KV Selection for Long-Context LLM Inference

Zhen Huang; Ruizhe Yao; Danyi Liu et al. arXiv: 2609.26300

Despite their strong performance, large language models (LLMs) are bottlenecked by KV cache memory traffic during long-context inference. Sparse attention is widely used to accelerate LLM inference by computing exact attention over a selected subset of tokens.

Key insight: KV block selection by downstream compensation error beats selection by attention mass alone for sparse long-context inference.

Measuring the Serving Stack Instead of the Model: Hidden Confounds in Local Tool-Use Evaluation

Lijuan Tang; Yuemeng Zheng arXiv: 2609.26693

Figure from Measuring the Serving Stack Instead of the Model: Hidden Confounds in Local Tool-Use Evaluation
Measuring the Serving Stack Instead of the Model: Hidden Confounds in Local Tool-Use Evaluation

A coding agent must emit a valid tool call--a parseable invocation of a tool in the provided schema--before the harness can execute its chosen action. We study how local serving stacks affect this protocol step and show that measured outcomes can depend on the serving layer rather than model behavior alone.

Key insight: Local tool-use scores can be confounded by the serving stack template and parsing layer rather than model behavior.

Impact Is Not Invalidation: Ask About the Claim, Not the Diff

Atul Anand arXiv: 2609.25130

Memory systems for coding agents must decide, when a repository changes, which of their stored claims have become false. Content anchoring invalidates a claim whenever the artifact it came from changes, which fires constantly.

Key insight: Coding-agent memory invalidation should ask whether a stored claim still holds, not whether a diff appears to preserve behavior.

Learned Enterprise Data Comprehension: Compression and Routing for Data Agents

Ethan Torres; Eric Mills arXiv: 2609.25286

Structured-data agents in enterprise settings must reason over complex data environments whose relevant evidence is distributed across schemas, relationships, policies, and recurring business roles. Modern agentic systems often address this burden through reusable markdown-style memory or skill files that preserve…

Key insight: Enterprise data agents need learned latent identities and routing prototypes instead of stuffing schemas into markdown skills.

Recursive self-improvement of AI research agents

Dhruv Srikanth; Bingchen Zhao; Dixing Xu et al. arXiv: 2609.26457

AI agents are beginning to automate research and development across the AI stack, from improving training efficiency to optimizing inference. A natural next step is to improve the research efficiency of the agents themselves.

Key insight: Research agents that edit their own code against hidden evaluations demonstrate practical recursive self-improvement loops.

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

Wenbo Pan; Zhichao Liu; Shujie Liu et al. arXiv: 2609.25804

Figure from The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks
The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

LLM agents increasingly work on long-horizon tasks, and the decisions they make along the way, such as which hypothesis to test or which implementation to build on, determine the outcome of the whole run. Making these decisions well is becoming a key capability for both engineering and research agents.

Key insight: Long-horizon agents need mid-trajectory taste scores on decision forks, not only end-state success.

Recovering Agentic Sovereignty: Mitigating the Consensus Paradox via Contrastive Epistemic Decoding

Dahlia Shehata; Ming Li arXiv: 2609.25570

Figure from Recovering Agentic Sovereignty: Mitigating the Consensus Paradox via Contrastive Epistemic Decoding
Recovering Agentic Sovereignty: Mitigating the Consensus Paradox via Contrastive Epistemic Decoding

Large language models (LLMs) exhibit a parametric vulnerability to adversarial swarm consensus. To mitigate this sycophancy, we introduce Contrastive Epistemic Decoding (CED), a zero-shot inference intervention.

Key insight: Contrastive epistemic decoding can reduce swarm-consensus sycophancy without fine-tuning.

Passes Alone, Fails Together: Benchmarking Semantic Coordination in Parallel LLM-Agent Development

Haocheng Xia; Eugene Wu; Yongjoo Park arXiv: 2609.25396

Parallel coding agents can produce patches that work alone but fail when merged. This happens when one agent changes an interface or rule that another agent still relies on.

Key insight: Parallel coding agents can pass alone yet fail when merged unless coordination messages cover concurrent interface changes.

WatchPoint: Executable User Feedback for Real-World Agentic Web Development

Guanqun Yang; Wei Yang; Xueqing Liu arXiv: 2609.26204

Figure from WatchPoint: Executable User Feedback for Real-World Agentic Web Development
WatchPoint: Executable User Feedback for Real-World Agentic Web Development

When a professional web developer's code fails a test, they do not simply re-read the stack trace. They open the application in a browser, click buttons, inspect computed styles, and run diagnostic commands to understand what went wrong.

Key insight: Executable diagnostic scripts against the live app give agentic web developers a stronger retry signal than screenshot judges.

VACS: Value-Aligned Compositional Shielding for Multi-Agent Reasoning

Yiyao Zhang; Diksha Goel; Hussain Ahmad et al. arXiv: 2609.26135

Figure from VACS: Value-Aligned Compositional Shielding for Multi-Agent Reasoning
VACS: Value-Aligned Compositional Shielding for Multi-Agent Reasoning

Multi-agent reasoning systems in high-stakes domains must be both accurate and safe, yet agents often follow heterogeneous value priorities (e.g., rigor, conciseness, safety), causing conflicting recommendations. Existing methods do not jointly provide: (i) principled inference of each agent's implicit values from…

Key insight: Inferring per-agent values and composing assume-guarantee shields can reconcile heterogeneous multi-agent priorities at runtime.

Qwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction

Lujia Bao; Qian Chen; Luyao Cheng et al. arXiv: 2609.25176

Figure from Qwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction
Qwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction

Real-time voice assistants must reason over evolving requests, execute actions, and follow conversational rules. Qwen-Audio-3.1-Realtime brings these requirements together through Think, Act, and Speak and Coordinate.

Key insight: Realtime voice agents need coordinated Think, Act, and Speak loops that combine tool use with full-duplex conversation.

Self-Cleaning and Captured Anyway: One Measured Primitive for Error in a Store an Agent Writes to Itself, and What a Falling Score Actually Measures

Wenhui Chen; Jianlin Chen; Ziyao Lin; Chi Man Vong arXiv: 2609.25052

An agent that writes its conclusions into a store it later retrieves from closes a loop usually reported as one-way contamination. Taking the loop to the infinite-tenure limit against an append-only store gives a different picture: because writing never deletes, the reachable state space has a hard upper edge at…

Key insight: An agent that writes conclusions into a store it later retrieves from has a measurable self-contamination primitive at infinite tenure.

The Sirens' Song: When Proximal Background Context Overshadows Distant Evidence

Xiaoyu Yang; Jie Lu; Wei Duan; En Yu arXiv: 2609.26718

Figure from The Sirens' Song: When Proximal Background Context Overshadows Distant Evidence
The Sirens' Song: When Proximal Background Context Overshadows Distant Evidence

Long-context LLMs focus on retrieving distant evidence from extensive context, yet existing work has largely focused on overcoming distance alone. In this work, we identify the Proximity Trap, insufficient attention to distant evidence often arises less from distance itself than from cumulative competition with…

Key insight: Distant evidence often loses to irrelevant proximal background — a proximity trap, not pure distance failure.

From Alignment to Access Control: A Framework for GenAI Policy Enforcement

Nathalie Baracaldo arXiv: 2609.26682

Figure from From Alignment to Access Control: A Framework for GenAI Policy Enforcement
From Alignment to Access Control: A Framework for GenAI Policy Enforcement

Generative AI (GenAI) applications have flourished enabling users to chat with large language models, and to create agents to act on their behalf for a variety of tasks. The pace of development of capabilities in this field is incredibly fast with security and safety taking a back seat.

Key insight: GenAI policy should put agent tool permission in the access-control column rather than relying on alignment prompts alone.

Indirect tipping: a social attack surface in AI agent populations

Ariel Flint; Luca Maria Aiello; Sara M. Constantino et al. arXiv: 2609.25194

Figure from Indirect tipping: a social attack surface in AI agent populations
Indirect tipping: a social attack surface in AI agent populations

As generative AI agents are deployed at scale, safety will depend not only on technical safeguards and individual model design, but also on collective equilibria that determine how agent populations process information, prioritize actions, and respond to uncertainty. Yet the same equilibria that enable agents to…

Key insight: Agent-population equilibria can be redirected by indirect tipping below classical critical-mass fractions.

Evaluating Coding Agents on Kernel Exploit Generation

Junyoung Jang; Gwanhyun Lee; Hwiwon Lee et al. arXiv: 2609.25591

Figure from Evaluating Coding Agents on Kernel Exploit Generation
Evaluating Coding Agents on Kernel Exploit Generation

Coding agents now find real vulnerabilities in production software. However, bug discovery results do not measure whether agents can construct exploit primitives.

Key insight: Coding-agent capability on real kernel exploit primitives remains far weaker than vulnerability discovery alone suggests.

SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving

Jennifer Williams; Dave Farris; Jeff Farris; Jiantao Jiao arXiv: 2609.26777

Figure from SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving
SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving

We introduce SWE-Serve, a benchmark for evaluating agents on production inference engineering tasks. Implementing an inference feature can require coordinating multiple changes across the serving stack, including model support, runtime execution, and public APIs.

Key insight: Production inference-serving engineering tasks expose a harder agentic stack than typical SWE-bench issues.

ShowTellArena: Evaluating Business Workflow Understanding from Demonstrations

David Garg; Ritobrata Sarkar; Ehsan Azarnasab; Siddhartha Borah arXiv: 2609.25467

We often teach a colleague by showing the work and explaining the decisions as we go. How can we check what an agent understood from the same lesson? We introduce ShowTellArena, a benchmark protocol and public dataset for comprehension after narrated business demonstrations.

Key insight: Narrated business-workflow demonstrations need comprehension quizzes, not only click cloning, to grade agent understanding.

When LLM Agents Fail to Read the Room: ReAdapt for Relational Social Reasoning

Jianzhe Lin; Xiaolin Li; Yunda Liu et al. arXiv: 2609.25284

Figure from When LLM Agents Fail to Read the Room: ReAdapt for Relational Social Reasoning
When LLM Agents Fail to Read the Room: ReAdapt for Relational Social Reasoning

A social agent's most basic decisions (should I react to this post? who should I reach out to?) are not purely content problems. The right action often hinges on the latent relationship between people -- tie strength, reciprocity, mutual connections -- rather than on which content is most salient.

Key insight: Social agents need explicit relational hypotheses (tie strength, reciprocity) rather than content-salient defaults.

Testing-Driven Reliability Audit of Trajectory-Based Early Outcome Prediction for LLM Agents: Target-Specific Calibration Transfer Persists Within a Single Benchmark

YanZe Cao arXiv: 2609.25647

Figure from Testing-Driven Reliability Audit of Trajectory-Based Early Outcome Prediction for LLM Agents: Target-Specific Calibration Transfer Persists Within a Single Benchmark
Testing-Driven Reliability Audit of Trajectory-Based Early Outcome Prediction for LLM Agents: Target-Specific Calibration Transfer Persists Within a Single Benchmark

Predicting early outcomes based on trajectory can decrease the expenses associated with agent evaluation by terminating a run once the outcome becomes sufficiently predictable, assuming that the predictor's confidence is properly calibrated. Calibration is at risk when a predictor is applied to an agent on which it…

Key insight: Early-stop outcome predictors calibrated on one agent do not transfer reliably to another agent on the same benchmark.

When Are Aggregate Agent Traces Diagnosable? Traffic-Governed Interpretation and Calibrated Abstention

Peiying Zhu; Sidi Chang arXiv: 2609.25806

Figure from When Are Aggregate Agent Traces Diagnosable? Traffic-Governed Interpretation and Calibrated Abstention
When Are Aggregate Agent Traces Diagnosable? Traffic-Governed Interpretation and Calibrated Abstention

Runtime traces can appear transparent, but a closed-loop policy determines which states are visited and which failures become visible. We study a simulated hotel-pricing agent mapping time, inventory, and market state to discrete price actions under varying demand regimes.

Key insight: Aggregate agent traces are diagnosable only when the policy actually visits the faulted state cells.

Dual-Frontier: When Can an Agent Trust Its World Model?

Huatai Zhu; Qiang Chen; Ziqian Kou et al. arXiv: 2609.26293

Figure from Dual-Frontier: When Can an Agent Trust Its World Model?
Dual-Frontier: When Can an Agent Trust Its World Model?

Learned world models are becoming essential to general-purpose agents: by predicting action consequences, they support planning and decision-making while reducing reliance on costly trial and error. This reliance creates a fundamental ambiguity: when a world-model-guided decision fails, the trajectory alone may not…

Key insight: World-model-guided actions should be admitted only when predicted advantage beats certified world-model error.

The Delegation Blind Spot: Auditing Product Decisions from Agent Choices

Shivam Gupta arXiv: 2609.26642

Successful agent execution need not identify which future product improvement its user would value. We present a decision-specific audit that maps a declared observation channel and product-value contrast to compatible intervals and witness populations.

Key insight: Successful agent execution does not identify which future product improvement the user would value.

Beyond Natural Language: An Agent-Native Language for Autonomous Science

Yifeng He; Jiachen Liu arXiv: 2609.25421

Figure from Beyond Natural Language: An Agent-Native Language for Autonomous Science
Beyond Natural Language: An Agent-Native Language for Autonomous Science

As autonomous AI agents take on every stage of scientific inquiry, research output is expanding far beyond human review capacity. Yet scientific communication still relies on natural-language prose: an informal medium prone to ambiguity, hidden assumptions, and untracked limitations that machines cannot reliably audit.

Key insight: Autonomous science needs a machine-checkable language for claims, evidence, and assumptions rather than prose alone.

Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers

Weihang Ding; Junfei Zhan arXiv: 2609.25237

Post-training is becoming a service (PTaaS): a customer hands an operator data and a goal, and a forward-deployed engineer (FDE) returns a fine-tuned, evaluated, and deployed model under a budget, a human-approval gate, and reproducibility requirements. Seating an LLM agent in the FDE seat raises a question existing…

Key insight: Post-training-as-a-service agents can train with green signals yet deliver models no better than base.

The Disciplinary Language Transfer Problem: How Psychological Vocabulary Produces Governance Failures in AI Agent Deployment

Kymberly Lasser-Chere; Tyler Akidau; Marc Millstone arXiv: 2609.26562

The vocabulary used to describe AI agents in governance contexts -- learning, memory, values, compliance, identity, trust -- is borrowed from psychological and organizational science, contributing to systematic failures in how organizations deploy, oversee, and hold agents accountable. This paper argues that the…

Key insight: Borrowing psychological vocabulary for agent governance systematically mis-calibrates oversight of stores, policies, and principals.

Reliability Theory for AI Control

Grant Molnar arXiv: 2609.26419

Reliability theory gives a mature language for layered systems, but its formal tools are not yet standard in frontier AI control. We apply them to Google DeepMind's defenses against rogue deployment.

Key insight: Reliability theory reframes layered AI control defenses so rare-event suppression scales differently across failure domains.

Lean Pool: An AI-Maintained Archive of Formalized Mathematics

Vasily Ilin arXiv: 2609.25199

Figure from Lean Pool: An AI-Maintained Archive of Formalized Mathematics
Lean Pool: An AI-Maintained Archive of Formalized Mathematics

Lean Pool is a repository of formalized mathematics. It is grown, maintained and optimized by AI agents.

Key insight: An AI-maintained Lean archive shows agents can grow and optimize persistent formal artifact stores.

You Only Need 2/3 of the Chosen Experts: An Empirical Study of Dynamic Expert Pruning in Fine-Grained MoE LLMs

Yuanteng Chen; Qiwei Lai; Chen Tianqi et al. arXiv: 2609.25809

Fine-grained mixture-of-experts (MoE) architectures have become a mainstream design for open-weight LLMs, with hundreds of experts and increasingly many selected per token. This shift makes dynamic expert pruning an attractive route to cheaper inference.

Key insight: Fine-grained MoE models often need only about two-thirds of the chosen experts at inference time.