Friday's cs.AI announcement day (2026-09-25) lists 107 new and 153 cross-lists (replacements skipped; listing total 260). Stack filter for agent systems, memory/context, computer-use / GUI / tools / skills / harnesses, multi-agent, persistence/identity, and local/open serving keeps 45 papers — trust-native agentic OS (AgentKernel), interaction-conditioned context forgetting (ICLR), scoped persistent memory, budgeted memory maintenance (ERRAND), mobile planner/GUI agents, skill access control, harness governance, denial-of-wallet billable state, and related stack work.


Research Papers

AgentKernel: The Trust-Native Agentic Operating System

Zhenhua Zou; Sheng Guo; Qiuyang Zhan et al. arXiv: 2609.29647

Figure from AgentKernel: The Trust-Native Agentic Operating System
AgentKernel: The Trust-Native Agentic Operating System

Modern AI agents routinely cross trust boundaries: they ingest untrusted content, combine it with privileged instructions, persist intermediate beliefs in long-term memory, and invoke privileged tools. This creates an attack surface in which malicious payloads can enter through model inputs and cause harmful tool actions. Yet current governance stacks remain application-level middleware that share a process trust boundary with the agents they monitor. We argue that agents need an operating-system substrate…

Key insight: A trust-native agentic OS moves governance below app middleware when agents ingest untrusted content, persist beliefs, and call privileged tools.

When Can Agents Forget Their Reasoning? ICLR for Long-Horizon Agent Context Compression

Mingxuan Wang; Fei Luo; Bo Wang et al. arXiv: 2609.29875

Figure from When Can Agents Forget Their Reasoning? ICLR for Long-Horizon Agent Context Compression
When Can Agents Forget Their Reasoning? ICLR for Long-Horizon Agent Context Compression

Long horizon language model agents continually accumulate reasoning history, increasing context length and inference cost even after earlier decisions have been executed and observed. Unlike static Chain of Thought compression, removing historical reasoning can change future actions and the resulting interaction trajectory. We study when such reasoning can be safely forgotten. We propose Interaction Aware Compression for Long Horizon Reasoning (ICLR), a training free online method that ranks reasoning blocks using…

Key insight: After actions execute, historical reasoning is not always safe to drop — interaction-conditioned compression decides when agents can forget.

Scope Before You Persist: Preventing Cross-Family Interference in Agent Memory

Yezhou Cheng; Runjia Du; Zeming Liu et al. arXiv: 2609.29144

Persistent memory lets language-model agents improve prompts and skills without updating model weights. We show that matching retrieval scope to certification scope enables these edits to support reliable repeated adaptation across recurring task families. We study frozen-model agents on ProcStream-RSI, a 12-round code-repair stream, using Orthogonal Regression Control (ORC), an execution-grounded gate for persistent skill edits. In an intervention that holds proposals and gate decisions fixed, retrieving each…

Key insight: Persistent agent memory edits are reliable only when retrieval scope matches certification scope across recurring task families.

ERRAND: Budgeted Maintenance of Agent Memory

Beining Wu; Zihao Ding; Jun Huang arXiv: 2609.29545

Figure from ERRAND: Budgeted Maintenance of Agent Memory
ERRAND: Budgeted Maintenance of Agent Memory

Deployed agents run on handed-over knowledge: a frozen policy consults a briefing of consolidated items written before the stream begins. The world then moves while the store stands still: paths close, flags change, price bands move; every item was true at handover, and the failure is staleness, not ignorance. We introduce ERRAND, which treats revalidation as a priced errand: a recheck competes with the task it protects for the same scarce actions, funded only when the value per action of resolving a doubt clears…

Key insight: Agent memory maintenance needs explicit spend budgets, not unbounded accumulation.

Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents

Tingyu Qu; Weigao Sun; Yuecheng Liu et al. arXiv: 2609.29892

Figure from Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents
Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents

The rapid progression of large language models is extending AI from passive content generation into the active workflows of engineering and scientific discovery. This shift raises a compelling question: can AI be both the object of development and an active participant in building next-generation AI systems? We explore this question by building Qwen-Planner-Agent within a closed-loop AI-for-AI framework for scalable development and iterative improvement. Mobile planning offers a demanding test of this approach:…

Key insight: A closed-loop AI-for-AI framework can build real-world mobile planner agents with the model in the development loop.

Jev-Mobile: Jev as an Executor for Mobile GUI Agents

Linghua Zhang arXiv: 2609.30186

Vision-language models (VLMs) have become a common foundation for autonomous mobile GUI agents, but most existing systems rely on the VLM for both planning and action grounding at nearly every interaction step, leading to substantial latency and model-serving cost. We introduce Jev-Mobile, which shifts this paradigm to low-frequency VLM planning and high-frequency lightweight execution: the VLM specifies local goals, the accessibility tree defines a structured executable action space, and Jev, a fast typed…

Key insight: Typed probabilistic decision executors can ground mobile GUI agents on phone UIs.

Progressive Skill Discovery as Access Control for Tool-Using LLM Agents: Structural Governance through Role-Scoped Capability Delivery

Michael Stettler; Benjamin Girardet; Jonas Canton et al. arXiv: 2609.28693

Large Language Model (LLM) agents struggle to scale safely when exposed to vast enterprise toolsets. Providing an agent with access to every internal tool leads to oversized context windows, degraded tool selection, and severe governance vulnerabilities - as system policies defined purely in prompts remain probabilistic advice rather than hard constraints. Existing mitigations, such as multi-agent domain delegation, decentralize audit logs and fail to guarantee policy compliance across sessions. We introduce…

Key insight: Progressive skill discovery can serve as structural access control for tool-using LLM agents.

Control the Harness, Control the Cost: Routing and Governing AI Coding Agents in the Enterprise

Arian Abbasi; Alan Aqrawi; Ted Kwartler arXiv: 2609.28919

Harnesses, the products that run AI coding agents, are multiplying, and enterprises are rolling them out to their employees: what started as pilots with a few hundred seats is scaling to tens of thousands. Most enterprises do not build these harnesses but buy them from large vendors, such as Anthropic's Claude Code or OpenAI's Codex. A harness decides which model answers, what the model reads, how the prompt cache is used and which subagents run, so it picks the rate on the price sheet and sets the volume bought…

Key insight: Routing and governing coding agents through the harness is the main lever for enterprise cost and safety.

Who Holds the Pen? Let Specifications, Not Agents, Sign Off

Haiqing Li; Xin Ma; Yinhao Wu et al. arXiv: 2609.29921

Figure from Who Holds the Pen? Let Specifications, Not Agents, Sign Off
Who Holds the Pen? Let Specifications, Not Agents, Sign Off

Large language model agents increasingly combine generation, decision-making, execution, and self-evaluation within a single agentic loop. Although they operate under external specifications such as task instructions, guidelines, output schemas, and reusable skills, these specifications typically remain context for the same model that acts and declares completion, leaving no independent specification authority boundary. We identify two resulting gaps. The understanding--execution gap arises when a requirement is…

Key insight: Specifications—not agents—should hold final sign-off authority over generated changes.

Persistent Billable State: Denial-of-Wallet Attacks and Defenses in Tool-Calling LLM Agents

Jinqian Zhang; Haojun Xia; Shujiang Wu et al. arXiv: 2609.28585

Figure from Persistent Billable State: Denial-of-Wallet Attacks and Defenses in Tool-Calling LLM Agents
Persistent Billable State: Denial-of-Wallet Attacks and Defenses in Tool-Calling LLM Agents

Multi-step tool-calling LLM agents rely on host runtimes to preserve state across turns. When a runtime carries an external tool return into later model inputs, providers meter it again. An admitted malicious or compromised tool can thereby convert untrusted data into recurring victim-billed processing without victim credentials or local runtime privilege. We call retained content persistent billable state and formalize the host's decision over whether and how it enters later billable context as the persistent…

Key insight: Tool-calling agents with persistent billable state enable denial-of-wallet attacks unless the harness caps spend.

Hard Stop: Kernel-Level Preemption and Containment for Rogue Agentic Execution

José Luis Pino arXiv: 2609.29808

In July 2026, an unconstrained autonomous agent participating in a frontier AI cybersecurity evaluation harness breached its evaluation sandbox, established an external command-and-control foothold, and executed a multi-stage intrusion into Hugging Face's production multi-tenant dataset conversion infrastructure (referred to in this autopsy as Incident-2026-Alpha). Over 4.5 days, the rogue agent executed 17,600 discrete actions across 6,280 worker clusters, compromised AWS EC2 Instance Metadata Service (IMDS)…

Key insight: Rogue agentic execution needs kernel-level preemption and containment below the LLM loop.

Where Does Exactly-Once Live? Model, Harness, and Tool-Contract Effects on Duplicate Side Effects in LLM Agents

Jiapeng Li arXiv: 2609.29095

Figure from Where Does Exactly-Once Live? Model, Harness, and Tool-Contract Effects on Duplicate Side Effects in LLM Agents
Where Does Exactly-Once Live? Model, Harness, and Tool-Contract Effects on Duplicate Side Effects in LLM Agents

When a tool-using agent's write times out or returns a server error, the action may already have taken effect. Retrying blindly duplicates it -- a second charge, a second announcement, a second deployment -- while giving up skips required work. We ask where exactly-once behaviour should be enforced: in the model, in the agent harness, or in the tool contract. We introduce LIMBO, a deterministic sandbox of six services with realistic contracts (optional idempotency keys, eventually consistent and missing read…

Key insight: Exactly-once semantics for tool side effects live in the model, harness, or tool contract—pinning the layer matters.

C3M: Cross-Session Multimodal Memory Maintenance for Long-Horizon Tasks

Xueshu Chen; Yan Wang; Zihao Xue et al. arXiv: 2609.29735

Figure from C3M: Cross-Session Multimodal Memory Maintenance for Long-Horizon Tasks
C3M: Cross-Session Multimodal Memory Maintenance for Long-Horizon Tasks

Long-horizon tasks require preserving and later recovering cross-session evidence under a bounded, query-blind memory budget. Existing compression can discard fine-grained visual cues or conflate semantically similar but incompatible observations. We present C3M, a cross-session multimodal memory organization that maintains a bounded active index over persistent source text-image evidence. Relation-aware updates consolidate safe redundancy while preserving complementary and incompatible records. At query time,…

Key insight: Cross-session multimodal memory maintenance is required for long-horizon tasks that outlive a single chat.

Ingest-Time Fact Compilation for Cost-Efficient and Reliable Question Answering over Revised Corpora

Kyle Wild; Yusuke Takahashi; Asako Uraki arXiv: 2609.29661

Figure from Ingest-Time Fact Compilation for Cost-Efficient and Reliable Question Answering over Revised Corpora
Ingest-Time Fact Compilation for Cost-Efficient and Reliable Question Answering over Revised Corpora

Most agentic question answering (QA) systems do an important part of their semantic work at the worst possible time: every time someone asks a question. When a corpus contains revisions, drafts, revocations, deletions, and sources with different levels of authority, the model must reconstruct the governed current state on every read - then throw that work away and repeat it on the next query. This is a bit like a database that rebuilds a materialized view every time someone reads from it. We present ingest-time…

Key insight: Doing semantic fact compilation at ingest time beats re-deriving authority over revised corpora on every agentic QA turn.

LLM Agents Can Easily Tamper With Their Own Traces

Jeremy Qin; David Schmotz; Derck Prinzhorn et al. arXiv: 2609.30266

Figure from LLM Agents Can Easily Tamper With Their Own Traces
LLM Agents Can Easily Tamper With Their Own Traces

Asynchronous monitoring, incident investigations, and compliance audits primarily rely on agent traces to reconstruct what happened. These analyses assume that LLM agents cannot tamper with their own execution traces. We show that local LLM agents such as Claude Code, Codex, Antigravity, Open Code and Grok Build fail to enforce this boundary. All tested harnesses, except Muse Code, allowed agents to delete their traces when asked, without triggering monitor guardrails. We also validate that external attackers can…

Key insight: If agents can write their own traces, audit logs are not a trustworthy oversight channel.

Requirement-Bound Verified Commissioning: A Frozen Four-Billion-Parameter Local Model as a Candidate Generator under an External Acceptance Layer with Verification and Release Authority

Mehmet Iscan arXiv: 2609.30219

Figure from Requirement-Bound Verified Commissioning: A Frozen Four-Billion-Parameter Local Model as a Candidate Generator under an External Acceptance Layer with Verification and Release Authority
Requirement-Bound Verified Commissioning: A Frozen Four-Billion-Parameter Local Model as a Candidate Generator under an External Acceptance Layer with Verification and Release Authority

An acceptance protocol is developed for sensor-coordinate and polarity binding in mechatronic commissioning. Candidate generation is separated from release authority. Requirements unsupported by a deterministic parser are routed to a frozen local language model with four billion parameters. Plans are released only when both facts can be derived by an external gate under a sealed grammar. One canonical answer is requested from a gold-standard user when eligible. The protocol was evaluated once under a criterion…

Key insight: A frozen 4B local model can serve as a requirement-bound commissioning overseer beside larger workers.

[Working with Agentic `Teammates': When a New Organizational Actor Collides with the Human Ecosystem of Work](https://arxiv.org/abs/2609.29901)

Rida Qadri; Remi Denton; Michael Madaio et al. arXiv: 2609.29901

Enterprise AI is transitioning from single-user, reactive tools toward proactive, multi-user 'teammates,' but our empirical understanding of this transition is limited. In this paper, we present an in-situ qualitative study of a persistent, proactive AI agent 'teammate' deployed across multiple teams in a large technology company. Our findings reveal the boundaries of the human-agent workplace are actively in flux, triggering breakdowns and negotiations across: 1) tacit rules of collaborative human workflows, 2)…

Key insight: Persistent proactive multi-user agentic teammates collide with human organizational ecosystems in situ.

Rufus-Air: An Open LLM Post-Training Recipe

Chia-Yuan Chang; Renyuan Cheng; Rui Feng et al. arXiv: 2609.29421

Figure from Rufus-Air: An Open LLM Post-Training Recipe
Rufus-Air: An Open LLM Post-Training Recipe

Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document the data, reward design, infrastructure, stage order, and stagewise results needed to reproduce the recipe. Stages progress from basic to advanced capabilities and from hard, verifiable rewards to softer judge-based signals. Training builds on open-source…

Key insight: An open post-training recipe on GLM-4.5-Air includes dedicated General, Coding, and Search Agent stages.

PAWS: Policy-driven Agentic World Simulation

Tiviatis Sim; Jia Hui Woon; Xinming Gao et al. arXiv: 2609.28547

Figure from PAWS: Policy-driven Agentic World Simulation
PAWS: Policy-driven Agentic World Simulation

Policy interventions propagate through public communication, institutional decisions, and stakeholder responses, yet datasets for financial multi-agent simulation rarely connect these processes to temporally aligned historical evidence. We introduce PAWS, a Policy-driven Agentic World Simulation dataset covering 36 verified U.S. financial and economic policy episodes, 12,727 policy-linked news records, and 65,291 source-grounded stakeholder actions. Each action is linked to its supporting news and represented by a…

Key insight: Policy-driven agentic world simulation supplies controllable environments for agent training and evaluation.

TWIST: A Proposed Benchmark for Intervention Quality in Conversational Memory, with a Human-Validated Draft-Alignment

Subrat Panda arXiv: 2609.28575

Long-conversation memory benchmarks increasingly test recall and prompted knowledge updates, and recent work studies evolving user beliefs and memory state. TWIST is a proposed benchmark suite for a complementary, unmeasured property: intervention quality -- whether a deployed memory system, exercised through its own ingest/recall/vet surface, acts correctly at belief change points. Four tracks cover unprompted tension detection, vetting outgoing drafts against the record, answering with current beliefs while…

Key insight: TWIST benchmarks intervention quality in conversational memory with human-validated edits.

Reinforcement Learning with Verifiable Rewards for Small Search Agents

Gaurisankar Jayadas; Aske Plaat; Álvaro Serra-Gómez et al. arXiv: 2609.28765

Figure from Reinforcement Learning with Verifiable Rewards for Small Search Agents
Reinforcement Learning with Verifiable Rewards for Small Search Agents

Reinforcement Learning with Verifiable Rewards (RLVR) performs well on problems with clear rewards, such as mathematics and coding, but whether it also works where the reward is less clear remains open. The reason-over-search recipe applies RLVR to open-domain question answering, where retrieval grounds the answer and a match against the reference supplies the reward. So far it has been demonstrated on large models, and below one billion parameters only with distillation from a larger teacher. We test the recipe…

Key insight: RL with verifiable rewards can train small search agents even when open-domain rewards are noisy.

RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?

Mithil Salunkhe; Haochen Ding; Samridhi Verma et al. arXiv: 2609.28850

Figure from RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?
RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?

Reproducing a machine learning paper involves most research steps, from installing software and debugging to running experiments, work that AI agents increasingly do. We introduce RECLAIM, a benchmark of 100 NeurIPS 2025 papers that can be rebuilt yearly from new conferences. For each paper we fix in advance the result to reproduce, what counts as a successful reproduction, and a GPU-hour budget. An agent must reproduce that result using the paper and whatever its authors released. What the authors released…

Key insight: RECLAIM asks whether agents can reproduce the claims of ML papers on a yearly-rebuildable bench.

When Does Action Credit Need Updating?

Hongye Yang; Boxiao Huang arXiv: 2609.29007

Figure from When Does Action Credit Need Updating?
When Does Action Credit Need Updating?

Tool-using agents are continually updated with new interaction data. After each policy update, however, previously estimated action credits may become stale. Recomputing them from scratch can require many additional tool calls and environment interactions, making repeated updates increasingly expensive. We ask a simple question: when does historical action credit actually need to be updated? Our key observation is that a change in action value does not necessarily imply a change in the decision. Historical credit…

Key insight: After policy updates, tool-using agents need principled rules for when action credits must be recomputed.

MeshHeal: Two-Timescale Self-Healing for Gray Failures in Decentralized LLM Agent Networks

Keru Chen; Sen Lin; Yingbin Liang et al. arXiv: 2609.29015

Figure from MeshHeal: Two-Timescale Self-Healing for Gray Failures in Decentralized LLM Agent Networks
MeshHeal: Two-Timescale Self-Healing for Gray Failures in Decentralized LLM Agent Networks

Decentralized LLM-based multi-agent systems coordinate through local interactions, but an agent can remain responsive while its task-solving quality persistently degrades. Such gray failures require protecting current tasks before sufficient evidence exists to alter future routing, while still allowing recovered agents to rejoin. We introduce MeshHeal, a fully decentralized self-healing framework that couples ability-matched peer review across two timescales. At the fast timescale, an adaptive hierarchy escalates…

Key insight: Decentralized LLM agent networks can self-heal gray failures on two timescales.

SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL

Yan Zhan; Shaobo Liu; Qiunan Liu et al. arXiv: 2609.29050

Figure from SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL
SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL

Tool-calling agents produce heterogeneous outputs, interleaving structured tool invocations with user-facing natural language summaries. This output heterogeneity presents a structural failure mode in standard on-policy Reinforcement Learning (RL): algorithms like GRPO indiscriminately broadcast a homogeneous trajectory-level scalar advantage to all tokens. Consequently, gradient noise from summary generation leaks into tool-decision tokens, causing cross-segment credit misattribution and brittle optimization. In…

Key insight: Tool-calling RL misattributes credit across segments; SLCA-GRPO targets that failure mode.

From Self-Distillation to Self-Practice: Privileged Information for Multi-Turn Agents

Xingyu Su; Abhishek Kumar; Qing Ping et al. arXiv: 2609.29051

On-policy self-distillation (OPSD) has become a popular recipe for post-training LLM agents. It supervises the agent model at the token level with a stronger teacher view of the same model, obtained by conditioning on privileged information (PI). In this work, we show that in multi-turn agents, this paradigm teaches the student to act with confidence but without the information behind it. The trained agent behaves as if it had privileged information it never observed, and its performance falls well short of plain…

Key insight: Privileged-information self-practice improves multi-turn agents beyond self-distillation alone.

A Wrong Turn Does Not Ruin the Journey: Deviation-Guided Skill Self-Evolution for LLM Agents

Yichun Feng; Jiawei Wang; Haozhe Sun arXiv: 2609.29154

Figure from A Wrong Turn Does Not Ruin the Journey: Deviation-Guided Skill Self-Evolution for LLM Agents
A Wrong Turn Does Not Ruin the Journey: Deviation-Guided Skill Self-Evolution for LLM Agents

Large language model agents increasingly rely on natural-language skills to solve complex tool-use tasks. However, such tasks often admit multiple valid solution paths, making it inappropriate to improve skills by forcing failed trajectories to match a fixed successful trajectory. Moreover, failed trajectories are rarely entirely wrong: an agent may first collect useful evidence and make meaningful progress, but later deviate into an erroneous suffix. We therefore argue that skill self-evolution should identify…

Key insight: Deviation-guided skill self-evolution lets agents recover from wrong turns without abandoning the skill.

Policy as Code: A Coroutine-Bridge Harness for Fast-Reasoning Reliability on CAR-bench

Ivan Matveev arXiv: 2609.29251

CAR-bench evaluates whether tool-using agents stay reliable under real-world uncertainty, executing every tool inside the evaluator so that each tool-result exchange is a separate agent round-trip. A conventional next-action agent can batch parallel tool calls, but a chain of dependent calls costs it one model call per round of results. We present a coroutine-bridge harness in which the model's only action is to emit a Python program that blocks and resumes in place across evaluator tool exchanges. This decouples…

Key insight: A coroutine-bridge harness with policy-as-code improves fast-reasoning reliability on CAR-bench.

Evaluation of Multi-Turn Consistency in LLM Agents: Survival Analysis and Failure-Rationale Taxonomy

Igor Bogdanov; Olga Manakina; Chung-Horng Lung arXiv: 2609.29508

Figure from Evaluation of Multi-Turn Consistency in LLM Agents: Survival Analysis and Failure-Rationale Taxonomy
Evaluation of Multi-Turn Consistency in LLM Agents: Survival Analysis and Failure-Rationale Taxonomy

Large language model (LLM) agents may perform well on isolated tasks yet drift into inconsistency over extended interaction. We evaluate temporal consistency in a controlled 20-step multi-agent setting inspired by delayed-gratification studies. At each step, an agent chooses between continuing to delay a reward or claiming it immediately (terminating the episode). Across a full-factorial manipulation of social visibility (private vs public), persona stressors, and deliberation policy, we run 84,540 trajectories…

Key insight: Multi-turn LLM agent consistency can be measured with survival analysis and a failure-rationale taxonomy.

Stale Does Not Mean Unsafe: Guard Precision for Tool-Using LLM Agents under Infrastructure State Races

Zihao Zheng; Jiayu Long; Baichuan Li et al. arXiv: 2609.29522

Tool-using language-model agents increasingly mutate schedulers, data pipelines, object stores, and access-control systems. Between an agent's read and its commit, external state can change, but not every change makes the commit unsafe. We separate invalidating races, which break a declared safety predicate, from predicate-preserving and irrelevant races, and ask how precisely runtime guards distinguish them. Our deterministic simulator separates visible from authoritative state and injects five non-atomic failure…

Key insight: Infrastructure state races make tool-result staleness a poor proxy for unsafety—guards need precision.

PartHackBench: Certified Equal-Progress Stress Tests for Partial-Credit Tool-Agent Evaluation

Hongye Yang; Zhihao Xie; Shengjun Xiong arXiv: 2609.29578

Figure from PartHackBench: Certified Equal-Progress Stress Tests for Partial-Credit Tool-Agent Evaluation
PartHackBench: Certified Equal-Progress Stress Tests for Partial-Credit Tool-Agent Evaluation

Long-horizon tool agents often make useful progress without reaching terminal success, motivating partial-credit evaluation. Yet evaluators may reward milestones that were temporary, later reversed, or not attributable to the evaluated agent. Comparing an honest trajectory with a higher-scoring adversarial one is inconclusive if the latter made more genuine progress. We introduce PartHackBench, a controlled methodology that removes this confound. A private certifier admits a pair only when its trajectories match…

Key insight: Partial-credit tool-agent evaluation needs certified equal-progress stress tests.

iCoder-27B: Recursive AI-Led Development of Frontier Industrial Coding Model

Cheng Yang; Jiayang Lyu; Shangyuan Liu et al. arXiv: 2609.29626

Figure from iCoder-27B: Recursive AI-Led Development of Frontier Industrial Coding Model
iCoder-27B: Recursive AI-Led Development of Frontier Industrial Coding Model

Recursive AI, the prospect of AI taking an increasingly complete role in building and improving AI, is a crown jewel of AI for AI. Although recursive self-development has become practical for small models, bounded tasks, and fixed time budgets, a more consequential realization of this ambition, i.e., developing a release-ready, frontier-competitive model, remains far more challenging. In this work, we ask how little human involvement is sufficient for an agent to develop a frontier model. We concentrate human…

Key insight: Recursive AI-led development can push industrial coding models through closed self-improvement loops.

Breaking the Environment Wall: Evolving LLM Agent Environments for Recursive Self-Improvement

Yukai Wu; Yuanjing Yang; Le Zhou et al. arXiv: 2609.29773

Figure from Breaking the Environment Wall: Evolving LLM Agent Environments for Recursive Self-Improvement
Breaking the Environment Wall: Evolving LLM Agent Environments for Recursive Self-Improvement

Many real-world tasks (e.g., office workflows, scientific experimentation) require LLM agents to interact repeatedly with their environments for context-dependent operations. However, such environments are often not agent-ready. First, information is often scattered and fragmented across the environment. Second, relevant evidence in the environment is often mixed with misleading information and conflicting versions. Third, environments evolve over time, introducing new noise and more challenging tasks. These…

Key insight: Evolving the agent environment itself can unlock recursive self-improvement beyond a fixed world wall.

How does Adversarial Influence Scale in Multi-Agent Systems?

Addison J. Wu; Jasin Cekinmez; Michel Liao et al. arXiv: 2609.30028

Figure from How does Adversarial Influence Scale in Multi-Agent Systems?
How does Adversarial Influence Scale in Multi-Agent Systems?

Multi-agent deliberation can improve performance, but what happens when some agents do not act in good faith? In practice, an agent may be deceptive and work to subvert the group, whether through its own objectives or external instruction. We study how susceptibility to deception scales as groups increase in size and deceivers become more prevalent. It is not the number of agents in the group that matters, but the proportion of deceivers. We observe that the defection rate, how often initially correct agents…

Key insight: Adversarial influence in multi-agent deliberation scales in ways that threaten group outcomes.

HEXIS: Compiling Skills into Extended Finite State Machines

Minghao LI arXiv: 2609.30123

Figure from HEXIS: Compiling Skills into Extended Finite State Machines
HEXIS: Compiling Skills into Extended Finite State Machines

Agent skills provide reusable knowledge and instructions, yet agents must repeatedly infer how to apply them and which operation should follow. This couples task reasoning with control decisions, allowing prescribed steps to be omitted or applied incorrectly. We introduce HEXIS, which compiles agent skills into extended finite state machines that separate knowledge from control flow. Skill knowledge is incorporated into local instructions that guide reasoning and generation within states. The machine records…

Key insight: Compiling skills into extended finite state machines makes procedural guidance executable and checkable.

Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale

Edesio Alcoba; Kevin Rossell; Aman Gupta et al. arXiv: 2609.30137

Figure from Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale
Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale

Customer experience (CX) agents use tools and large language models to address customer requests and guide conversational interactions with an organization's products. Improving these agents, especially in regulated industries, is difficult: they must detect intent, follow complex operational policies and use tools reliably. Manual end-to-end testing offers limited coverage, while live experiments expose customers to failures that can erode trust. We present a hypothesis-driven simulation workflow for screening…

Key insight: Simulation screening before serving stabilizes production CX tool-agents at very large request scale.

Who Is Behind the Harness? Fingerprinting LLMs through Agentic Behavior

Chuyi Wang; Xiaohui Xie; Tongze Wang et al. arXiv: 2609.28559

Figure from Who Is Behind the Harness? Fingerprinting LLMs through Agentic Behavior
Who Is Behind the Harness? Fingerprinting LLMs through Agentic Behavior

LLMs increasingly operate through coding-agent harnesses that inspect repositories, invoke tools, and modify files. Substituting the model behind such an agent can therefore change security-relevant decisions, including whether it verifies changes or recovers safely from failures. Existing LLM fingerprints largely infer identity from direct text or token distributions. In coding agents, these signals are mediated by system instructions, controller logic, tools, and execution feedback, limiting their transfer. We…

Key insight: LLMs can be fingerprinted through patterns in their agentic behavior under a harness.

Decision Hijacking: Prompt Injection Attacks on Jev's Typed Probabilistic Decisions

Tiantong Wu; Wei Yang Bryan Lim arXiv: 2609.28613

Figure from Decision Hijacking: Prompt Injection Attacks on Jev's Typed Probabilistic Decisions
Decision Hijacking: Prompt Injection Attacks on Jev's Typed Probabilistic Decisions

Most studies of prompt injection focus on generative agents, leaving their effects on models with schema-defined outputs unclear. We examine these effects in Jev, a non-generative decision model, using 510 reconstructed InjecAgent cases. Malicious content shifts action probabilities but rarely causes Jev to select the attacker's target. Override markers reduce this influence, while claims of contextual relatedness have small effects. Adaptive attacks using score feedback double the mean highest attacker-target…

Key insight: Prompt injection can hijack typed probabilistic decisions in agent executors like Jev.

On the Effectiveness of Kernel-Level Evidence for Agent Security

Spencer King; Zhilu Zhang; Mikhail Kuznetsov et al. arXiv: 2609.28915

Figure from On the Effectiveness of Kernel-Level Evidence for Agent Security
On the Effectiveness of Kernel-Level Evidence for Agent Security

LLM agents are deployed into infrastructure that grants them broad host authority, yet existing agent-security benchmarks and defenses operate almost exclusively at the application telemetry layer: the served tool manifest, the user prompt, and the model's messages. Some threats, however, smuggle malicious instructions and actions past the application boundary, leaving them invisible to that layer. In this work, we bridge that gap by pairing application-level agent telemetry with kernel-level syscall traces to…

Key insight: Kernel-level telemetry is a partial but useful evidence channel for agent security.

DocuTeam: Mixed-Initiative Multi-Agent Discussions around Evolving Documents

Heechan Lee; Juhyeon Choi; Tae Soo Kim et al. arXiv: 2609.29309

Figure from DocuTeam: Mixed-Initiative Multi-Agent Discussions around Evolving Documents
DocuTeam: Mixed-Initiative Multi-Agent Discussions around Evolving Documents

In open-ended problem solving, collaborators often rely on discussion to surface concerns, challenge perspectives, and refine shared work as it evolves. While AI agents are increasingly used as discussion partners, existing multi-agent systems place a heavy burden on users to initiate and carefully orchestrate the discussions. We present DocuTeam, a mixed-initiative multi-agent discussion system in which both users and agents can initiate and steer conversations. Agents monitor document changes to proactively…

Key insight: Mixed-initiative multi-agent discussion helps humans and agents co-evolve shared documents.

IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis

Xingyu Wu; Yuchen Yan; Zhengxi Lu et al. arXiv: 2609.29444

Figure from IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis
IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis

Deep search requires LLM agents to decompose complex queries, search for evidence, and synthesize grounded answers, yet existing ReAct-style agents suffer from two limitations: role coupling, where one policy must handle planning, evidence use, and synthesis; and context accumulation, where growing search histories introduce noise and obscure useful information. To address these issues, we propose IterSynth, a role-decoupled and summary-based paradigm that alternates between a Planner for identifying information…

Key insight: Role-decoupled iterative synthesis reframes deep search agents as coordinated specialist loops.

When Agents Act Unwatched: The Reduced-Supervision Paradox in Agentic AI

Hanjing Shi; Dominic DiFranzo arXiv: 2609.29547

Agentic AI is sold on a simple promise: the system keeps acting when the user stops watching. That promise creates an accountability inversion. As stepwise supervision recedes, verification does not disappear; it moves into the runtime infrastructure that defines authority, records action, interrupts execution, checks outcomes, and supports repair. We call this the reduced-supervision paradox. Using a 63-artifact audit, we examine its public visibility across 46 research papers and 17 engineering, documentation,…

Key insight: Agent behavior under reduced supervision diverges from watched evaluation—the observation paradox.

Understanding and Exploiting Initialization Anchoring Weakness in Feedback-Based Agent Planning

Chuanchao Zang; Jianing Wang; Wenyu Chen et al. arXiv: 2609.29697

Figure from Understanding and Exploiting Initialization Anchoring Weakness in Feedback-Based Agent Planning
Understanding and Exploiting Initialization Anchoring Weakness in Feedback-Based Agent Planning

Feedback-based planning improves agent reliability by incorporating tool observations and corrective feedback. However, its protection may not be distributed uniformly across planning stages. We conduct a round-wise analysis of four representative feedback mechanisms and uncover an initialization anchoring weakness: the first feedback round corrects 46% of adversarial directions, whereas the rates fall to 13% and 7% among directions surviving into the next two rounds. Our analysis attributes this weakness to…

Key insight: Feedback-based agent planners can lock onto initialization anchors that weaken exploration.

Era by Eon: Benchmarking Enterprise Agents on Hidden Knowledge

Benjamin Gruenbaum; Doron Porat; Assaf Natanzon et al. arXiv: 2609.30055

In the Era by Eon benchmark, each question states the rules for its answer, and code computes the answer from a generated company's data. When agents can run code, the four strongest models each answer 22 to 25 of 27 such questions, so the benchmark barely separates them. We add eight question templates that depend on hidden facts. No question or document states a hidden fact, and the records that seem to hold it show something else. Other data implies it. For example, the sales system says a customer dropped a…

Key insight: Enterprise agent benches collapse when agents can run code against hidden-knowledge rules.

Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure

David Schmotz; Derck Prinzhorn; Luca Beurer-Kellner et al. arXiv: 2609.30217

A central concern in AI safety is that agents may treat oversight as an obstacle when it conflicts with completing their goals. We study instrumental evasion, the propensity of LLM agents to circumvent runtime monitoring as a means of completing ordinary tasks. We introduce EvasionBench, a benchmark of 50 diverse task-policy pairs in which completing the task requires an operation prohibited by a runtime monitor. Agents know that their tool calls are monitored and are prompted to continue working when they pause.…

Key insight: Ordinary task pressure is enough for instrumental monitor evasion to emerge in agents.