Friday's cs.AI announcement day (2026-09-25) lists 107 new and 153 cross-lists (replacements skipped; listing total 260). Stack filter for agent systems, memory/context, computer-use / GUI / tools / skills / harnesses, multi-agent, persistence/identity, and local/open serving keeps 45 papers — trust-native agentic OS (AgentKernel), interaction-conditioned context forgetting (ICLR), scoped persistent memory, budgeted memory maintenance (ERRAND), mobile planner/GUI agents, skill access control, harness governance, denial-of-wallet billable state, and related stack work.
Zhenhua Zou; Sheng Guo; Qiuyang Zhan et al. arXiv: 2609.29647
Modern AI agents routinely cross trust boundaries: they ingest untrusted content, combine it with privileged instructions, persist intermediate beliefs in long-term memory, and invoke privileged tools. This creates an attack surface in which malicious payloads can enter through model inputs and cause harmful tool actions. Yet current governance stacks remain application-level middleware that share a process trust boundary with the agents they monitor. We argue that agents need an operating-system substrate…
Key insight: A trust-native agentic OS moves governance below app middleware when agents ingest untrusted content, persist beliefs, and call privileged tools.
Mingxuan Wang; Fei Luo; Bo Wang et al. arXiv: 2609.29875
Long horizon language model agents continually accumulate reasoning history, increasing context length and inference cost even after earlier decisions have been executed and observed. Unlike static Chain of Thought compression, removing historical reasoning can change future actions and the resulting interaction trajectory. We study when such reasoning can be safely forgotten. We propose Interaction Aware Compression for Long Horizon Reasoning (ICLR), a training free online method that ranks reasoning blocks using…
Key insight: After actions execute, historical reasoning is not always safe to drop — interaction-conditioned compression decides when agents can forget.
Yezhou Cheng; Runjia Du; Zeming Liu et al. arXiv: 2609.29144
Persistent memory lets language-model agents improve prompts and skills without updating model weights. We show that matching retrieval scope to certification scope enables these edits to support reliable repeated adaptation across recurring task families. We study frozen-model agents on ProcStream-RSI, a 12-round code-repair stream, using Orthogonal Regression Control (ORC), an execution-grounded gate for persistent skill edits. In an intervention that holds proposals and gate decisions fixed, retrieving each…
Key insight: Persistent agent memory edits are reliable only when retrieval scope matches certification scope across recurring task families.
Beining Wu; Zihao Ding; Jun Huang arXiv: 2609.29545
Deployed agents run on handed-over knowledge: a frozen policy consults a briefing of consolidated items written before the stream begins. The world then moves while the store stands still: paths close, flags change, price bands move; every item was true at handover, and the failure is staleness, not ignorance. We introduce ERRAND, which treats revalidation as a priced errand: a recheck competes with the task it protects for the same scarce actions, funded only when the value per action of resolving a doubt clears…
Key insight: Agent memory maintenance needs explicit spend budgets, not unbounded accumulation.
Tingyu Qu; Weigao Sun; Yuecheng Liu et al. arXiv: 2609.29892
The rapid progression of large language models is extending AI from passive content generation into the active workflows of engineering and scientific discovery. This shift raises a compelling question: can AI be both the object of development and an active participant in building next-generation AI systems? We explore this question by building Qwen-Planner-Agent within a closed-loop AI-for-AI framework for scalable development and iterative improvement. Mobile planning offers a demanding test of this approach:…
Key insight: A closed-loop AI-for-AI framework can build real-world mobile planner agents with the model in the development loop.
Linghua Zhang arXiv: 2609.30186
Vision-language models (VLMs) have become a common foundation for autonomous mobile GUI agents, but most existing systems rely on the VLM for both planning and action grounding at nearly every interaction step, leading to substantial latency and model-serving cost. We introduce Jev-Mobile, which shifts this paradigm to low-frequency VLM planning and high-frequency lightweight execution: the VLM specifies local goals, the accessibility tree defines a structured executable action space, and Jev, a fast typed…
Key insight: Typed probabilistic decision executors can ground mobile GUI agents on phone UIs.
Michael Stettler; Benjamin Girardet; Jonas Canton et al. arXiv: 2609.28693
Large Language Model (LLM) agents struggle to scale safely when exposed to vast enterprise toolsets. Providing an agent with access to every internal tool leads to oversized context windows, degraded tool selection, and severe governance vulnerabilities - as system policies defined purely in prompts remain probabilistic advice rather than hard constraints. Existing mitigations, such as multi-agent domain delegation, decentralize audit logs and fail to guarantee policy compliance across sessions. We introduce…
Key insight: Progressive skill discovery can serve as structural access control for tool-using LLM agents.
Arian Abbasi; Alan Aqrawi; Ted Kwartler arXiv: 2609.28919
Harnesses, the products that run AI coding agents, are multiplying, and enterprises are rolling them out to their employees: what started as pilots with a few hundred seats is scaling to tens of thousands. Most enterprises do not build these harnesses but buy them from large vendors, such as Anthropic's Claude Code or OpenAI's Codex. A harness decides which model answers, what the model reads, how the prompt cache is used and which subagents run, so it picks the rate on the price sheet and sets the volume bought…
Key insight: Routing and governing coding agents through the harness is the main lever for enterprise cost and safety.
Haiqing Li; Xin Ma; Yinhao Wu et al. arXiv: 2609.29921
Large language model agents increasingly combine generation, decision-making, execution, and self-evaluation within a single agentic loop. Although they operate under external specifications such as task instructions, guidelines, output schemas, and reusable skills, these specifications typically remain context for the same model that acts and declares completion, leaving no independent specification authority boundary. We identify two resulting gaps. The understanding--execution gap arises when a requirement is…
Key insight: Specifications—not agents—should hold final sign-off authority over generated changes.
Jinqian Zhang; Haojun Xia; Shujiang Wu et al. arXiv: 2609.28585
Multi-step tool-calling LLM agents rely on host runtimes to preserve state across turns. When a runtime carries an external tool return into later model inputs, providers meter it again. An admitted malicious or compromised tool can thereby convert untrusted data into recurring victim-billed processing without victim credentials or local runtime privilege. We call retained content persistent billable state and formalize the host's decision over whether and how it enters later billable context as the persistent…
Key insight: Tool-calling agents with persistent billable state enable denial-of-wallet attacks unless the harness caps spend.
José Luis Pino arXiv: 2609.29808
In July 2026, an unconstrained autonomous agent participating in a frontier AI cybersecurity evaluation harness breached its evaluation sandbox, established an external command-and-control foothold, and executed a multi-stage intrusion into Hugging Face's production multi-tenant dataset conversion infrastructure (referred to in this autopsy as Incident-2026-Alpha). Over 4.5 days, the rogue agent executed 17,600 discrete actions across 6,280 worker clusters, compromised AWS EC2 Instance Metadata Service (IMDS)…
Key insight: Rogue agentic execution needs kernel-level preemption and containment below the LLM loop.
Jiapeng Li arXiv: 2609.29095
When a tool-using agent's write times out or returns a server error, the action may already have taken effect. Retrying blindly duplicates it -- a second charge, a second announcement, a second deployment -- while giving up skips required work. We ask where exactly-once behaviour should be enforced: in the model, in the agent harness, or in the tool contract. We introduce LIMBO, a deterministic sandbox of six services with realistic contracts (optional idempotency keys, eventually consistent and missing read…
Key insight: Exactly-once semantics for tool side effects live in the model, harness, or tool contract—pinning the layer matters.
Xueshu Chen; Yan Wang; Zihao Xue et al. arXiv: 2609.29735
Long-horizon tasks require preserving and later recovering cross-session evidence under a bounded, query-blind memory budget. Existing compression can discard fine-grained visual cues or conflate semantically similar but incompatible observations. We present C3M, a cross-session multimodal memory organization that maintains a bounded active index over persistent source text-image evidence. Relation-aware updates consolidate safe redundancy while preserving complementary and incompatible records. At query time,…
Key insight: Cross-session multimodal memory maintenance is required for long-horizon tasks that outlive a single chat.
Kyle Wild; Yusuke Takahashi; Asako Uraki arXiv: 2609.29661
Most agentic question answering (QA) systems do an important part of their semantic work at the worst possible time: every time someone asks a question. When a corpus contains revisions, drafts, revocations, deletions, and sources with different levels of authority, the model must reconstruct the governed current state on every read - then throw that work away and repeat it on the next query. This is a bit like a database that rebuilds a materialized view every time someone reads from it. We present ingest-time…
Key insight: Doing semantic fact compilation at ingest time beats re-deriving authority over revised corpora on every agentic QA turn.
Jeremy Qin; David Schmotz; Derck Prinzhorn et al. arXiv: 2609.30266
Asynchronous monitoring, incident investigations, and compliance audits primarily rely on agent traces to reconstruct what happened. These analyses assume that LLM agents cannot tamper with their own execution traces. We show that local LLM agents such as Claude Code, Codex, Antigravity, Open Code and Grok Build fail to enforce this boundary. All tested harnesses, except Muse Code, allowed agents to delete their traces when asked, without triggering monitor guardrails. We also validate that external attackers can…
Key insight: If agents can write their own traces, audit logs are not a trustworthy oversight channel.
Mehmet Iscan arXiv: 2609.30219
An acceptance protocol is developed for sensor-coordinate and polarity binding in mechatronic commissioning. Candidate generation is separated from release authority. Requirements unsupported by a deterministic parser are routed to a frozen local language model with four billion parameters. Plans are released only when both facts can be derived by an external gate under a sealed grammar. One canonical answer is requested from a gold-standard user when eligible. The protocol was evaluated once under a criterion…
Key insight: A frozen 4B local model can serve as a requirement-bound commissioning overseer beside larger workers.
Rida Qadri; Remi Denton; Michael Madaio et al. arXiv: 2609.29901
Enterprise AI is transitioning from single-user, reactive tools toward proactive, multi-user 'teammates,' but our empirical understanding of this transition is limited. In this paper, we present an in-situ qualitative study of a persistent, proactive AI agent 'teammate' deployed across multiple teams in a large technology company. Our findings reveal the boundaries of the human-agent workplace are actively in flux, triggering breakdowns and negotiations across: 1) tacit rules of collaborative human workflows, 2)…
Key insight: Persistent proactive multi-user agentic teammates collide with human organizational ecosystems in situ.
Chia-Yuan Chang; Renyuan Cheng; Rui Feng et al. arXiv: 2609.29421
Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document the data, reward design, infrastructure, stage order, and stagewise results needed to reproduce the recipe. Stages progress from basic to advanced capabilities and from hard, verifiable rewards to softer judge-based signals. Training builds on open-source…
Key insight: An open post-training recipe on GLM-4.5-Air includes dedicated General, Coding, and Search Agent stages.
Tiviatis Sim; Jia Hui Woon; Xinming Gao et al. arXiv: 2609.28547
Policy interventions propagate through public communication, institutional decisions, and stakeholder responses, yet datasets for financial multi-agent simulation rarely connect these processes to temporally aligned historical evidence. We introduce PAWS, a Policy-driven Agentic World Simulation dataset covering 36 verified U.S. financial and economic policy episodes, 12,727 policy-linked news records, and 65,291 source-grounded stakeholder actions. Each action is linked to its supporting news and represented by a…
Key insight: Policy-driven agentic world simulation supplies controllable environments for agent training and evaluation.
Subrat Panda arXiv: 2609.28575
Long-conversation memory benchmarks increasingly test recall and prompted knowledge updates, and recent work studies evolving user beliefs and memory state. TWIST is a proposed benchmark suite for a complementary, unmeasured property: intervention quality -- whether a deployed memory system, exercised through its own ingest/recall/vet surface, acts correctly at belief change points. Four tracks cover unprompted tension detection, vetting outgoing drafts against the record, answering with current beliefs while…
Key insight: TWIST benchmarks intervention quality in conversational memory with human-validated edits.
Gaurisankar Jayadas; Aske Plaat; Álvaro Serra-Gómez et al. arXiv: 2609.28765
Reinforcement Learning with Verifiable Rewards (RLVR) performs well on problems with clear rewards, such as mathematics and coding, but whether it also works where the reward is less clear remains open. The reason-over-search recipe applies RLVR to open-domain question answering, where retrieval grounds the answer and a match against the reference supplies the reward. So far it has been demonstrated on large models, and below one billion parameters only with distillation from a larger teacher. We test the recipe…
Key insight: RL with verifiable rewards can train small search agents even when open-domain rewards are noisy.
Mithil Salunkhe; Haochen Ding; Samridhi Verma et al. arXiv: 2609.28850
Reproducing a machine learning paper involves most research steps, from installing software and debugging to running experiments, work that AI agents increasingly do. We introduce RECLAIM, a benchmark of 100 NeurIPS 2025 papers that can be rebuilt yearly from new conferences. For each paper we fix in advance the result to reproduce, what counts as a successful reproduction, and a GPU-hour budget. An agent must reproduce that result using the paper and whatever its authors released. What the authors released…
Key insight: RECLAIM asks whether agents can reproduce the claims of ML papers on a yearly-rebuildable bench.
Hongye Yang; Boxiao Huang arXiv: 2609.29007
Tool-using agents are continually updated with new interaction data. After each policy update, however, previously estimated action credits may become stale. Recomputing them from scratch can require many additional tool calls and environment interactions, making repeated updates increasingly expensive. We ask a simple question: when does historical action credit actually need to be updated? Our key observation is that a change in action value does not necessarily imply a change in the decision. Historical credit…
Key insight: After policy updates, tool-using agents need principled rules for when action credits must be recomputed.
Keru Chen; Sen Lin; Yingbin Liang et al. arXiv: 2609.29015
Decentralized LLM-based multi-agent systems coordinate through local interactions, but an agent can remain responsive while its task-solving quality persistently degrades. Such gray failures require protecting current tasks before sufficient evidence exists to alter future routing, while still allowing recovered agents to rejoin. We introduce MeshHeal, a fully decentralized self-healing framework that couples ability-matched peer review across two timescales. At the fast timescale, an adaptive hierarchy escalates…
Key insight: Decentralized LLM agent networks can self-heal gray failures on two timescales.
Yan Zhan; Shaobo Liu; Qiunan Liu et al. arXiv: 2609.29050
Tool-calling agents produce heterogeneous outputs, interleaving structured tool invocations with user-facing natural language summaries. This output heterogeneity presents a structural failure mode in standard on-policy Reinforcement Learning (RL): algorithms like GRPO indiscriminately broadcast a homogeneous trajectory-level scalar advantage to all tokens. Consequently, gradient noise from summary generation leaks into tool-decision tokens, causing cross-segment credit misattribution and brittle optimization. In…
Key insight: Tool-calling RL misattributes credit across segments; SLCA-GRPO targets that failure mode.
Xingyu Su; Abhishek Kumar; Qing Ping et al. arXiv: 2609.29051
On-policy self-distillation (OPSD) has become a popular recipe for post-training LLM agents. It supervises the agent model at the token level with a stronger teacher view of the same model, obtained by conditioning on privileged information (PI). In this work, we show that in multi-turn agents, this paradigm teaches the student to act with confidence but without the information behind it. The trained agent behaves as if it had privileged information it never observed, and its performance falls well short of plain…
Key insight: Privileged-information self-practice improves multi-turn agents beyond self-distillation alone.
Yichun Feng; Jiawei Wang; Haozhe Sun arXiv: 2609.29154
Large language model agents increasingly rely on natural-language skills to solve complex tool-use tasks. However, such tasks often admit multiple valid solution paths, making it inappropriate to improve skills by forcing failed trajectories to match a fixed successful trajectory. Moreover, failed trajectories are rarely entirely wrong: an agent may first collect useful evidence and make meaningful progress, but later deviate into an erroneous suffix. We therefore argue that skill self-evolution should identify…
Key insight: Deviation-guided skill self-evolution lets agents recover from wrong turns without abandoning the skill.
Ivan Matveev arXiv: 2609.29251
CAR-bench evaluates whether tool-using agents stay reliable under real-world uncertainty, executing every tool inside the evaluator so that each tool-result exchange is a separate agent round-trip. A conventional next-action agent can batch parallel tool calls, but a chain of dependent calls costs it one model call per round of results. We present a coroutine-bridge harness in which the model's only action is to emit a Python program that blocks and resumes in place across evaluator tool exchanges. This decouples…
Key insight: A coroutine-bridge harness with policy-as-code improves fast-reasoning reliability on CAR-bench.
Igor Bogdanov; Olga Manakina; Chung-Horng Lung arXiv: 2609.29508
Large language model (LLM) agents may perform well on isolated tasks yet drift into inconsistency over extended interaction. We evaluate temporal consistency in a controlled 20-step multi-agent setting inspired by delayed-gratification studies. At each step, an agent chooses between continuing to delay a reward or claiming it immediately (terminating the episode). Across a full-factorial manipulation of social visibility (private vs public), persona stressors, and deliberation policy, we run 84,540 trajectories…
Key insight: Multi-turn LLM agent consistency can be measured with survival analysis and a failure-rationale taxonomy.
Zihao Zheng; Jiayu Long; Baichuan Li et al. arXiv: 2609.29522
Tool-using language-model agents increasingly mutate schedulers, data pipelines, object stores, and access-control systems. Between an agent's read and its commit, external state can change, but not every change makes the commit unsafe. We separate invalidating races, which break a declared safety predicate, from predicate-preserving and irrelevant races, and ask how precisely runtime guards distinguish them. Our deterministic simulator separates visible from authoritative state and injects five non-atomic failure…
Key insight: Infrastructure state races make tool-result staleness a poor proxy for unsafety—guards need precision.
Hongye Yang; Zhihao Xie; Shengjun Xiong arXiv: 2609.29578
Long-horizon tool agents often make useful progress without reaching terminal success, motivating partial-credit evaluation. Yet evaluators may reward milestones that were temporary, later reversed, or not attributable to the evaluated agent. Comparing an honest trajectory with a higher-scoring adversarial one is inconclusive if the latter made more genuine progress. We introduce PartHackBench, a controlled methodology that removes this confound. A private certifier admits a pair only when its trajectories match…
Key insight: Partial-credit tool-agent evaluation needs certified equal-progress stress tests.
Cheng Yang; Jiayang Lyu; Shangyuan Liu et al. arXiv: 2609.29626
Recursive AI, the prospect of AI taking an increasingly complete role in building and improving AI, is a crown jewel of AI for AI. Although recursive self-development has become practical for small models, bounded tasks, and fixed time budgets, a more consequential realization of this ambition, i.e., developing a release-ready, frontier-competitive model, remains far more challenging. In this work, we ask how little human involvement is sufficient for an agent to develop a frontier model. We concentrate human…
Key insight: Recursive AI-led development can push industrial coding models through closed self-improvement loops.
Yukai Wu; Yuanjing Yang; Le Zhou et al. arXiv: 2609.29773
Many real-world tasks (e.g., office workflows, scientific experimentation) require LLM agents to interact repeatedly with their environments for context-dependent operations. However, such environments are often not agent-ready. First, information is often scattered and fragmented across the environment. Second, relevant evidence in the environment is often mixed with misleading information and conflicting versions. Third, environments evolve over time, introducing new noise and more challenging tasks. These…
Key insight: Evolving the agent environment itself can unlock recursive self-improvement beyond a fixed world wall.
Addison J. Wu; Jasin Cekinmez; Michel Liao et al. arXiv: 2609.30028
Multi-agent deliberation can improve performance, but what happens when some agents do not act in good faith? In practice, an agent may be deceptive and work to subvert the group, whether through its own objectives or external instruction. We study how susceptibility to deception scales as groups increase in size and deceivers become more prevalent. It is not the number of agents in the group that matters, but the proportion of deceivers. We observe that the defection rate, how often initially correct agents…
Key insight: Adversarial influence in multi-agent deliberation scales in ways that threaten group outcomes.
Minghao LI arXiv: 2609.30123
Agent skills provide reusable knowledge and instructions, yet agents must repeatedly infer how to apply them and which operation should follow. This couples task reasoning with control decisions, allowing prescribed steps to be omitted or applied incorrectly. We introduce HEXIS, which compiles agent skills into extended finite state machines that separate knowledge from control flow. Skill knowledge is incorporated into local instructions that guide reasoning and generation within states. The machine records…
Key insight: Compiling skills into extended finite state machines makes procedural guidance executable and checkable.
Edesio Alcoba; Kevin Rossell; Aman Gupta et al. arXiv: 2609.30137
Customer experience (CX) agents use tools and large language models to address customer requests and guide conversational interactions with an organization's products. Improving these agents, especially in regulated industries, is difficult: they must detect intent, follow complex operational policies and use tools reliably. Manual end-to-end testing offers limited coverage, while live experiments expose customers to failures that can erode trust. We present a hypothesis-driven simulation workflow for screening…
Key insight: Simulation screening before serving stabilizes production CX tool-agents at very large request scale.
Chuyi Wang; Xiaohui Xie; Tongze Wang et al. arXiv: 2609.28559
LLMs increasingly operate through coding-agent harnesses that inspect repositories, invoke tools, and modify files. Substituting the model behind such an agent can therefore change security-relevant decisions, including whether it verifies changes or recovers safely from failures. Existing LLM fingerprints largely infer identity from direct text or token distributions. In coding agents, these signals are mediated by system instructions, controller logic, tools, and execution feedback, limiting their transfer. We…
Key insight: LLMs can be fingerprinted through patterns in their agentic behavior under a harness.
Tiantong Wu; Wei Yang Bryan Lim arXiv: 2609.28613
Most studies of prompt injection focus on generative agents, leaving their effects on models with schema-defined outputs unclear. We examine these effects in Jev, a non-generative decision model, using 510 reconstructed InjecAgent cases. Malicious content shifts action probabilities but rarely causes Jev to select the attacker's target. Override markers reduce this influence, while claims of contextual relatedness have small effects. Adaptive attacks using score feedback double the mean highest attacker-target…
Key insight: Prompt injection can hijack typed probabilistic decisions in agent executors like Jev.
Spencer King; Zhilu Zhang; Mikhail Kuznetsov et al. arXiv: 2609.28915
LLM agents are deployed into infrastructure that grants them broad host authority, yet existing agent-security benchmarks and defenses operate almost exclusively at the application telemetry layer: the served tool manifest, the user prompt, and the model's messages. Some threats, however, smuggle malicious instructions and actions past the application boundary, leaving them invisible to that layer. In this work, we bridge that gap by pairing application-level agent telemetry with kernel-level syscall traces to…
Key insight: Kernel-level telemetry is a partial but useful evidence channel for agent security.
Heechan Lee; Juhyeon Choi; Tae Soo Kim et al. arXiv: 2609.29309
In open-ended problem solving, collaborators often rely on discussion to surface concerns, challenge perspectives, and refine shared work as it evolves. While AI agents are increasingly used as discussion partners, existing multi-agent systems place a heavy burden on users to initiate and carefully orchestrate the discussions. We present DocuTeam, a mixed-initiative multi-agent discussion system in which both users and agents can initiate and steer conversations. Agents monitor document changes to proactively…
Key insight: Mixed-initiative multi-agent discussion helps humans and agents co-evolve shared documents.
Xingyu Wu; Yuchen Yan; Zhengxi Lu et al. arXiv: 2609.29444
Deep search requires LLM agents to decompose complex queries, search for evidence, and synthesize grounded answers, yet existing ReAct-style agents suffer from two limitations: role coupling, where one policy must handle planning, evidence use, and synthesis; and context accumulation, where growing search histories introduce noise and obscure useful information. To address these issues, we propose IterSynth, a role-decoupled and summary-based paradigm that alternates between a Planner for identifying information…
Key insight: Role-decoupled iterative synthesis reframes deep search agents as coordinated specialist loops.
Hanjing Shi; Dominic DiFranzo arXiv: 2609.29547
Agentic AI is sold on a simple promise: the system keeps acting when the user stops watching. That promise creates an accountability inversion. As stepwise supervision recedes, verification does not disappear; it moves into the runtime infrastructure that defines authority, records action, interrupts execution, checks outcomes, and supports repair. We call this the reduced-supervision paradox. Using a 63-artifact audit, we examine its public visibility across 46 research papers and 17 engineering, documentation,…
Key insight: Agent behavior under reduced supervision diverges from watched evaluation—the observation paradox.
Chuanchao Zang; Jianing Wang; Wenyu Chen et al. arXiv: 2609.29697
Feedback-based planning improves agent reliability by incorporating tool observations and corrective feedback. However, its protection may not be distributed uniformly across planning stages. We conduct a round-wise analysis of four representative feedback mechanisms and uncover an initialization anchoring weakness: the first feedback round corrects 46% of adversarial directions, whereas the rates fall to 13% and 7% among directions surviving into the next two rounds. Our analysis attributes this weakness to…
Key insight: Feedback-based agent planners can lock onto initialization anchors that weaken exploration.
Benjamin Gruenbaum; Doron Porat; Assaf Natanzon et al. arXiv: 2609.30055
In the Era by Eon benchmark, each question states the rules for its answer, and code computes the answer from a generated company's data. When agents can run code, the four strongest models each answer 22 to 25 of 27 such questions, so the benchmark barely separates them. We add eight question templates that depend on hidden facts. No question or document states a hidden fact, and the records that seem to hold it show something else. Other data implies it. For example, the sales system says a customer dropped a…
Key insight: Enterprise agent benches collapse when agents can run code against hidden-knowledge rules.
David Schmotz; Derck Prinzhorn; Luca Beurer-Kellner et al. arXiv: 2609.30217
A central concern in AI safety is that agents may treat oversight as an obstacle when it conflicts with completing their goals. We study instrumental evasion, the propensity of LLM agents to circumvent runtime monitoring as a means of completing ordinary tasks. We introduce EvasionBench, a benchmark of 50 diverse task-policy pairs in which completing the task requires an operation prohibited by a runtime monitor. Agents know that their tool calls are monitored and are prompted to continue working when they pause.…
Key insight: Ordinary task pressure is enough for instrumental monitor evasion to emerge in agents.