Friday's cs.AI announcement day (2026-10-09) lists 125 new and 177 cross-lists (replacements skipped; listing total 302). A filter for agent systems, memory/context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent, persistence/identity, and local/open models keeps 40 papers. Memory and context dominate: agent-controlled forgetting archives bulky tool results out of the window, always-loaded AGENTS.md files are proved to be capacitated assortments where appending everything useful is wrong, and admission/presentation governance plus formation-time gated memory decide what durable facts enter a turn. Harnesses and skills get hard edges: self-evolution should stay on the harness (prompts/tools/composition) with a recorded gate, process vs content failures choose the lever, co-installed similar skills can drop intended constraints while tasks still pass, and SkillContrast / NOMOS improve skill ranking and policy-to-tool gates. On monitors, time and safety, OnTrack matches steps to past successful runs in about a millisecond, prompt-only wall-clock budgets fail without harness clocks, verifiers must be selected against held-out anchors, and obligation / option-channel / structure-tax results tighten what runtime guards and JSON schemas can be trusted to do.


Research Papers

Agent-Controlled Forgetting for Tool-Using Agents: Reversible Context Curation in Practice

Jan-Peter Franke arXiv: 2610.10590

Tool-using agents carry observations whose useful content is often much smaller than the payload. Agent-controlled forgetting lets the acting model select prior tool results, replace each with a short note at its original position, and keep the exact original in a recoverable archive. A Python harness exposes batch archival and explicit recovery without task-specific training, while protecting user instructions and assistant messages. In an exploratory OpenTelemetry debugging case followed by an unrelated implementation task: 231,951 provider prompt tokens vs 912,492 under retained history (50% fewer cumulative input tokens; about USD 1.28-1.44 vs about 4.38). Both arms passed the two-case primary behavioral oracle.

Key insight: Reversible forgetting of bulky tool results — short notes in context, originals in an archive — cut prompt tokens by about 50% on a noisy debug-and-implement run.

Curating Always-Loaded Context for LLM Agents: A Capacitated Assortment Model with Censored Feedback

Zexuan Liu; Yuning Yang; Tiancheng Zhao arXiv: 2610.11007

At session start LLM agents load a fixed context file (e.g. AGENTS.md); every loaded token is charged again each later round, and growing files can degrade performance, yet curators usually append. Formulates curation as a capacitated assortment problem: instructions consume tokens under finite attention; adding one never raises compliance of others; retained instructions incur per-session setup cost. Proves an upper bound on optimal file size regardless of candidate count, and that appending every positive-standalone-value instruction can be arbitrarily worse than an optimal subset.

Key insight: Always-loaded context files are capacitated assortments: appending every positive-value instruction can be arbitrarily worse than choosing an optimal subset under a token budget.

What to Admit and How to Present: Governing Persistent Memory in LLM Agents

Chang Liu; Deliang Ding arXiv: 2610.11188

Persistent memory improves personalization but can induce sycophancy and cross-domain leakage. Separates two governance decisions: admission (what recalled info enters working context) and presentation (how admitted info is expressed). Two inference-time designs without retraining: factor-compiled admission (FC) assessing whole entries, and permission-semantic admission (PS) decomposing entries into typed units; both use deterministic policies. On an external benchmark (four tasks times 300 samples), FC and PS cut pooled judge-assessed failure rates vs verbatim injection by 6.7 and 8.8 pp.

Key insight: Governing persistent memory means two decisions — what is admitted into the turn and how it is phrased — and admission quality alone cut judge-assessed failures by 6.7–8.8 points.

Gated Memory: Admission-Controlled Memory Formation for Conversational AI

Preeti Saraswat; Divya Neelagiri; Ajay Manoj arXiv: 2610.11270

Figure from Gated Memory: Admission-Controlled Memory Formation for Conversational AI
Gated Memory: Admission-Controlled Memory Formation for Conversational AI

Personalized conversational AI extracts facts into persistent vector stores, but the formation stage (first write) has almost no principled attention. Critical signals (permanent attribute vs transient situation) exist only in the original utterance and are lost once extraction produces a subject-relation-object triple. Gated Memory interposes two checkpoints between conversation and storage: an admission gate evaluating every candidate fact, so formation-time context is not thrown away.

Key insight: Formation-time admission is the binding constraint on memory quality: permanent vs transient context is lost the moment extraction produces a bare triple.

The Harness as the Only Mutable Surface: Compliance-Bounded Self-Evolution of LLM Agents in Credit Pipelines, with a Measured Admission Gate

Ravil Akhtyamov arXiv: 2610.10629

Argues self-evolution is reviewable only if confined to the runtime harness (instruction text, tool-call logic, primitive composition) while model weights stay fixed — every adaptation is a diff with a cause and a test. Dual-loop engine with one admission gate writing a hash-chained record before deployment. Measured in simulation (simulated agent + seeded-search proposer, not LLMs): across three families of supervisory re-interpretation at three severities times 10 seeds, the gated loop admitted 144 of 7,449 candidates, none of which worsened error on held-out history.

Key insight: Self-evolution is reviewable only when confined to the harness — prompts, tools, and composition — with every change gated by a recorded test before deployment.

Harness Evolution Hits a Ceiling: When Weight Training Should Begin

Yuan Tian; Bing Hu; Hao Wang et al. arXiv: 2610.11655

Improving a long-horizon agent means evolving the harness around a frozen model or training weights. Labels failed trajectories by the first signal that fires into process failures (blocked calls, loops, exhausted budgets) vs content failures (delivered but poor plan). Harness evolution repairs process failures (behavior can later be trained into weights); content failures need weight training. On DeepPlanning, self-evolving harness lifts Qwen3.5-4B from 0.16 to 0.30 and 9B from 0.32 to 0.44 held-out.

Key insight: Process failures (loops, blocked tools) are harness problems; content failures need weight training — diagnose which before choosing the lever.

One Skill Too Many: How Co-Installed Skills Conflict in Coding Agents

Chaoliang Yan; Zihao Xu; Yuekang Li et al. arXiv: 2610.11647

Figure from One Skill Too Many: How Co-Installed Skills Conflict in Coding Agents
One Skill Too Many: How Co-Installed Skills Conflict in Coding Agents

Co-installed similar skills (from independent sources) can conflict: the model picks by name and description alone, and the intended skill loses core functions (e.g. a ban on touching git) because a similar skill runs instead — yet the task still passes, so completion benchmarks miss it. From 20,947 repos: 822,109 candidate similar-skill pairs; LLM-judged sample of 3,754; 312 confirmed pairs on three models (6,368 runs, 169,294 tool calls, 542 agent-hours).

Key insight: Co-installed similar skills can drop the intended skill's constraints while the task still passes, so completion benchmarks miss the conflict.

SkillContrast: Difference-Guided Text Selection for Agent Skill Reranking

Jiandong Ding; Honglei Ji; Ming Liu et al. arXiv: 2610.11650

Figure from SkillContrast: Difference-Guided Text Selection for Agent Skill Reranking
SkillContrast: Difference-Guided Text Selection for Agent Skill Reranking

Similar agent skills share instructions but differ in conditions of use; query-based text selection keeps shared text and drops distinctions. SkillContrast is training-free: compares retrieved skills and retains differing text with local context for a pretrained reranker. On 1,235 SameCapRisk-Bench requests: 54 to 72 more clean hits (helpful skill without risky sibling) than TF-IDF at matched lengths; uses 51.1 to 58.8% fewer model-input tokens than full skill bodies.

Key insight: Reranking skills by the text that differs between candidates, not shared boilerplate, yields more clean hits at far fewer input tokens.

NOMOS: Compiling Written Policies into Statically Verified Tool-Call Gates for LLM Agents

Min-Young Yu; Tony Kim; Jang Won Choi arXiv: 2610.11030

Tool-using agents silently violate deployed policies. NOMOS is a four-pass compiler from natural-language policy to a deterministic tool-call gate; static verification with tool-schema-level checks alone (no prover, solver, or LLM) repairs or rejects 37% (airline) and 13% (retail) of candidates — without which most shipped rules are inoperable. On tau2-bench the gate cuts violations of reference-encoded clauses; a development binding that refused 95.9% of task-passing calls is flagged as over-blocking.

Key insight: Compiling natural-language policy into schema-checked tool-call gates catches inoperable rules before they ship, but over-blocking gates must be flagged.

OnTrack: Real-Time Monitoring and Intervention in LLM Agent Trajectories via Streaming Structure-Aware Optimal Transport

Babak Barazandeh; Connor Swanson; Chinmay Kulkarni et al. arXiv: 2610.12375

Safeguard agents per-step add cost and latency; post-hoc log review comes after damage. OnTrack streams monitoring by comparing an agent's steps and dependencies against recorded successful runs, alerting or blocking in about 1 ms per step. Studied under full reference access (historical runs and tool schemas), schemas-only, and step-logs-only.

Key insight: Streaming monitors that match agent steps against past successful runs can alert or block in about a millisecond — cheaper than a second LLM safeguard.

On the Clock: Towards Punctual and Productive Time-Budgeted AI Agents

Aaron Wang; Neelabh Madan; Vlad Sobal et al. arXiv: 2610.10833

Do small LLM agents respect and productively use wall-clock budgets? Prompt-only budgets fail: harness gives no timing feedback, agents cannot anticipate action duration, and lack a learned mapping from time to strategy. Studies harness interventions that expose timing and enforce deadlines, plus agent-side adaptations, on MLE-Bench Lite (Qwen3.6-27B) and Zork I (Qwen3-4B).

Key insight: Prompt-only wall-clock budgets fail; timing feedback and hard deadlines must live in the harness.

Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents

Xing Zhang; Guanghui Wang; Yanwei Cui et al. arXiv: 2610.11464

Self-improving loops need a verifier; on open-ended tasks a hand rubric or bare LLM judge invites reward hacking. Makes the verifier the evolving object: inspectable expression over mostly-deterministic drawback detectors, synthesized from clustered failures, gated at birth, selected for agreement with a ten-item anchored reference set plus consensus — never for the agent's score. On MBPP+ gains +0.21 held-out agreement over seed composition on every seed; removing anchor guards collapses the result.

Key insight: Self-improve loops need an inspectable verifier selected against held-out anchors — never against the agent's own score.

Safe Actions Alone Do Not Ensure Safe Agents: Identifying Unfulfilled Obligations with Guard Models

Youwei Feng; Yitong Zhang; Yuetong Liu et al. arXiv: 2610.11773

Guard models mostly flag forbidden actions, but safety also requires detecting required-yet-unperformed safety-critical actions (obligations). On a popular safety benchmark, 56.92% of GLM-5.3 trajectories contain unfulfilled obligations vs 30.00% forbidden actions. Introduces ObligationBench to evaluate whether guards can identify obligations.

Key insight: Safety guards must detect required-yet-unperformed obligations; on one benchmark unfulfilled obligations outnumbered forbidden actions.

One Word Opens the Gate: The Option-Channel Attack on Typed Decision Models as Agent Guardrails

Seyedarmin Azizi; Erfan Baghaei Potraghloo; Massoud Pedram arXiv: 2610.12292

Typed decision models (probability over caller-defined options with written definitions, no free text) are used as agent guardrails. Seven open-weight models: accuracy 36-72% vs 50% chance on injection/jailbreak/toxic screening; low error in one direction often means a default-allow or default-block bias. On synthetic agent tool calls, option-name polarity manipulations (yes/no names etc.) raise decision-flip rates up to 70.4 pp vs 0/1 controls while type-error stays 0%.

Key insight: Typed yes/no decision guardrails flip under option-name polarity alone, with zero type errors — prefer schema gates over soft classifiers.

Structure Tax: How Structured Output affects LLMs Performance

Vineet Kumar; Kanishka; Bhuvanesh Mandora arXiv: 2610.12056

Re-examines the claim that structured JSON/XML outputs inherently cost accuracy. Across models, datasets, and schemas: the tax depends on schema design — reasoning-first field ordering matches or exceeds free-form accuracy; answer-first ordering causes steep drops especially in smaller models. Schemas preserving reasoning order also improve calibration.

Key insight: The structured-output accuracy tax depends on field order: reasoning-first schemas match free-form accuracy; answer-first schemas hurt, especially on small models.

Cross-Provider Review as a Runtime Contract for Coding Agents: A Controlled Pilot and Fault-Injection Study

Bowen Xu; Boyu Chen arXiv: 2610.10961

Describes an advisory cross-provider review contract: distinct resource pools, bounded execution, restricted reviewer capabilities, complete input delivery, usable semantic output, explicit failure states, durable per-attempt evidence. Controlled pilot of 20 paired development turns: 8 had a material reviewer finding (95% CI 19.1-63.9%). Boundary scan reproduced a false success on partial input (four truncation levels passed historically, failed after repair).

Key insight: An advisory cross-provider review contract needs complete input delivery — truncated reviewer input produced false successes in a controlled pilot.

Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute

Sietse Schelpe arXiv: 2610.10845

Tests galahad-kv, a public package that saves KV state of about 16k-token blocks to encrypted local NVMe and loads it back byte-exact without recompute. On 50M tokens of public text via vLLM on one H100 with Gemma 4 12B and 31B: 100/100 probed blocks loaded with no recompute at depths 0 to 50M; load 2.8 to 4.3x faster than recompute and 8.8 to 12.3x less GPU energy; GPU memory stayed flat. The 12B model answered facts planted millions of tokens earlier at 82% (per abstract).

Key insight: Persisting KV-cache blocks to encrypted local NVMe gives a 50M-token reusable window that loads 2.8–4.3× faster than recompute with flat GPU memory.

Memory Type Varies: Empowering LLM Agents for Long-Term Memory with Diverse Strategies

Yi Wen; Derong Xu; Pengyue Jia et al. arXiv: 2610.11573

Figure from Memory Type Varies: Empowering LLM Agents for Long-Term Memory with Diverse Strategies
Memory Type Varies: Empowering LLM Agents for Long-Term Memory with Diverse Strategies

Retrieval-based memory usually uses one strategy for all memories. Proposes TriMEM, a multi-class memory dataset with precise type annotations, and MemoType, which adaptively recognizes each memory and query and selects type-appropriate strategies. Addresses topic-rich, scenario-complex, boundary-blurred memory scenarios where precise classification is hard.

Key insight: Typed memories need type-appropriate retrieval strategies; a single strategy over a multi-class corpus has a fundamental precision ceiling.

DeltaReplay: Task-Relative Memory Reuse for Mobile GUI Agents

Yudong Bai; Yihong Chen; Quanming Yao et al. arXiv: 2610.11707

Figure from DeltaReplay: Task-Relative Memory Reuse for Mobile GUI Agents
DeltaReplay: Task-Relative Memory Reuse for Mobile GUI Agents

Memory-augmented mobile GUI agents store successful trajectories, but a stored trajectory rarely matches a new task exactly (different parameters, partial overlap, or no relevant record). Forcing irrelevant memory misleads; discarding useful memory wastes experience. DeltaReplay is a step-level reuse framework that decides how to use existing memory without modifying it, via page-level consistency and action-level generality over a transition-graph of trajectories.

Key insight: Mobile GUI agents should reuse memory at the step level relative to the new task, not force whole-trajectory replay of partially matching records.

Use and Disuse: Intent-Structured Experience Consolidation for Memory and Learning in LLM Agents

Xiangyi Zeng; Baihang Liu; Xutong Wang et al. arXiv: 2610.12124

Figure from Use and Disuse: Intent-Structured Experience Consolidation for Memory and Learning in LLM Agents
Use and Disuse: Intent-Structured Experience Consolidation for Memory and Learning in LLM Agents

Hippocam is a hierarchical memory and continual-learning architecture inspired by goal-selective maintenance and gradual consolidation. Structures ongoing work as nested intents: active context stays on the current intent; completed intents consolidate into task-relevant outcomes and state rather than full working details. A recursive prefix consolidation mechanism repeatedly consolidates experience into reusable knowledge for long-term autonomous operation.

Key insight: Structuring work as nested intents and consolidating completed ones into outcomes (not full working details) turns continuous experience into reusable knowledge without weight updates.

MemoWM: How World Models Change What Agents Need to Remember

Bingfan Zeng; Zhisheng Chen; Chenbo Sang et al. arXiv: 2610.10778

Long-term agents face growing storage as they accumulate experience. World models capture reusable regularities that can cut per-experience storage. MemoWM allocates memory conditioned on a world model: shared predictions compress retained info and reconstruct omitted content; a task-aware rule balances reconstruction-error impact vs storage cost. Across five long-term agent-memory benchmarks: 42.42% average answer accuracy (+2.62 pp over strongest baseline) while cutting average experience-specific storage 53.9% vs MIRIX.

Key insight: Conditioning memory on a world model lets agents store only residuals the prior cannot predict, cutting experience-specific storage 53.9% while gaining accuracy.

REMORY: Learning Residual Memory for Context Compaction

Hanchen Xia; Baoyou Chen; Yutang Ge et al. arXiv: 2610.11287

Long-horizon agents compact history to fit a finite window, but a textual summary alone may not support every later decision. REMORY supplements the summary with a bounded sequence of soft memory tokens that help a frozen LLM approximate the continuation it would produce with full history — an analogue of a residual connection along the sequence. On SummHay, improves source attribution at nearly unchanged insight coverage using only 5.2% of input positions; gains on Qwen3.8-27B and GLM-5.3-Flash agent benchmarks.

Key insight: Supplementing a textual history summary with a small residual of soft memory tokens recovers continuation quality at a few percent of input positions.

TaReD: Tool-Aware Recursive Decomposition for Long-Horizon Tasks

Wei-Xiang Mao; Zhi-Kai Chen; De-Chuan Zhan et al. arXiv: 2610.11268

Single-chain interleaved reasoning and actions becomes unreliable on complex tasks as histories grow and early planning errors propagate. Recursive decomposition helps only if each subtask matches available tools; full tool libraries are too large to expose. TaReD organizes tools by functional relationships and does tool-aware recursive decomposition so subtasks align with what tools can execute without injecting every tool description.

Key insight: Recursive decomposition only helps when each subtask matches available tools; organizing tools by function keeps catalogs out of the prompt.

Safe, Persistent, and Evolving Agent Harness for Understanding Partially Observable Worlds

Yisen Gao; Yue Guo; Qing Zong et al. arXiv: 2610.11552

Figure from Safe, Persistent, and Evolving Agent Harness for Understanding Partially Observable Worlds
Safe, Persistent, and Evolving Agent Harness for Understanding Partially Observable Worlds

Enterprise agents need policy compliance, handling of hidden side effects under partial observability, and persistent state across records. E-Ledger: multi-agent harness with a code approval layer checking every proposed action against policy before execution, plus a world ledger of verified hidden rules and evidence-backed dynamic state. WorldAbduct evolves the harness by diagnosing trajectories across complementary views to discover hidden rules a priori unknown.

Key insight: Enterprise agents need a code-approval layer against policy before every action, plus an evidence-backed world ledger of hidden environment rules.

Trajectory-Guided Fault Localization for Agent Skill Evolution

Yu Ge; Linna Xie; Zhong Li et al. arXiv: 2610.11858

Incomplete skills impair code agents; LLM revisions from execution feedback are hard to ground in behavioral evidence. SkillMorph links execution evidence to specific skill contents before revising: compares failure vs success evidence in abstracted trajectories across repeated runs and tasks, localizes suspicious actions to edit sites in skills, then generates revisions.

Key insight: Skill revisions should be grounded in failure-vs-success trajectory evidence that localizes which skill lines to edit.

Code Understanding is a Bottleneck for Coding Agents

Nishant Balepur; Kiran Tomlinson; Tobias Schnabel arXiv: 2610.10610

SWE-bench-style datasets poorly control which abilities drive errors. CABRA builds tasks from scratch as call-graph transformations scaling difficulty on four axes: function traversal, search, runtime resolution, instruction following. 6,840 tasks across 8 LLMs and 6 coding agents: LLM accuracy falls as task size grows, but agents stay near-perfect by offloading to tools (e.g. grep); larger tasks elicit more read/analysis tool calls, and on SWE-bench Verified those tool-call counts predict accuracy better than lines of code edited.

Key insight: Coding-agent accuracy on code-understanding tasks tracks tool-call counts for reading better than lines of code edited.

TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator Evolution

Shuangjie Yao; Hao Wang; Koushik Sen et al. arXiv: 2610.10619

Nearly all coding benchmarks still use passes-fixed-unit-tests — tests often check only part of the requirement, so agents reward-hack or miss required behavior while passing. TestJack generates tests per trial targeting prompt requirements and observed failure modes (evaluator evolution), rather than one-shot static augmentation before any trial is seen.

Key insight: Static unit tests miss requirement gaps; evolving tests against observed agent failure modes exposes reward hacking that fixed suites miss.

Runnable Commit Untangling for Coding Agents

Jinfeng Jiang; Dongsun Kim; Dayi Lin et al. arXiv: 2610.11593

Coding agents produce large tangled patches. Existing untangling ignores that untangled commits must be ordered and leave code runnable. RucTangle is the first agentic method that untangles while keeping code runnable after each commit; TangleEval measures whether untangled histories help agents repair bugs. On 131 agent-generated patches all RucTangle histories are runnable vs 20.6-37.4% unrunnable for baselines; on 453 regression-introducing patches, RucTangle histories give +5.2% absolute pass@1 for repair agents.

Key insight: Untangling agent patches into ordered, runnable commits improves both human review and later agent repair of regression-introducing histories.

SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction

Hexuan Deng; Yue Wang; Wenyu Jiang et al. arXiv: 2610.11559

Claude Code / Codex-style assistants need long chains of work in evolving repos plus multi-turn clarification. SWE-Journey: weak-to-strong synthesis of long-horizon coding tasks, plus four user personas mined from real interaction data driving a user-simulation agent. Addresses both the task-horizon gap and the interaction gap relative to single-issue benchmarks.

Key insight: Coding assistants need evaluation on multi-turn, long-horizon journeys with simulated user personas, not only single-issue SWE-bench patches.

Evaluating Local Language Model Agents for Reproducible Data Engineering: An Empirical Software Engineering Study of Mobility Workflows

Jorge García-Carrasco; Javier Sanchis; Alejandro Reina-Reina et al. arXiv: 2610.11482

Evaluates whether locally deployable open-weight agents can produce correct, reproducible data-engineering artifacts. Benchmark of 15 mobility-workflow tasks (discovery, connectors, feeds, enrichment, features, validation, viz, reporting) with deterministic checkers on scripts, tables, files, and figures. Quantifies closed-loop workspace effects and trade-offs in model scale, architecture, quantization, runtime, tool use, and failure modes.

Key insight: Local open-weight agents for data engineering should be judged by deterministic artifact checkers on scripts and tables, not chat text.

AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks

Xing Han Lù; Dheeraj Vattikonda; Sina Hajimiri et al. arXiv: 2610.11050

Figure from AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks
AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks

Automatic judges for computer-use agents may call a long screenshot-and-action trajectory a success while it violates instruction constraints or causes side effects. AgentHorizon: 1,373 computer-use tasks from 166 hours of human-recorded trajectories across three OSes; constructs negatives by swapping instructions across related trajectories so judges must check the trajectory against the instruction.

Key insight: Automatic judges for computer-use trajectories need instruction-swapped negatives so they check the trajectory against the instruction, not just apparent success.

Error-Propagation Modeling for Failure Attribution in LLM-Based Multi-Agent Systems

Jiaqi Liao; Yuanzhao Zhai; Huanxi Liu et al. arXiv: 2610.11600

In multi-agent systems the observed failure often is not the decisive error. Defines the attribution target as the agent-step pair whose correction would recover the failed execution. EMFA builds a structured representation of the failed trajectory modeling how errors propagate and persist in unresolved loops, to separate decisive errors from downstream symptoms.

Key insight: Multi-agent failure attribution should target the agent-step whose correction would recover the run, not the last noisy symptom.

ReCast: Attribution-Oriented Step Representation Learning for LLM-Based Agent Systems

Weilin Jin; Mingyu Wang; Taiyu Zhu et al. arXiv: 2610.11334

Failure attribution needs the earliest responsible step; hidden-state step representations from frozen LLMs show limited separation of root-cause steps. ReCast learns attribution-oriented step representations: selects relevant layers, builds complementary pattern and deviation features, transforming frozen-LLM hidden states for better root-cause identification.

Key insight: Root-cause step identification needs attribution-oriented encodings, not raw frozen-LLM hidden states.

AgentEvolver: System-Wide Self-Evolution Through Task Execution

Wentao Zhang; Fuchao Yang; Yilei Zhao et al. arXiv: 2610.11613

Completing a task does not improve how the agent works. AgentEvolver develops capabilities during task execution with a fixed foundation model: eight entity families (operations, methods, agents, control flow, interfaces, supporting state) share a versioned lifecycle; Runtime coordinates work with persistent planning and recoverable context. Reports 82.08% SWE-bench Pro Public resolution with evolution vs a lower baseline without; six application cases show retained capabilities reused later.

Key insight: System-wide self-evolution of harness entities under a versioned lifecycle can retain and reuse capabilities across tasks without changing weights.

What Output-Only Review Cannot Verify: Study Contracts for Research Agents

Eitan Waks; Ben Glocker arXiv: 2610.11754

Some defects in AI-generated studies need knowledge of what was approved before execution. Study contracts bind declared experimental choices, run obligations and claim scope to recorded execution evidence. Deterministic checker detected all eight registered mutations in a diagnostic of clean/mutated pairs; LLM judges given packages without pair context flagged 104 of 144 mutated cases (32 abstentions, 8 terminal failures, no explicit clean on mutated).

Key insight: Binding declared experimental choices and claim scope to recorded execution evidence catches mutations that package-only LLM judges miss.

When Should Agents Think? Adaptive Reasoning via Cross-Turn Estimation

Yiruo Cheng; Shen Huang; Xiaoshuai Song et al. arXiv: 2610.12061

Agents usually reason before every action, but earlier reasoning may still support later actions. Finds that drops in likelihood of subsequent reference actions after removing additional reasoning track whether those actions remain recoverable — a lightweight signal for cross-turn action support. Proposes RACE: train adaptive reasoning that skips unnecessary think-steps.

Key insight: Agents need not reason before every action; drops in likelihood of later reference actions after removing reasoning signal when think-steps can be skipped.

Accurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge Conflict

Kaiser Sun; Bernal Jimenez Gutierrez; Hongjun Liu et al. arXiv: 2610.12360

When retrieved evidence contradicts an agent's priors, does it revise, acknowledge uncertainty, or persist wrongly? Evaluates epistemic humility via Identify, Solve, Escalate (ISE) under controlled and naturally occurring knowledge conflicts during multi-step agentic execution (also on HF Daily Papers). Focuses on willingness to recognize, act on, and communicate uncertainty — not only task success.

Key insight: Under knowledge conflict, measure whether agents escalate uncertainty — not only whether the final answer is right.

BRANCH: Bypassing Multi-Scanner AI Guardrails

William Hackett; Peter Garraghan arXiv: 2610.10742

Multi-scanner guardrail systems share latent representations across classification boundaries, defeating single-scanner bypasses. BRANCH uses branching tree search applying adversarial perturbations against individual scanners with optimization based on overall improvement — designed specifically for multi-scanner stacks.

Key insight: Multi-scanner guardrail stacks share latent boundaries; red-team the ensemble with branching search, not each scanner alone.

Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized Deception

Oskar J. Hollinsworth; Alex F. Spies; Tigist Diriba et al. arXiv: 2610.12445

White-box deception probes trained on the largest deception dataset to date, with a probe architecture aggregating across layers and tokens. 98.8% AUC on SHADE-Arena, beating an Opus 5.5 text-monitoring baseline; efficacy improves as the underlying model scales. Also tests introspective deception where ground truth needs elicitation or training-data knowledge.

Key insight: White-box deception probes aggregating across layers and tokens beat text monitors on SHADE-Arena as the underlying model scales.

Plan-and-Patch: Diffusion Language Models for Agentic Planning

Syamantak Kumar; Jiang Guo; Hassan Hamad et al. arXiv: 2610.10786

Long-horizon agents must revise plans when environment or tools invalidate assumptions; revisions often affect only a region. Plan-and-Patch uses a diffusion LM to generate a structured program-like plan via parallel unmasking and repair selected regions while keeping prefix and suffix fixed — avoiding full-plan regeneration.

Key insight: When tools invalidate one plan region, patch that region with diffusion unmasking rather than regenerating the whole plan.