Friday's cs.AI announcement day (2026-10-09) lists 125 new and 177 cross-lists (replacements skipped; listing total 302). A filter for agent systems, memory/context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent, persistence/identity, and local/open models keeps 40 papers. Memory and context dominate: agent-controlled forgetting archives bulky tool results out of the window, always-loaded AGENTS.md files are proved to be capacitated assortments where appending everything useful is wrong, and admission/presentation governance plus formation-time gated memory decide what durable facts enter a turn. Harnesses and skills get hard edges: self-evolution should stay on the harness (prompts/tools/composition) with a recorded gate, process vs content failures choose the lever, co-installed similar skills can drop intended constraints while tasks still pass, and SkillContrast / NOMOS improve skill ranking and policy-to-tool gates. On monitors, time and safety, OnTrack matches steps to past successful runs in about a millisecond, prompt-only wall-clock budgets fail without harness clocks, verifiers must be selected against held-out anchors, and obligation / option-channel / structure-tax results tighten what runtime guards and JSON schemas can be trusted to do.
Jan-Peter Franke arXiv: 2610.10590
Tool-using agents carry observations whose useful content is often much smaller than the payload. Agent-controlled forgetting lets the acting model select prior tool results, replace each with a short note at its original position, and keep the exact original in a recoverable archive. A Python harness exposes batch archival and explicit recovery without task-specific training, while protecting user instructions and assistant messages. In an exploratory OpenTelemetry debugging case followed by an unrelated implementation task: 231,951 provider prompt tokens vs 912,492 under retained history (50% fewer cumulative input tokens; about USD 1.28-1.44 vs about 4.38). Both arms passed the two-case primary behavioral oracle.
Key insight: Reversible forgetting of bulky tool results — short notes in context, originals in an archive — cut prompt tokens by about 50% on a noisy debug-and-implement run.
Zexuan Liu; Yuning Yang; Tiancheng Zhao arXiv: 2610.11007
At session start LLM agents load a fixed context file (e.g. AGENTS.md); every loaded token is charged again each later round, and growing files can degrade performance, yet curators usually append. Formulates curation as a capacitated assortment problem: instructions consume tokens under finite attention; adding one never raises compliance of others; retained instructions incur per-session setup cost. Proves an upper bound on optimal file size regardless of candidate count, and that appending every positive-standalone-value instruction can be arbitrarily worse than an optimal subset.
Key insight: Always-loaded context files are capacitated assortments: appending every positive-value instruction can be arbitrarily worse than choosing an optimal subset under a token budget.
Chang Liu; Deliang Ding arXiv: 2610.11188
Persistent memory improves personalization but can induce sycophancy and cross-domain leakage. Separates two governance decisions: admission (what recalled info enters working context) and presentation (how admitted info is expressed). Two inference-time designs without retraining: factor-compiled admission (FC) assessing whole entries, and permission-semantic admission (PS) decomposing entries into typed units; both use deterministic policies. On an external benchmark (four tasks times 300 samples), FC and PS cut pooled judge-assessed failure rates vs verbatim injection by 6.7 and 8.8 pp.
Key insight: Governing persistent memory means two decisions — what is admitted into the turn and how it is phrased — and admission quality alone cut judge-assessed failures by 6.7–8.8 points.
Preeti Saraswat; Divya Neelagiri; Ajay Manoj arXiv: 2610.11270
Personalized conversational AI extracts facts into persistent vector stores, but the formation stage (first write) has almost no principled attention. Critical signals (permanent attribute vs transient situation) exist only in the original utterance and are lost once extraction produces a subject-relation-object triple. Gated Memory interposes two checkpoints between conversation and storage: an admission gate evaluating every candidate fact, so formation-time context is not thrown away.
Key insight: Formation-time admission is the binding constraint on memory quality: permanent vs transient context is lost the moment extraction produces a bare triple.
Ravil Akhtyamov arXiv: 2610.10629
Argues self-evolution is reviewable only if confined to the runtime harness (instruction text, tool-call logic, primitive composition) while model weights stay fixed — every adaptation is a diff with a cause and a test. Dual-loop engine with one admission gate writing a hash-chained record before deployment. Measured in simulation (simulated agent + seeded-search proposer, not LLMs): across three families of supervisory re-interpretation at three severities times 10 seeds, the gated loop admitted 144 of 7,449 candidates, none of which worsened error on held-out history.
Key insight: Self-evolution is reviewable only when confined to the harness — prompts, tools, and composition — with every change gated by a recorded test before deployment.
Yuan Tian; Bing Hu; Hao Wang et al. arXiv: 2610.11655
Improving a long-horizon agent means evolving the harness around a frozen model or training weights. Labels failed trajectories by the first signal that fires into process failures (blocked calls, loops, exhausted budgets) vs content failures (delivered but poor plan). Harness evolution repairs process failures (behavior can later be trained into weights); content failures need weight training. On DeepPlanning, self-evolving harness lifts Qwen3.5-4B from 0.16 to 0.30 and 9B from 0.32 to 0.44 held-out.
Key insight: Process failures (loops, blocked tools) are harness problems; content failures need weight training — diagnose which before choosing the lever.
Chaoliang Yan; Zihao Xu; Yuekang Li et al. arXiv: 2610.11647
Co-installed similar skills (from independent sources) can conflict: the model picks by name and description alone, and the intended skill loses core functions (e.g. a ban on touching git) because a similar skill runs instead — yet the task still passes, so completion benchmarks miss it. From 20,947 repos: 822,109 candidate similar-skill pairs; LLM-judged sample of 3,754; 312 confirmed pairs on three models (6,368 runs, 169,294 tool calls, 542 agent-hours).
Key insight: Co-installed similar skills can drop the intended skill's constraints while the task still passes, so completion benchmarks miss the conflict.
Jiandong Ding; Honglei Ji; Ming Liu et al. arXiv: 2610.11650
Similar agent skills share instructions but differ in conditions of use; query-based text selection keeps shared text and drops distinctions. SkillContrast is training-free: compares retrieved skills and retains differing text with local context for a pretrained reranker. On 1,235 SameCapRisk-Bench requests: 54 to 72 more clean hits (helpful skill without risky sibling) than TF-IDF at matched lengths; uses 51.1 to 58.8% fewer model-input tokens than full skill bodies.
Key insight: Reranking skills by the text that differs between candidates, not shared boilerplate, yields more clean hits at far fewer input tokens.
Min-Young Yu; Tony Kim; Jang Won Choi arXiv: 2610.11030
Tool-using agents silently violate deployed policies. NOMOS is a four-pass compiler from natural-language policy to a deterministic tool-call gate; static verification with tool-schema-level checks alone (no prover, solver, or LLM) repairs or rejects 37% (airline) and 13% (retail) of candidates — without which most shipped rules are inoperable. On tau2-bench the gate cuts violations of reference-encoded clauses; a development binding that refused 95.9% of task-passing calls is flagged as over-blocking.
Key insight: Compiling natural-language policy into schema-checked tool-call gates catches inoperable rules before they ship, but over-blocking gates must be flagged.
Babak Barazandeh; Connor Swanson; Chinmay Kulkarni et al. arXiv: 2610.12375
Safeguard agents per-step add cost and latency; post-hoc log review comes after damage. OnTrack streams monitoring by comparing an agent's steps and dependencies against recorded successful runs, alerting or blocking in about 1 ms per step. Studied under full reference access (historical runs and tool schemas), schemas-only, and step-logs-only.
Key insight: Streaming monitors that match agent steps against past successful runs can alert or block in about a millisecond — cheaper than a second LLM safeguard.
Aaron Wang; Neelabh Madan; Vlad Sobal et al. arXiv: 2610.10833
Do small LLM agents respect and productively use wall-clock budgets? Prompt-only budgets fail: harness gives no timing feedback, agents cannot anticipate action duration, and lack a learned mapping from time to strategy. Studies harness interventions that expose timing and enforce deadlines, plus agent-side adaptations, on MLE-Bench Lite (Qwen3.6-27B) and Zork I (Qwen3-4B).
Key insight: Prompt-only wall-clock budgets fail; timing feedback and hard deadlines must live in the harness.
Xing Zhang; Guanghui Wang; Yanwei Cui et al. arXiv: 2610.11464
Self-improving loops need a verifier; on open-ended tasks a hand rubric or bare LLM judge invites reward hacking. Makes the verifier the evolving object: inspectable expression over mostly-deterministic drawback detectors, synthesized from clustered failures, gated at birth, selected for agreement with a ten-item anchored reference set plus consensus — never for the agent's score. On MBPP+ gains +0.21 held-out agreement over seed composition on every seed; removing anchor guards collapses the result.
Key insight: Self-improve loops need an inspectable verifier selected against held-out anchors — never against the agent's own score.
Youwei Feng; Yitong Zhang; Yuetong Liu et al. arXiv: 2610.11773
Guard models mostly flag forbidden actions, but safety also requires detecting required-yet-unperformed safety-critical actions (obligations). On a popular safety benchmark, 56.92% of GLM-5.3 trajectories contain unfulfilled obligations vs 30.00% forbidden actions. Introduces ObligationBench to evaluate whether guards can identify obligations.
Key insight: Safety guards must detect required-yet-unperformed obligations; on one benchmark unfulfilled obligations outnumbered forbidden actions.
Seyedarmin Azizi; Erfan Baghaei Potraghloo; Massoud Pedram arXiv: 2610.12292
Typed decision models (probability over caller-defined options with written definitions, no free text) are used as agent guardrails. Seven open-weight models: accuracy 36-72% vs 50% chance on injection/jailbreak/toxic screening; low error in one direction often means a default-allow or default-block bias. On synthetic agent tool calls, option-name polarity manipulations (yes/no names etc.) raise decision-flip rates up to 70.4 pp vs 0/1 controls while type-error stays 0%.
Key insight: Typed yes/no decision guardrails flip under option-name polarity alone, with zero type errors — prefer schema gates over soft classifiers.
Vineet Kumar; Kanishka; Bhuvanesh Mandora arXiv: 2610.12056
Re-examines the claim that structured JSON/XML outputs inherently cost accuracy. Across models, datasets, and schemas: the tax depends on schema design — reasoning-first field ordering matches or exceeds free-form accuracy; answer-first ordering causes steep drops especially in smaller models. Schemas preserving reasoning order also improve calibration.
Key insight: The structured-output accuracy tax depends on field order: reasoning-first schemas match free-form accuracy; answer-first schemas hurt, especially on small models.
Bowen Xu; Boyu Chen arXiv: 2610.10961
Describes an advisory cross-provider review contract: distinct resource pools, bounded execution, restricted reviewer capabilities, complete input delivery, usable semantic output, explicit failure states, durable per-attempt evidence. Controlled pilot of 20 paired development turns: 8 had a material reviewer finding (95% CI 19.1-63.9%). Boundary scan reproduced a false success on partial input (four truncation levels passed historically, failed after repair).
Key insight: An advisory cross-provider review contract needs complete input delivery — truncated reviewer input produced false successes in a controlled pilot.
Sietse Schelpe arXiv: 2610.10845
Tests galahad-kv, a public package that saves KV state of about 16k-token blocks to encrypted local NVMe and loads it back byte-exact without recompute. On 50M tokens of public text via vLLM on one H100 with Gemma 4 12B and 31B: 100/100 probed blocks loaded with no recompute at depths 0 to 50M; load 2.8 to 4.3x faster than recompute and 8.8 to 12.3x less GPU energy; GPU memory stayed flat. The 12B model answered facts planted millions of tokens earlier at 82% (per abstract).
Key insight: Persisting KV-cache blocks to encrypted local NVMe gives a 50M-token reusable window that loads 2.8–4.3× faster than recompute with flat GPU memory.
Yi Wen; Derong Xu; Pengyue Jia et al. arXiv: 2610.11573
Retrieval-based memory usually uses one strategy for all memories. Proposes TriMEM, a multi-class memory dataset with precise type annotations, and MemoType, which adaptively recognizes each memory and query and selects type-appropriate strategies. Addresses topic-rich, scenario-complex, boundary-blurred memory scenarios where precise classification is hard.
Key insight: Typed memories need type-appropriate retrieval strategies; a single strategy over a multi-class corpus has a fundamental precision ceiling.
Yudong Bai; Yihong Chen; Quanming Yao et al. arXiv: 2610.11707
Memory-augmented mobile GUI agents store successful trajectories, but a stored trajectory rarely matches a new task exactly (different parameters, partial overlap, or no relevant record). Forcing irrelevant memory misleads; discarding useful memory wastes experience. DeltaReplay is a step-level reuse framework that decides how to use existing memory without modifying it, via page-level consistency and action-level generality over a transition-graph of trajectories.
Key insight: Mobile GUI agents should reuse memory at the step level relative to the new task, not force whole-trajectory replay of partially matching records.
Xiangyi Zeng; Baihang Liu; Xutong Wang et al. arXiv: 2610.12124
Hippocam is a hierarchical memory and continual-learning architecture inspired by goal-selective maintenance and gradual consolidation. Structures ongoing work as nested intents: active context stays on the current intent; completed intents consolidate into task-relevant outcomes and state rather than full working details. A recursive prefix consolidation mechanism repeatedly consolidates experience into reusable knowledge for long-term autonomous operation.
Key insight: Structuring work as nested intents and consolidating completed ones into outcomes (not full working details) turns continuous experience into reusable knowledge without weight updates.
Bingfan Zeng; Zhisheng Chen; Chenbo Sang et al. arXiv: 2610.10778
Long-term agents face growing storage as they accumulate experience. World models capture reusable regularities that can cut per-experience storage. MemoWM allocates memory conditioned on a world model: shared predictions compress retained info and reconstruct omitted content; a task-aware rule balances reconstruction-error impact vs storage cost. Across five long-term agent-memory benchmarks: 42.42% average answer accuracy (+2.62 pp over strongest baseline) while cutting average experience-specific storage 53.9% vs MIRIX.
Key insight: Conditioning memory on a world model lets agents store only residuals the prior cannot predict, cutting experience-specific storage 53.9% while gaining accuracy.
Hanchen Xia; Baoyou Chen; Yutang Ge et al. arXiv: 2610.11287
Long-horizon agents compact history to fit a finite window, but a textual summary alone may not support every later decision. REMORY supplements the summary with a bounded sequence of soft memory tokens that help a frozen LLM approximate the continuation it would produce with full history — an analogue of a residual connection along the sequence. On SummHay, improves source attribution at nearly unchanged insight coverage using only 5.2% of input positions; gains on Qwen3.8-27B and GLM-5.3-Flash agent benchmarks.
Key insight: Supplementing a textual history summary with a small residual of soft memory tokens recovers continuation quality at a few percent of input positions.
Wei-Xiang Mao; Zhi-Kai Chen; De-Chuan Zhan et al. arXiv: 2610.11268
Single-chain interleaved reasoning and actions becomes unreliable on complex tasks as histories grow and early planning errors propagate. Recursive decomposition helps only if each subtask matches available tools; full tool libraries are too large to expose. TaReD organizes tools by functional relationships and does tool-aware recursive decomposition so subtasks align with what tools can execute without injecting every tool description.
Key insight: Recursive decomposition only helps when each subtask matches available tools; organizing tools by function keeps catalogs out of the prompt.
Yisen Gao; Yue Guo; Qing Zong et al. arXiv: 2610.11552
Enterprise agents need policy compliance, handling of hidden side effects under partial observability, and persistent state across records. E-Ledger: multi-agent harness with a code approval layer checking every proposed action against policy before execution, plus a world ledger of verified hidden rules and evidence-backed dynamic state. WorldAbduct evolves the harness by diagnosing trajectories across complementary views to discover hidden rules a priori unknown.
Key insight: Enterprise agents need a code-approval layer against policy before every action, plus an evidence-backed world ledger of hidden environment rules.
Yu Ge; Linna Xie; Zhong Li et al. arXiv: 2610.11858
Incomplete skills impair code agents; LLM revisions from execution feedback are hard to ground in behavioral evidence. SkillMorph links execution evidence to specific skill contents before revising: compares failure vs success evidence in abstracted trajectories across repeated runs and tasks, localizes suspicious actions to edit sites in skills, then generates revisions.
Key insight: Skill revisions should be grounded in failure-vs-success trajectory evidence that localizes which skill lines to edit.
Nishant Balepur; Kiran Tomlinson; Tobias Schnabel arXiv: 2610.10610
SWE-bench-style datasets poorly control which abilities drive errors. CABRA builds tasks from scratch as call-graph transformations scaling difficulty on four axes: function traversal, search, runtime resolution, instruction following. 6,840 tasks across 8 LLMs and 6 coding agents: LLM accuracy falls as task size grows, but agents stay near-perfect by offloading to tools (e.g. grep); larger tasks elicit more read/analysis tool calls, and on SWE-bench Verified those tool-call counts predict accuracy better than lines of code edited.
Key insight: Coding-agent accuracy on code-understanding tasks tracks tool-call counts for reading better than lines of code edited.
Shuangjie Yao; Hao Wang; Koushik Sen et al. arXiv: 2610.10619
Nearly all coding benchmarks still use passes-fixed-unit-tests — tests often check only part of the requirement, so agents reward-hack or miss required behavior while passing. TestJack generates tests per trial targeting prompt requirements and observed failure modes (evaluator evolution), rather than one-shot static augmentation before any trial is seen.
Key insight: Static unit tests miss requirement gaps; evolving tests against observed agent failure modes exposes reward hacking that fixed suites miss.
Jinfeng Jiang; Dongsun Kim; Dayi Lin et al. arXiv: 2610.11593
Coding agents produce large tangled patches. Existing untangling ignores that untangled commits must be ordered and leave code runnable. RucTangle is the first agentic method that untangles while keeping code runnable after each commit; TangleEval measures whether untangled histories help agents repair bugs. On 131 agent-generated patches all RucTangle histories are runnable vs 20.6-37.4% unrunnable for baselines; on 453 regression-introducing patches, RucTangle histories give +5.2% absolute pass@1 for repair agents.
Key insight: Untangling agent patches into ordered, runnable commits improves both human review and later agent repair of regression-introducing histories.
Hexuan Deng; Yue Wang; Wenyu Jiang et al. arXiv: 2610.11559
Claude Code / Codex-style assistants need long chains of work in evolving repos plus multi-turn clarification. SWE-Journey: weak-to-strong synthesis of long-horizon coding tasks, plus four user personas mined from real interaction data driving a user-simulation agent. Addresses both the task-horizon gap and the interaction gap relative to single-issue benchmarks.
Key insight: Coding assistants need evaluation on multi-turn, long-horizon journeys with simulated user personas, not only single-issue SWE-bench patches.
Jorge García-Carrasco; Javier Sanchis; Alejandro Reina-Reina et al. arXiv: 2610.11482
Evaluates whether locally deployable open-weight agents can produce correct, reproducible data-engineering artifacts. Benchmark of 15 mobility-workflow tasks (discovery, connectors, feeds, enrichment, features, validation, viz, reporting) with deterministic checkers on scripts, tables, files, and figures. Quantifies closed-loop workspace effects and trade-offs in model scale, architecture, quantization, runtime, tool use, and failure modes.
Key insight: Local open-weight agents for data engineering should be judged by deterministic artifact checkers on scripts and tables, not chat text.
Xing Han Lù; Dheeraj Vattikonda; Sina Hajimiri et al. arXiv: 2610.11050
Automatic judges for computer-use agents may call a long screenshot-and-action trajectory a success while it violates instruction constraints or causes side effects. AgentHorizon: 1,373 computer-use tasks from 166 hours of human-recorded trajectories across three OSes; constructs negatives by swapping instructions across related trajectories so judges must check the trajectory against the instruction.
Key insight: Automatic judges for computer-use trajectories need instruction-swapped negatives so they check the trajectory against the instruction, not just apparent success.
Jiaqi Liao; Yuanzhao Zhai; Huanxi Liu et al. arXiv: 2610.11600
In multi-agent systems the observed failure often is not the decisive error. Defines the attribution target as the agent-step pair whose correction would recover the failed execution. EMFA builds a structured representation of the failed trajectory modeling how errors propagate and persist in unresolved loops, to separate decisive errors from downstream symptoms.
Key insight: Multi-agent failure attribution should target the agent-step whose correction would recover the run, not the last noisy symptom.
Weilin Jin; Mingyu Wang; Taiyu Zhu et al. arXiv: 2610.11334
Failure attribution needs the earliest responsible step; hidden-state step representations from frozen LLMs show limited separation of root-cause steps. ReCast learns attribution-oriented step representations: selects relevant layers, builds complementary pattern and deviation features, transforming frozen-LLM hidden states for better root-cause identification.
Key insight: Root-cause step identification needs attribution-oriented encodings, not raw frozen-LLM hidden states.
Wentao Zhang; Fuchao Yang; Yilei Zhao et al. arXiv: 2610.11613
Completing a task does not improve how the agent works. AgentEvolver develops capabilities during task execution with a fixed foundation model: eight entity families (operations, methods, agents, control flow, interfaces, supporting state) share a versioned lifecycle; Runtime coordinates work with persistent planning and recoverable context. Reports 82.08% SWE-bench Pro Public resolution with evolution vs a lower baseline without; six application cases show retained capabilities reused later.
Key insight: System-wide self-evolution of harness entities under a versioned lifecycle can retain and reuse capabilities across tasks without changing weights.
Eitan Waks; Ben Glocker arXiv: 2610.11754
Some defects in AI-generated studies need knowledge of what was approved before execution. Study contracts bind declared experimental choices, run obligations and claim scope to recorded execution evidence. Deterministic checker detected all eight registered mutations in a diagnostic of clean/mutated pairs; LLM judges given packages without pair context flagged 104 of 144 mutated cases (32 abstentions, 8 terminal failures, no explicit clean on mutated).
Key insight: Binding declared experimental choices and claim scope to recorded execution evidence catches mutations that package-only LLM judges miss.
Yiruo Cheng; Shen Huang; Xiaoshuai Song et al. arXiv: 2610.12061
Agents usually reason before every action, but earlier reasoning may still support later actions. Finds that drops in likelihood of subsequent reference actions after removing additional reasoning track whether those actions remain recoverable — a lightweight signal for cross-turn action support. Proposes RACE: train adaptive reasoning that skips unnecessary think-steps.
Key insight: Agents need not reason before every action; drops in likelihood of later reference actions after removing reasoning signal when think-steps can be skipped.
Kaiser Sun; Bernal Jimenez Gutierrez; Hongjun Liu et al. arXiv: 2610.12360
When retrieved evidence contradicts an agent's priors, does it revise, acknowledge uncertainty, or persist wrongly? Evaluates epistemic humility via Identify, Solve, Escalate (ISE) under controlled and naturally occurring knowledge conflicts during multi-step agentic execution (also on HF Daily Papers). Focuses on willingness to recognize, act on, and communicate uncertainty — not only task success.
Key insight: Under knowledge conflict, measure whether agents escalate uncertainty — not only whether the final answer is right.
William Hackett; Peter Garraghan arXiv: 2610.10742
Multi-scanner guardrail systems share latent representations across classification boundaries, defeating single-scanner bypasses. BRANCH uses branching tree search applying adversarial perturbations against individual scanners with optimization based on overall improvement — designed specifically for multi-scanner stacks.
Key insight: Multi-scanner guardrail stacks share latent boundaries; red-team the ensemble with branching search, not each scanner alone.
Oskar J. Hollinsworth; Alex F. Spies; Tigist Diriba et al. arXiv: 2610.12445
White-box deception probes trained on the largest deception dataset to date, with a probe architecture aggregating across layers and tokens. 98.8% AUC on SHADE-Arena, beating an Opus 5.5 text-monitoring baseline; efficacy improves as the underlying model scales. Also tests introspective deception where ground truth needs elicitation or training-data knowledge.
Key insight: White-box deception probes aggregating across layers and tokens beat text monitors on SHADE-Arena as the underlying model scales.
Syamantak Kumar; Jiang Guo; Hassan Hamad et al. arXiv: 2610.10786
Long-horizon agents must revise plans when environment or tools invalidate assumptions; revisions often affect only a region. Plan-and-Patch uses a diffusion LM to generate a structured program-like plan via parallel unmasking and repair selected regions while keeping prefix and suffix fixed — avoiding full-plan regeneration.
Key insight: When tools invalidate one plan region, patch that region with diffusion unmasking rather than regenerating the whole plan.