Monday's cs.AI announcement day (2026-10-05) lists 124 new and 142 cross-lists (replacements skipped; listing total 266). A filter for agent systems, memory/context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent, persistence/identity, and local/open models keeps 39 papers. The day's through-line is that the system around the model keeps deciding outcomes: WebFovea lifts a live-web agent from 31.0 to 57.0 with the same model, VERSE finds harness self-evolution works only with execution-based verification, and FinSkillBench shows curated tools and skills — not self-written ones — carry the gains. Memory work shifts toward the write path: DyadMem measures how an agent should work with a specific user, APDMem reads memory coarse-to-fine, and Sentry keeps failure lessons out of context until a failure actually occurs. On the safety side, GHOST, COBRA and Pincer each show a different way long-running or computer-use agents drift past constraints, and propose layered fixes.
Yifei Tao; Xinyu Zhong; Henry Hengyuan Zhao et al. arXiv: 2610.03020
Long-term memory benchmark for User-conditioned Relational Agent Memory (how this agent should work with this user): 3,065 episodes, 50,961 sessions, 61,210 QA with Capture/Update/Recall gold labels. Across 16 open-weight and 4 proprietary models, Gold-Memory QA is strong but Full-Pipeline QA drops sharply; low capture recall and unsafe deletion even in frontier models.
Key insight: Long-term agents need memory of how to work with a specific user, and the capture and update stages — not just recall — are where pipelines fail.
Jiangang Han arXiv: 2610.03036
2nd-place WebRetriever Challenge 2026 vision web agent (57.0/100). Same model across all four submissions; hidden-set score rose 31.0 → 57.0 from harness changes (up to run-to-run variance on live sites). Failures traced to parse → act → report → show stages: a coordinate-space mismatch put every click at 3/4 of intended coordinates, native dropdowns/iframes/text boxes failed silently, chat-template tokens contaminated 4.9% of episodes.
Key insight: On live websites, many agent failures sit in the plumbing between model and page — parsing, action effect, reporting and observation — not in the model's reasoning.
Changxiu Ji; Amy Lu; Qizheng Zhang et al. arXiv: 2610.02994
Failure-management layer beside the agent: on detected failure it retrieves matching lessons from an external playbook, verifies recovery without task rewards, and stores a lesson only if recovery worked; the full playbook never enters context. Beats the strongest runtime-intervention baseline by 37% on average and the strongest context-evolution baseline by 39%.
Key insight: Failure lessons are conditional knowledge: surfacing them only when a failure is detected works better than keeping the whole playbook in context.
A. Said Gurbuz; Ahmed Nassar; Sunghwan Hong et al. arXiv: 2610.02320
Controllable desktop environment composing real apps to produce DeskForge-1M (1.2M annotated observations, 159.7M elements). Fine-tuning Qwen3.5-4B on 200K grounding examples: +11.51 pp ScreenSpot-Pro, +10.11 OSWorld-G; WebArena-Infinity 31 → 50 of 119 tasks, OpenApps 3 → 15 of 100. Code, data and model released.
Key insight: Composing real desktop applications under controlled variation yields dense grounding supervision that transfers to long-horizon computer use.
Chin-Lun Fu; Anagha Kulkarni; Hong Ni et al. arXiv: 2610.02472
Four-layer memory (thematic summaries, personalized key facts, turn-level evidence notes, raw messages) with a controller that drills down only when needed and a note synthesizer that orders events and flags contradictions. Strong LongMemEval performance while accessing only 8% of conversations (EMNLP 2026 Industry).
Key insight: Reading memory coarse-to-fine lets simple queries stop early while hard temporal or multi-hop queries drill into raw evidence.
XinPeng Shen; Lan Zhang; Yixiao Huang et al. arXiv: 2610.02664
Names the failure where a long-horizon agent violates a safety constraint stated many turns earlier under benign conditions: 11.5% occurrence on GPT-5.5. STAR-Guard restores applicable historical constraints and adds a deterministic pre-execution audit; no GHOST events observed under the GPT-5.5 setup.
Key insight: Long interaction histories let agents violate constraints stated many turns earlier, even under benign conditions; restoring and auditing those constraints prevents it.
Giulio Zingrillo; Hanna Foerster; Ilia Shumailov et al. arXiv: 2610.03089
Shows Dual-LLM CUAs are vulnerable to branch steering — untrusted page data pushes the agent down a hazardous pre-approved branch. STEER-Bench (101 tasks): 94.4% ASR on standard and 89.5% on vanilla Dual-LLM CUAs. COBRA pairs branching plans with ahead-of-time capability constraints: 0% ASR, 97% benign utility.
Key insight: Plan-first defenses for computer-use agents leak through branch steering; bounding each branch's parameters and destinations closes the gap.
Mayank Rathee; Alexander Stepanov; Shalin Madabhavi et al. arXiv: 2610.02569
Resource-layer defense for coding agents: an isolated-context digital twin continually learns user-specific least-privilege policies from multi-day interaction and answers the agent's permission requests as the user's proxy, alongside tool-call-layer auto modes. Outperforms LLM-judge baselines and Conseca adaptations.
Key insight: A continually learning digital twin of the user can answer an agent's permission requests with user-specific least privilege.
Zekai Wang; Yingqiang Ge; Zekun Wang et al. arXiv: 2610.02616
Self-evolving harness optimizer that tests draft edits, replays failures and perturbs suspected steps before submitting, and revises both the executor harness and its own prompts, skills, tools, hooks and notes (weights fixed). Best validation-selected harness: 42.3% / 37.7% on held-out SWE-rebench and newer OOD tasks vs 39.2% / 29.3% for the strongest baselines.
Key insight: Letting a harness optimizer rewrite its own tools only pays off when every edit is checked by execution, not by the optimizer's own judgment.
Jermyn Zhen Yong Bek; Zhuang Qiang Bok; Zhongtian Sun arXiv: 2610.03564
2,603 point-in-time finance episodes, 17,820 runs across 9 models. Curated skill packages add +16.2 points (0.366 → 0.528); skills generated within one episode add only +0.5 while costing more tokens/turns. Docs alone +5.6, tools alone +19.5, combined subadditive; a second harness reproduces direction but not magnitude.
Key insight: Curated skills and executable tools carry most of the measured 'skill premium'; skills an agent writes for itself mid-episode add almost nothing.
Moonseok Choi; Taehong Moon; Giung Nam et al. arXiv: 2610.02858
HAD distills a teacher agent into a smaller student while the harness stays in place: action preferences contrast the teacher's actions with vs without harness info, plus a validity check against harness records; no task rewards needed. Beats on-policy distillation with the same fixed harness; fewer unproductive loops, more error recovery.
Key insight: When the harness stays in place, distillation should target what the teacher adds beyond the harness rather than imitating its full outputs.
Zongxia Li; Yucheng Shi; Zhongzhi Li et al. arXiv: 2610.02826
One Qwen-3.8-27B base model solves tasks under diverse harnesses, then a planner/critic/executor rewrites successes into runbook-style trajectories for a general harness: 2,001 source trajectories → 11,094 rewrites. Terminal-Bench 2 pass@3 57.0% → 74.2%; open weights released (IntelligenceLab/RSR-27B).
Key insight: Successes found under specialized harnesses can be rewritten into training trajectories that work under a plain, general harness.
Xi Qin; Isabel Kurth; Xin Cui et al. arXiv: 2610.02405
Diagnoses meta-agent (Claude Opus) generated terminal tasks/verifiers: benchmark invalidity, harness brittleness, reward misalignment. Prompt redesign and context extension raise solvability 5.6×; a 9B model saturates at 81.3% mean pass@2, dropping to 20.6% when hard tasks are added.
Key insight: A runnable Docker image and test suite do not make a faithful training pipeline; solvability bands and verifier audits must be measured explicitly.
Yu Li; Guangfeng Cai; Long-Fei Li et al. arXiv: 2610.03634
Dependency-Aware Group Policy Optimization builds a command dependency graph from terminal traces and traces back from what the verifier inspects, assigning credit to relevant writes and supporting reads. Improves performance and training stability on complex terminal tasks.
Key insight: Tracing read-write dependencies between terminal commands gives sharper credit assignment than trajectory- or step-level rewards.
Jiawei Li arXiv: 2610.02267
Paired evaluation of an open-weight (Laya) and hosted (Jev) single-forward-pass classifier on 11 harness decision points (routing, tool choice, RAG gating, injection). Jev wins 9/11 (+10.8 to +46.0 pp); neither beats chance on zero-shot model routing; Laya flips 30% of answers on option-order reversal and hits 31% at 50 nearest-neighbour tools. Self-audit: a reported 23.9% saving was actually 4.3%.
Key insight: Single-pass decision classifiers for harness gates can be order-sensitive and collapse on large tool catalogs, and cost-saving claims are easy to overstate.
Albert Sadowski; Jarosław A. Chudziak arXiv: 2610.02897
Compares three summarization write policies for a multi-goal assistant: goal-neutral, one all-goal summary, or one summary per goal. Per-goal summaries win on relevance, completeness and accuracy; the all-goal summary loses even to the neutral one written at a fraction of its budget.
Key insight: Summaries written per standing goal beat a single all-purpose summary, because goals disagree about what was worth keeping.
Minji Park; Seunghyun Yoon; Hyuk Lim arXiv: 2610.02736
Benchmark for 'turning-point eviction': compressors keep dialogue facts but drop the turn where the user revised them. At retained fraction 0.30 every compressed method stays below full context; deleting the update turn sharply lowers current-value accuracy; recency is the best compressed method on LongMemEval-KU and RiSAWOZ.
Key insight: Dialogue compressors can keep the facts but drop the turn where the user changed them; overall retention scores hide this.
Yehya Farhat; Michael Desmond; Anastasios Kyrillidis arXiv: 2610.02687
Frames agent memory updates as optimization over the model's context; GraphMemory accumulates and connects reusable strategies and retrieves only the relevant subgraph per query, keeping retrieved memory constant as examples grow. Competitive accuracy with roughly 81-85% fewer memory-construction tokens.
Key insight: Treating memory updates as context optimization and retrieving only a relevant subgraph keeps memory cost bounded as experience grows.
Jingyu Liu; Zhiwen Wang; Yuxin Jing et al. arXiv: 2610.02769
Shows agents barely use action–outcome correspondence in history: much of history's benefit survives shuffling past actions. Explicitly labeling each observation as the outcome of the preceding action improves success and reduces repetition; a learned calibrator that reassesses past actions helps further.
Key insight: Agents often fail to connect past actions with their outcomes; explicitly labeling each observation with its action measurably helps.
Junyi Zhang; Jinxi Yu; Eric Hanchen Jiang et al. arXiv: 2610.02945
Math research agent built on a cross-problem graph memory of facts, plans and counterexamples with dependency-aware retrieval, an evidence-sensitive curator and scoped recall of negative findings. Reports closure on all ten First Proof Second Batch problems and solutions to the Jamison caterpillar conjecture and Erdős Problems 289, 348 and 488.
Key insight: A typed graph memory of facts, plans and counterexamples — including negative results — lets parallel research agents sustain and resume long searches.
Ankur Samanta; Yonathan Efroni; Paul Sajda et al. arXiv: 2610.02525
Hierarchical research agent: an outer meta-reasoner curates context from a persistent research record and writes work orders; a fresh inner executor runs each. A generative critic forecasts remaining return at decision boundaries; MIRA-AC trains on meta-decisions only and improves gold performance across four autoresearch environments.
Key insight: Separating what to investigate next from how to execute it turns long-horizon research steering into a learnable policy.
Johannes Wesch; Danni Liu; Jan Niehues arXiv: 2610.03198
Query-agnostic KV-cache compression for a prefilled context reused across many queries: a proxy scorer picks informative tokens, then only that subset is reprocessed for eviction scores. On RULER 16K at a 2% budget it beats the next-best baseline by more than 40 points.
Key insight: A reusable prefilled context can be compressed accurately without reprocessing the whole prompt, by rescoring only a proxy-selected subset.
Yulong Ming; Jie Xu; Zihan Wu et al. arXiv: 2610.02932
Payback-Aware Compilation from Experience: measures costs of compiling repeated GUI procedures into programs (including failed attempts) and decides online when to compile. Payback of 2-16 uses for successful compilations; 17.3% token reduction vs ReAct and 24.9% vs an AutoRPA adaptation.
Key insight: Compiling repeated GUI procedures into programs is worth it only when expected reuse beats compile cost, including failed attempts.
Alham Fikri Aji; Faiz Rizki Ramadhan; Zayd M. K. Zuhri et al. arXiv: 2610.03574
423 hard human-validated browsing questions across 13 languages requiring obscure evidence in videos, scans, images or maps; questions answerable without internet are filtered out. Evaluates provider-native search and a shared retrieval harness under a common agent protocol, with a human baseline.
Key insight: Multilingual, multimodal web research remains far from solved when evidence is obscure and spread across videos, scans and maps.
Ajay Vohra; Tao Chen; Neeti Narayan et al. arXiv: 2610.02351
Splits ReAct into a Brain plus an action-validating Critic and a Context Manager that reconstructs environment-supported state and certifies completion. On GAIA and SWE-bench Verified: +6.5-7.0 Pass@1 for Qwen3-Coder-480B and +4.2-5.2 for Claude Sonnet 4.5; comparable with Claude Opus 4.5 but more evidence-complete trajectories.
Key insight: Externalizing action validation and completion certification from the main policy helps weaker models most and improves grounding for strong ones.
Yu Li; Zheng Zhang; Xin Liu et al. arXiv: 2610.02330
Trains a Comparative Inference Model to estimate how likely a candidate next tool call is to support final success, using observed tool behavior, a Bayesian tool-graph simulator and LLM comparisons. Improves Tool F1 and task success across three tool-use benchmarks (NeurIPS 2026).
Key insight: Tool-use agents can learn to estimate the long-horizon value of a candidate tool call before executing it.
Chiara Troiani; Arash Salarian; Majed El Helou et al. arXiv: 2610.03213
SLM classifier that checks each selected tool call against the task intent for low-latency, on-prem per-call oversight; dataset of multi-tool tasks spanning distinct MCP servers; optimized via prompt optimization, SFT and GRPO.
Key insight: Small models can check whether each tool call matches the task's intent, a layer of oversight that authorization alone cannot provide.
Zhen Xu; Qizheng Zhang; Gerry Wan et al. arXiv: 2610.02670
Latency framework for speculative action proposals plus a small 0.6B drafter trained on target action sequences; up to 60% faster end-to-end with no systematic change in task success; drafter can be trained online with no prior trace collection.
Key insight: A small drafter trained on the target agent's own actions can speculate tool actions accurately enough to cut wall-clock time substantially.
Linh-An Phan; MingXue Wang; Guangyu Wu et al. arXiv: 2610.03315
Budget-bounded trajectory evaluation: offline rule profiles, online heuristic failure marking, fixed-budget serialization and one rubric-guided judge. +20-35 pp failure-localization alignment on Magentic-One and up to +23 on τ-retail vs AgentRx at ~6× lower cost and >8× faster; deployed in an enterprise platform.
Key insight: Budgeted preprocessing plus one rubric-guided judge localizes agent failures better and far cheaper than heavier trajectory evaluators.
Feng Chen; Ritam Dutt; Atnaz Taheri et al. arXiv: 2610.02627
Same email tasks rephrased along five style axes and four dialects across a RAG pipeline and two tool-using agents. Indirect requests hurt all three; formal requests hurt both agentic ones; failures are mostly omitted required actions rather than extra unsupported actions (NeurIPS 2026 workshop).
Key insight: Rephrasing the same email request indirectly or formally makes agents silently skip required actions, even when retrieval succeeds.
Nathan Conklin; Miranda Capra; Chris North arXiv: 2610.02369
Position paper: encode Nielsen heuristics, affordances, WCAG criteria and mixed-initiative principles as machine-readable skill files the UI-generating agent loads at runtime, turning the dialogue into a 'Space to Think'.
Key insight: Design and accessibility knowledge can ship as runtime skill files that UI-generating agents load, making quality a property of the process.
Songtao Wei; Yi Li; Zhichun Guo et al. arXiv: 2610.02396
Test-time evolution of multi-agent workflows with workflow inheritance (edit from the latest candidate, prune unhelpful nodes) and execution inheritance (reuse stored results only when full request and context match). GPT-4o-mini workers: 55.4% on WorkBench, 49.7% joint F1 on HotpotQA FullWiki; execution inheritance cuts worker tokens 29.1% / 34.6%.
Key insight: Inheriting both workflow structure and exact-match execution results across revisions makes multi-agent evolution cheaper and steadier.
Ryuichi Yamafuji Lun; Jingzhen Wang; Shreyas Kolte et al. arXiv: 2610.02349
Communication-layer integrity for multi-agent systems against Agent-in-the-Middle attacks: replicate a canonical payload across k routes and accept only on strict-majority digest agreement. 0% ASR below the α < 0.5 route-compromise threshold at 1× token cost; LLM-as-judge costs 35× and blocks up to 44.2% of benign outputs.
Key insight: Majority agreement across multiple message routes blocks agent-in-the-middle tampering without paying for an LLM judge.
Jabin Koo; Soheil Abbasloo; Sungjae Lee et al. arXiv: 2610.02951
For MoE backbones serving many agent roles, a lightweight predictor turns each agent's system and task prompts into a per-request expert mask in one forward pass. Beats static pruning/merging, generalizes to unseen workflows, largest margin when few experts are retained.
Key insight: An agent's own system and task prompts are enough to predict which MoE experts it needs, enabling per-request pruning.
Shiyi Kuang; Xuemei Luo; Kun Liu et al. arXiv: 2610.03153
450 adversarial workspace-agent tasks across six scenarios, verified via runtime traces and environment state, across 3 models × 3 harnesses (Claude Code, Codex, OpenClaw). Worst config (Codex + DeepSeek-V4-Pro-0813) hits 68.44% ASR; ASR varies more across models than harnesses.
Key insight: Runtime security risk in workspace agents varies more with the model than with the harness wrapped around it.
Xiqiao Xiong; Moxin Li; Zhixin Ma et al. arXiv: 2610.02920
Multi-agent framework that evolves safety harnesses from sparse threat evidence (short reports, a few examples) via adversarial safety-spec and attack-case generation. Consistently reduces attack success while preserving benign utility.
Key insight: Safety harnesses can be evolved from sparse threat reports by pitting specification generation against attack generation.
Hang Cui arXiv: 2610.03014
Security-aware static analysis building a dependency graph per sensitive operation (source evidence, trust boundaries, guards, external effects). AgentSecBench: 67 agent repos, 37,542 files, 23,866 candidates analyzed in 50.8 minutes; 22 established behaviors including one confirmed vulnerability; 91.1% context preserved vs 20.0% for sink-only.
Key insight: Whether an agent's sensitive operation is dangerous depends on its dependency, trust-boundary and guard context, not its identity alone.
Dimitrios Prasakis arXiv: 2610.02456
Open-source, local microVM sandbox for AI coding agents on macOS. Survey finds fewer than 40% of coding-agent users run agents in a sandbox; compared on 23 capability tests, Docker Sandboxes and SideKernel score highest on usability features.
Key insight: Most coding-agent users skip sandboxing; a usable local microVM sandbox for macOS targets the barriers they report.
Rudrendu Kumar Paul; Sourav Nandy arXiv: 2610.02503
150 production incidents → 23 failure modes in five categories (retrieval, generation, tool, orchestration, integration). Fault injection: circuit breakers cut cascade propagation 89%, quality gates catch 73% of silent degradation, isolation cuts blast radius 64%; 3+ patterns cut MTTR 71% (ICML 2026 workshop).
Key insight: Compound AI failures cluster at component boundaries, and a handful of resilience patterns sharply reduce cascades and recovery time.