Monday's cs.AI announcement day (2026-10-05) lists 124 new and 142 cross-lists (replacements skipped; listing total 266). A filter for agent systems, memory/context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent, persistence/identity, and local/open models keeps 39 papers. The day's through-line is that the system around the model keeps deciding outcomes: WebFovea lifts a live-web agent from 31.0 to 57.0 with the same model, VERSE finds harness self-evolution works only with execution-based verification, and FinSkillBench shows curated tools and skills — not self-written ones — carry the gains. Memory work shifts toward the write path: DyadMem measures how an agent should work with a specific user, APDMem reads memory coarse-to-fine, and Sentry keeps failure lessons out of context until a failure actually occurs. On the safety side, GHOST, COBRA and Pincer each show a different way long-running or computer-use agents drift past constraints, and propose layered fixes.


Research Papers

DyadMem: A Long-Term Memory Benchmark of How Agents Work with Users

Yifei Tao; Xinyu Zhong; Henry Hengyuan Zhao et al. arXiv: 2610.03020

Figure from DyadMem: A Long-Term Memory Benchmark of How Agents Work with Users
DyadMem: A Long-Term Memory Benchmark of How Agents Work with Users

Long-term memory benchmark for User-conditioned Relational Agent Memory (how this agent should work with this user): 3,065 episodes, 50,961 sessions, 61,210 QA with Capture/Update/Recall gold labels. Across 16 open-weight and 4 proprietary models, Gold-Memory QA is strong but Full-Pipeline QA drops sharply; low capture recall and unsafe deletion even in frontier models.

Key insight: Long-term agents need memory of how to work with a specific user, and the capture and update stages — not just recall — are where pipelines fail.

WebFovea: When the Model Is Right but the Click Is Wrong -- Reliable Round Trips for Vision-Based Web Agents on Live Websites

Jiangang Han arXiv: 2610.03036

Figure from WebFovea: When the Model Is Right but the Click Is Wrong -- Reliable Round Trips for Vision-Based Web Agents on Live Websites
WebFovea: When the Model Is Right but the Click Is Wrong -- Reliable Round Trips for Vision-Based Web Agents on Live Websites

2nd-place WebRetriever Challenge 2026 vision web agent (57.0/100). Same model across all four submissions; hidden-set score rose 31.0 → 57.0 from harness changes (up to run-to-run variance on live sites). Failures traced to parse → act → report → show stages: a coordinate-space mismatch put every click at 3/4 of intended coordinates, native dropdowns/iframes/text boxes failed silently, chat-template tokens contaminated 4.9% of episodes.

Key insight: On live websites, many agent failures sit in the plumbing between model and page — parsing, action effect, reporting and observation — not in the model's reasoning.

Sentry: Learning to Recover from LLM Agent Failures at Test Time

Changxiu Ji; Amy Lu; Qizheng Zhang et al. arXiv: 2610.02994

Figure from Sentry: Learning to Recover from LLM Agent Failures at Test Time
Sentry: Learning to Recover from LLM Agent Failures at Test Time

Failure-management layer beside the agent: on detected failure it retrieves matching lessons from an external playbook, verifies recovery without task rewards, and stores a lesson only if recovery worked; the full playbook never enters context. Beats the strongest runtime-intervention baseline by 37% on average and the strongest context-evolution baseline by 39%.

Key insight: Failure lessons are conditional knowledge: surfacing them only when a failure is detected works better than keeping the whole playbook in context.

DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents

A. Said Gurbuz; Ahmed Nassar; Sunghwan Hong et al. arXiv: 2610.02320

Figure from DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents
DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents

Controllable desktop environment composing real apps to produce DeskForge-1M (1.2M annotated observations, 159.7M elements). Fine-tuning Qwen3.5-4B on 200K grounding examples: +11.51 pp ScreenSpot-Pro, +10.11 OSWorld-G; WebArena-Infinity 31 → 50 of 119 tasks, OpenApps 3 → 15 of 100. Code, data and model released.

Key insight: Composing real desktop applications under controlled variation yields dense grounding supervision that transfers to long-horizon computer use.

APDMem: Agent-Controlled Progressive Disclosure for Query-Adaptive Long-Term Memory

Chin-Lun Fu; Anagha Kulkarni; Hong Ni et al. arXiv: 2610.02472

Figure from APDMem: Agent-Controlled Progressive Disclosure for Query-Adaptive Long-Term Memory
APDMem: Agent-Controlled Progressive Disclosure for Query-Adaptive Long-Term Memory

Four-layer memory (thematic summaries, personalized key facts, turn-level evidence notes, raw messages) with a controller that drills down only when needed and a note synthesizer that orders events and flags contradictions. Strong LongMemEval performance while accessing only 8% of conversations (EMNLP 2026 Industry).

Key insight: Reading memory coarse-to-fine lets simple queries stop early while hard temporal or multi-hop queries drill into raw evidence.

A GHOST in Long-Horizon Agents: Governance Hazard from Overlooked Safety Constraints across Turns

XinPeng Shen; Lan Zhang; Yixiao Huang et al. arXiv: 2610.02664

Figure from A GHOST in Long-Horizon Agents: Governance Hazard from Overlooked Safety Constraints across Turns
A GHOST in Long-Horizon Agents: Governance Hazard from Overlooked Safety Constraints across Turns

Names the failure where a long-horizon agent violates a safety constraint stated many turns earlier under benign conditions: 11.5% occurrence on GPT-5.5. STAR-Guard restores applicable historical constraints and adds a deterministic pre-execution audit; no GHOST events observed under the GPT-5.5 setup.

Key insight: Long interaction histories let agents violate constraints stated many turns earlier, even under benign conditions; restoring and auditing those constraints prevents it.

Securing Computer-Use Agents Against Branch Steering Attacks

Giulio Zingrillo; Hanna Foerster; Ilia Shumailov et al. arXiv: 2610.03089

Figure from Securing Computer-Use Agents Against Branch Steering Attacks
Securing Computer-Use Agents Against Branch Steering Attacks

Shows Dual-LLM CUAs are vulnerable to branch steering — untrusted page data pushes the agent down a hazardous pre-approved branch. STEER-Bench (101 tasks): 94.4% ASR on standard and 89.5% on vanilla Dual-LLM CUAs. COBRA pairs branching plans with ahead-of-time capability constraints: 0% ASR, 97% benign utility.

Key insight: Plan-first defenses for computer-use agents leak through branch steering; bounding each branch's parameters and destinations closes the gap.

Pincer: Resource Authorization for Agents using a Digital Twin

Mayank Rathee; Alexander Stepanov; Shalin Madabhavi et al. arXiv: 2610.02569

Figure from Pincer: Resource Authorization for Agents using a Digital Twin
Pincer: Resource Authorization for Agents using a Digital Twin

Resource-layer defense for coding agents: an isolated-context digital twin continually learns user-specific least-privilege policies from multi-day interaction and answers the agent's permission requests as the user's proxy, alongside tool-call-layer auto modes. Outperforms LLM-judge baselines and Conseca adaptations.

Key insight: A continually learning digital twin of the user can answer an agent's permission requests with user-specific least privilege.

VERSE: Verified Self-Evolving Optimizer for Agent Harnesses

Zekai Wang; Yingqiang Ge; Zekun Wang et al. arXiv: 2610.02616

Self-evolving harness optimizer that tests draft edits, replays failures and perturbs suspected steps before submitting, and revises both the executor harness and its own prompts, skills, tools, hooks and notes (weights fixed). Best validation-selected harness: 42.3% / 37.7% on held-out SWE-rebench and newer OOD tasks vs 39.2% / 29.3% for the strongest baselines.

Key insight: Letting a harness optimizer rewrite its own tools only pays off when every edit is checked by execution, not by the optimizer's own judgment.

Knowledge or Calculator? Decomposing the Skill Premium in Verifiable Financial Agent Workflows

Jermyn Zhen Yong Bek; Zhuang Qiang Bok; Zhongtian Sun arXiv: 2610.03564

2,603 point-in-time finance episodes, 17,820 runs across 9 models. Curated skill packages add +16.2 points (0.366 → 0.528); skills generated within one episode add only +0.5 while costing more tokens/turns. Docs alone +5.6, tools alone +19.5, combined subadditive; a second harness reproduces direction but not magnitude.

Key insight: Curated skills and executable tools carry most of the measured 'skill premium'; skills an agent writes for itself mid-episode add almost nothing.

Harness-Aware Distillation for Small Language Model Agents

Moonseok Choi; Taehong Moon; Giung Nam et al. arXiv: 2610.02858

HAD distills a teacher agent into a smaller student while the harness stays in place: action preferences contrast the teacher's actions with vs without harness info, plus a validity check against harness records; no task rewards needed. Beats on-policy distillation with the same fixed harness; fewer unproductive loops, more error recovery.

Key insight: When the harness stays in place, distillation should target what the teacher adds beyond the harness rather than imitating its full outputs.

Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite

Zongxia Li; Yucheng Shi; Zhongzhi Li et al. arXiv: 2610.02826

One Qwen-3.8-27B base model solves tasks under diverse harnesses, then a planner/critic/executor rewrites successes into runbook-style trajectories for a general harness: 2,001 source trajectories → 11,094 rewrites. Terminal-Bench 2 pass@3 57.0% → 74.2%; open weights released (IntelligenceLab/RSR-27B).

Key insight: Successes found under specialized harnesses can be rewritten into training trajectories that work under a plain, general harness.

When Terminal-Agent Training Stalls: Demystifying Data Generation and Verification Challenge

Xi Qin; Isabel Kurth; Xin Cui et al. arXiv: 2610.02405

Diagnoses meta-agent (Claude Opus) generated terminal tasks/verifiers: benchmark invalidity, harness brittleness, reward misalignment. Prompt redesign and context extension raise solvability 5.6×; a 9B model saturates at 81.3% mean pass@2, dropping to 20.6% when hard tasks are added.

Key insight: A runnable Docker image and test suite do not make a faithful training pipeline; solvability bands and verifier audits must be measured explicitly.

Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents

Yu Li; Guangfeng Cai; Long-Fei Li et al. arXiv: 2610.03634

Dependency-Aware Group Policy Optimization builds a command dependency graph from terminal traces and traces back from what the verifier inspects, assigning credit to relevant writes and supporting reads. Improves performance and training stability on complex terminal tasks.

Key insight: Tracing read-write dependencies between terminal commands gives sharper credit assignment than trajectory- or step-level rewards.

Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses

Jiawei Li arXiv: 2610.02267

Paired evaluation of an open-weight (Laya) and hosted (Jev) single-forward-pass classifier on 11 harness decision points (routing, tool choice, RAG gating, injection). Jev wins 9/11 (+10.8 to +46.0 pp); neither beats chance on zero-shot model routing; Laya flips 30% of answers on option-order reversal and hits 31% at 50 nearest-neighbour tools. Self-audit: a reported 23.9% saving was actually 4.3%.

Key insight: Single-pass decision classifiers for harness gates can be order-sensitive and collapse on large tool catalogs, and cost-saving claims are easy to overstate.

Interpreting at Write Time: A Policy Ablation for Multi-Goal Agent Memory

Albert Sadowski; Jarosław A. Chudziak arXiv: 2610.02897

Compares three summarization write policies for a multi-goal assistant: goal-neutral, one all-goal summary, or one summary per goal. Per-goal summaries win on relevance, completeness and accuracy; the all-goal summary loses even to the neutral one written at a fraction of its budget.

Key insight: Summaries written per standing goal beat a single all-purpose summary, because goals disagree about what was worth keeping.

TPBench: A Turning-Point Benchmark for Dialogue Compression

Minji Park; Seunghyun Yoon; Hyuk Lim arXiv: 2610.02736

Benchmark for 'turning-point eviction': compressors keep dialogue facts but drop the turn where the user revised them. At retained fraction 0.30 every compressed method stays below full context; deleting the update turn sharply lowers current-value accuracy; recency is the best compressed method on LongMemEval-KU and RiSAWOZ.

Key insight: Dialogue compressors can keep the facts but drop the turn where the user changed them; overall retention scores hide this.

Decoupling Memory from Context: Structured Memory for Token-Efficient Test-Time Continual Learning

Yehya Farhat; Michael Desmond; Anastasios Kyrillidis arXiv: 2610.02687

Frames agent memory updates as optimization over the model's context; GraphMemory accumulates and connects reusable strategies and retrieves only the relevant subgraph per query, keeping retrieved memory constant as examples grow. Competitive accuracy with roughly 81-85% fewer memory-construction tokens.

Key insight: Treating memory updates as context optimization and retrieving only a relevant subgraph keeps memory cost bounded as experience grows.

When History Fails to Become Experience: Action Calibration in Language Agents

Jingyu Liu; Zhiwen Wang; Yuxin Jing et al. arXiv: 2610.02769

Shows agents barely use action–outcome correspondence in history: much of history's benefit survives shuffling past actions. Explicitly labeling each observation as the outcome of the preceding action improves success and reduces repetition; a learned calibrator that reassesses past actions helps further.

Key insight: Agents often fail to connect past actions with their outcomes; explicitly labeling each observation with its action measurably helps.

Continual Graph Memory for Mathematical Research Agents

Junyi Zhang; Jinxi Yu; Eric Hanchen Jiang et al. arXiv: 2610.02945

Math research agent built on a cross-problem graph memory of facts, plans and counterexamples with dependency-aware retrieval, an evidence-sensitive curator and scoped recall of negative findings. Reports closure on all ten First Proof Second Batch problems and solutions to the Jamison caterpillar conjecture and Erdős Problems 289, 348 and 488.

Key insight: A typed graph memory of facts, plans and counterexamples — including negative results — lets parallel research agents sustain and resume long searches.

Learning What to Investigate Next: Meta-Reasoning for Long-Horizon Research Agents

Ankur Samanta; Yonathan Efroni; Paul Sajda et al. arXiv: 2610.02525

Hierarchical research agent: an outer meta-reasoner curates context from a persistent research record and writes work orders; a fresh inner executor runs each. A generative critic forecasts remaining return at decision boundaries; MIRA-AC trains on meta-decisions only and improves gold performance across four autoresearch environments.

Key insight: Separating what to investigate next from how to execute it turns long-horizon research steering into a learnable policy.

KV$^2$: A Self-Refining KV Cache

Johannes Wesch; Danni Liu; Jan Niehues arXiv: 2610.03198

Query-agnostic KV-cache compression for a prefilled context reused across many queries: a proxy scorer picks informative tokens, then only that subset is reprocessed for eviction scores. On RULER 16K at a 2% budget it beats the next-best baseline by more than 40 points.

Key insight: A reusable prefilled context can be compressed accurately without reprocessing the whole prompt, by rescoring only a proxy-selected subset.

When to Compile a Computer-Use Agent? Measuring Payback and Making Compilation Decisions for Token Efficiency

Yulong Ming; Jie Xu; Zihan Wu et al. arXiv: 2610.02932

Payback-Aware Compilation from Experience: measures costs of compiling repeated GUI procedures into programs (including failed attempts) and decides online when to compile. Payback of 2-16 uses for successful compilations; 17.3% token reduction vs ReAct and 24.9% vs an AutoRPA adaptation.

Key insight: Compiling repeated GUI procedures into programs is worth it only when expected reuse beats compile cost, including failed attempts.

HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents

Alham Fikri Aji; Faiz Rizki Ramadhan; Zayd M. K. Zuhri et al. arXiv: 2610.03574

423 hard human-validated browsing questions across 13 languages requiring obscure evidence in videos, scans, images or maps; questions answerable without internet are filtered out. Evaluates provider-native search and a shared retrieval harness under a common agent protocol, with a human baseline.

Key insight: Multilingual, multimodal web research remains far from solved when evidence is obscure and spread across videos, scans and maps.

DeReAct: Decomposed Reasoning and Acting for Reliable AI Agents

Ajay Vohra; Tao Chen; Neeti Narayan et al. arXiv: 2610.02351

Splits ReAct into a Brain plus an action-validating Critic and a Context Manager that reconstructs environment-supported state and certifies completion. On GAIA and SWE-bench Verified: +6.5-7.0 Pass@1 for Qwen3-Coder-480B and +4.2-5.2 for Claude Sonnet 4.5; comparable with Claude Opus 4.5 but more evidence-complete trajectories.

Key insight: Externalizing action validation and completion certification from the main policy helps weaker models most and improves grounding for strong ones.

Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents

Yu Li; Zheng Zhang; Xin Liu et al. arXiv: 2610.02330

Trains a Comparative Inference Model to estimate how likely a candidate next tool call is to support final success, using observed tool behavior, a Bayesian tool-graph simulator and LLM comparisons. Improves Tool F1 and task success across three tool-use benchmarks (NeurIPS 2026).

Key insight: Tool-use agents can learn to estimate the long-horizon value of a candidate tool call before executing it.

Toward SLM-based agentic task-tool intent matching

Chiara Troiani; Arash Salarian; Majed El Helou et al. arXiv: 2610.03213

SLM classifier that checks each selected tool call against the task intent for low-latency, on-prem per-call oversight; dataset of multi-tool tasks spanning distinct MCP servers; optimized via prompt optimization, SFT and GRPO.

Key insight: Small models can check whether each tool call matches the task's intent, a layer of oversight that authorization alone cannot provide.

LEAP: Learning Efficient Action Proposals For LLM Agents

Zhen Xu; Qizheng Zhang; Gerry Wan et al. arXiv: 2610.02670

Latency framework for speculative action proposals plus a small 0.6B drafter trained on target action sequences; up to 60% faster end-to-end with no systematic change in task success; drafter can be trained online with no prior trace collection.

Key insight: A small drafter trained on the target agent's own actions can speculate tool actions accurately enough to cut wall-clock time substantially.

Lightweight, Rubric-Guided Trajectory Evaluation for Production AI Agents

Linh-An Phan; MingXue Wang; Guangyu Wu et al. arXiv: 2610.03315

Budget-bounded trajectory evaluation: offline rule profiles, online heuristic failure marking, fixed-budget serialization and one rubric-guided judge. +20-35 pp failure-localization alignment on Magentic-One and up to +23 on τ-retail vs AgentRx at ~6× lower cost and >8× faster; deployed in an enterprise platform.

Key insight: Budgeted preprocessing plus one rubric-guided judge localizes agent failures better and far cheaper than heavier trajectory evaluators.

Lost in the Request: How Communication Variation Disrupts Retrieval and Action in Email Agents

Feng Chen; Ritam Dutt; Atnaz Taheri et al. arXiv: 2610.02627

Same email tasks rephrased along five style axes and four dialects across a RAG pipeline and two tool-using agents. Indirect requests hurt all three; formal requests hurt both agentic ones; failures are mostly omitted required actions rather than extra unsupported actions (NeurIPS 2026 workshop).

Key insight: Rephrasing the same email request indirectly or formally makes agents silently skip required actions, even when retrieval succeeds.

Automating the Application of HCI Principles: Skills for On-Demand UI Construction, the Human-AI Space to Think, and the Future of HCI

Nathan Conklin; Miranda Capra; Chris North arXiv: 2610.02369

Position paper: encode Nielsen heuristics, affordances, WCAG criteria and mixed-initiative principles as machine-readable skill files the UI-generating agent loads at runtime, turning the dialogue into a 'Space to Think'.

Key insight: Design and accessibility knowledge can ship as runtime skill files that UI-generating agents load, making quality a property of the process.

Inherit-MAS: Test-Time Evolution of Multi-Agent Systems through Workflow and Execution Inheritance

Songtao Wei; Yi Li; Zhichun Guo et al. arXiv: 2610.02396

Test-time evolution of multi-agent workflows with workflow inheritance (edit from the latest candidate, prune unhelpful nodes) and execution inheritance (reuse stored results only when full request and context match). GPT-4o-mini workers: 55.4% on WorkBench, 49.7% joint F1 on HotpotQA FullWiki; execution inheritance cuts worker tokens 29.1% / 34.6%.

Key insight: Inheriting both workflow structure and exact-match execution results across revisions makes multi-agent evolution cheaper and steadier.

MIRROR: Multipath Quorum Integrity for LLM Multi-Agent Communication

Ryuichi Yamafuji Lun; Jingzhen Wang; Shreyas Kolte et al. arXiv: 2610.02349

Communication-layer integrity for multi-agent systems against Agent-in-the-Middle attacks: replicate a canonical payload across k routes and accept only on strict-majority digest agreement. 0% ASR below the α < 0.5 route-compromise threshold at 1× token cost; LLM-as-judge costs 35× and blocks up to 44.2% of benign outputs.

Key insight: Majority agreement across multiple message routes blocks agent-in-the-middle tampering without paying for an LLM judge.

Dynamic Expert Pruning for Multi-Agent Systems

Jabin Koo; Soheil Abbasloo; Sungjae Lee et al. arXiv: 2610.02951

For MoE backbones serving many agent roles, a lightweight predictor turns each agent's system and task prompts into a per-request expert mask in one forward pass. Beats static pruning/merging, generalizes to unseen workflows, largest margin when few experts are retained.

Key insight: An agent's own system and task prompts are enough to predict which MoE experts it needs, enabling per-request pruning.

EvoRiskBench: An Evolving Benchmark for Runtime Security Risks in Workspace Agents

Shiyi Kuang; Xuemei Luo; Kun Liu et al. arXiv: 2610.03153

450 adversarial workspace-agent tasks across six scenarios, verified via runtime traces and environment state, across 3 models × 3 harnesses (Claude Code, Codex, OpenClaw). Worst config (Codex + DeepSeek-V4-Pro-0813) hits 68.44% ASR; ASR varies more across models than harnesses.

Key insight: Runtime security risk in workspace agents varies more with the model than with the harness wrapped around it.

HASTE: Evolving Agent Harnesses Against Emerging Attacks Using Sparse Evidence

Xiqiao Xiong; Moxin Li; Zhixin Ma et al. arXiv: 2610.02920

Multi-agent framework that evolves safety harnesses from sparse threat evidence (short reports, a few examples) via adversarial safety-spec and attack-case generation. Consistently reduces attack success while preserving benign utility.

Key insight: Safety harnesses can be evolved from sparse threat reports by pitting specification generation against attack generation.

Beyond Predefined Sinks: Security-Aware Dependency Analysis for LLM Agents

Hang Cui arXiv: 2610.03014

Security-aware static analysis building a dependency graph per sensitive operation (source evidence, trust boundaries, guards, external effects). AgentSecBench: 67 agent repos, 37,542 files, 23,866 candidates analyzed in 50.8 minutes; 22 established behaviors including one confirmed vulnerability; 91.1% context preserved vs 20.0% for sink-only.

Key insight: Whether an agent's sensitive operation is dangerous depends on its dependency, trust-boundary and guard context, not its identity alone.

SideKernel: A Usable microVM Sandbox for AI Coding Agents on macOS

Dimitrios Prasakis arXiv: 2610.02456

Open-source, local microVM sandbox for AI coding agents on macOS. Survey finds fewer than 40% of coding-agent users run agents in a sandbox; compared on 23 capability tests, Docker Sandboxes and SideKernel score highest on usability features.

Key insight: Most coding-agent users skip sandboxing; a usable local microVM sandbox for macOS targets the barriers they report.

Compound AI System Reliability: A Failure Taxonomy and Resilience Pattern Catalog from 150 Production Incidents

Rudrendu Kumar Paul; Sourav Nandy arXiv: 2610.02503

150 production incidents → 23 failure modes in five categories (retrieval, generation, tool, orchestration, integration). Fault injection: circuit breakers cut cascade propagation 89%, quality gates catch 73% of silent degradation, isolation cuts blast radius 64%; 3+ patterns cut MTTR 71% (ICML 2026 workshop).

Key insight: Compound AI failures cluster at component boundaries, and a handful of resilience patterns sharply reduce cascades and recovery time.