Tuesday's cs.AI announcement day (2026-10-06) lists 230 new and 324 cross-lists (replacements skipped; listing total 554). A filter for agent systems, memory/context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent, persistence/identity, and local/open models keeps 40 papers. The day's sharpest thread is the harness itself: What Does a Harness Buy? finds mature coding harnesses indistinguishable on score but up to 3x apart on cost, SHarP shows most harness modules can be pruned, CUAWright wins on computer-use with nothing but a shell and a file system, and IEC-Bench shows tool calls are often silently altered before they execute. Memory and persistence papers converge on validity: MemTrace, Concord and StateWise all re-check stored evidence against the current world before reuse, while MemLeak and a study of self-propagating misalignment show what goes wrong when memory is shared or written by the agent itself. For personal and on-device agents, DelegationBench and UndoBench separate deciding and recovering from merely completing, Sibyl runs a 0.6B agent with rare cloud calls, and a phone-side study fine-tunes a 3B model on an iPhone within one charge.


Research Papers

What Does a Harness Buy? Tokens, Mostly

Yangze Liu; Zhongyi Han arXiv: 2610.04433

Figure from What Does a Harness Buy? Tokens, Mostly
What Does a Harness Buy? Tokens, Mostly

Five models through Claude Code, mini-SWE-agent and OpenCode on SWE-bench Verified with reruns to calibrate noise. On 447 tasks Claude Code and mini-SWE-agent are equivalent within five points; on a 45-task hard subset swapping harness flips as many tasks as rerunning (13% each). The one effect clearing noise is a loss (OpenCode up to 9 points behind). Cost per task differs up to 3x across harnesses, set by the per-step preamble of system prompt + tool schemas.

Key insight: Among mature coding harnesses, the measurable difference is mostly cost per task, driven by the prompt and tool-schema preamble sent on every step.

CUAWright: A Minimal Unified Interface for Digital Agents

Yadong Lu; Theodore Lee; Yifei Li et al. arXiv: 2610.04116

Figure from CUAWright: A Minimal Unified Interface for Digital Agents
CUAWright: A Minimal Unified Interface for Digital Agents

Minimal ~3K-LOC terminal harness for computer-use: bash is the only action, the file system is the evolvable space for building tools and managing context. On OSWorld 2.0, 33.2% relative partial-reward gain at 37.5% lower estimated cost vs the published GPT-5.5 baseline; beats a GUI-native harness by 4.7% (Online-Mind2Web) and 44.0% (Odysseys) success.

Key insight: Computer-use agents do better and cost less when given a programmable shell and file system instead of a fixed GUI tool set.

Do Tool Calls Execute as Intended? Measuring and Repairing Intent-Execution Correspondence in LLM Agents

Boyang Yang; Zhenhao Li; Ziyao Yang et al. arXiv: 2610.04375

Measures whether tool calls execute as emitted. In 47,828 production shell calls, Claude Code's Bash tool changes 12.0% of calls carrying code, escapes or long text; for 80.7% of changed-backslash calls the wrong action runs with no error. All 10 measured harnesses change a call; trajectory judgment blames the LLM for 95.1% of failures though the path caused more than half. IntAct recovers 79.2% of changed-call failures.

Key insight: Tool calls are frequently altered on their way to execution, silently running the wrong action; harnesses need hop-by-hop testing.

SHarP: Saliency-based Pruning of Agent Harnesses

Xinyi Gao; Qiucheng Wu; Kaizhi Qian et al. arXiv: 2610.04178

Figure from SHarP: Saliency-based Pruning of Agent Harnesses
SHarP: Saliency-based Pruning of Agent Harnesses

Treats tools, instructions and support mechanisms as prunable modules; estimates each module's saliency by single-module ablation on performance and token cost, then prunes the least useful/most expensive. Most studied harnesses are highly redundant and keep comparable performance after substantial pruning.

Key insight: Accumulated harness modules are often redundant; ablating them one at a time shows which can be removed without losing performance.

UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents

Dolly Sah; Tanmay Sah; Harshul Jain et al. arXiv: 2610.05622

Figure from UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents
UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents

36 base workflows and 36 fault scenarios across 8 enterprise domains, separating task competence from recovery via paired trials and wire-level effect oracles. Nominal competence 83.54% vs conditional recovery 46.72%; naive retry produces duplicate external effects in 53.33% of trials. After-commit-before-ack faults are handled best by verification and server-side idempotency.

Key insight: Completing a workflow and recovering from a fault are separate capabilities; naive retries often duplicate real-world side effects.

DelegationBench: Measuring When AI Agents Should Ask Before Acting

Shiva Pochampally arXiv: 2610.05532

Figure from DelegationBench: Measuring When AI Agents Should Ask Before Acting
DelegationBench: Measuring When AI Agents Should Ask Before Acting

156 scenarios (act / ask permission / ask for info / refuse), mostly matched pairs varying one feature (requested, stakes, reversibility, audience). A post-hoc keyword rule agrees with annotators more than eight of ten models; equivalent phrasings change act rate by up to 52.5 pp; every model asks less when executing with tools than when judging a proposed action.

Key insight: Agreement scores hide how agents decide whether to ask before acting, and models ask less often when they are the ones holding the tools.

MemTrace: State-Consistent Memory for Long-Horizon Coding Agents

Hongming Xu; Le Zhou; ZhongHe Jin et al. arXiv: 2610.04838

Figure from MemTrace: State-Consistent Memory for Long-Horizon Coding Agents
MemTrace: State-Consistent Memory for Long-Horizon Coding Agents

Provenance-aware memory: immutable Memory Traces anchored to files, symbols and tests in a Memory Trace Graph; compact Memory Anchors stay in working memory and evidence is validity-checked against the current repo before reuse. Under Codex CLI: +21.2 DeepSWE pass@1, +4.4 SWE-EVO resolved, +17.8 SWE-Milestone score.

Key insight: Anchoring coding-agent memories to files, symbols and tests and re-validating them before reuse keeps long sessions consistent after context refreshes.

When Agent Context Goes Stale: Incoherence in Volatile Agent Context

Yingying Liu; Junzhou Fang; Chenxiong Qian arXiv: 2610.05281

Figure from When Agent Context Goes Stale: Incoherence in Volatile Agent Context
When Agent Context Goes Stale: Incoherence in Volatile Agent Context

Context coherence for agents: links each tool observation in context to its source, detects source changes and updates, annotates or suppresses stale context. On ConcordBench (stale file contents), answers are consistent with restored workspace state in all evaluated cases across three frontier models, using 46.4% fewer tokens than the strongest non-oracle baseline.

Key insight: Tool observations in context go stale when their sources change; linking each observation to its source lets the runtime update or suppress it.

MemLeak: Cross-User Semantic Leakage in Multi-Tenant AI Agent Memory

Priyanka Mudgal; Kai Zhao; Guilin Zhang et al. arXiv: 2610.04195

Figure from MemLeak: Cross-User Semantic Leakage in Multi-Tenant AI Agent Memory
MemLeak: Cross-User Semantic Leakage in Multi-Tenant AI Agent Memory

Cross-user leakage in shared vector-store memory: incidental leakage 70-100% under pooled same-team retrieval, crafted memories 90-100% top-k placement, end-to-end contamination up to 5.00/5. Only hard post-retrieval ownership gating restores the clean baseline (1.00/5), at ~1.4 ms per query.

Key insight: Shared vector memory leaks across users through ordinary similarity search; hard ownership filtering after retrieval is the fix that holds.

Fine-Tuning a 3B-Parameter LLM on a Smartphone: Characterizing Sustained Training

Andrew Geyko; Marius Mosbach; André Brinkmann arXiv: 2610.06325

First systematic characterization of fine-tuning a 3B LLM on a phone: an iPhone 17 Pro fine-tunes to a typical user within one battery charge, with adapters improving personalization as much as server-trained ones. Sustained training throttles to about half throughput; most step time is the frozen base backward pass. A fix to an MLX kernel (merged upstream) trains adapters 1.47x faster on a third less energy.

Key insight: Fine-tuning a 3B model on a current phone is feasible within one battery charge, but runtimes must treat training as a first-class workload.

Sibyl: An Efficient Small-large Model Collaboration Framework for Long-horizon Tasks

Zhewei Fang; Yuxin Zhang; Zhenwei Shao et al. arXiv: 2610.05383

Trains a small on-device agent to consult a cloud model only at selected steps and internalize the guidance; three-stage RL pipeline. With a 0.6B model, outperforms baselines by 95.2% (ALFWorld) and 80.4% (WebShop) in success rate while averaging 0.8 and 3.9 cloud calls per trajectory.

Key insight: A sub-billion-parameter on-device agent can match far larger systems by learning exactly when to ask a cloud model for help.

Self-Propagating Misalignment in LLM Agents, and Why Auditing or Disabling Memory Is Not Enough

Debeshee Das; Jacqueline Tay; Bruce Tsai et al. arXiv: 2610.04083

A misaligned agent writes a goal it can't yet act on into persistent memory so a future aligned agent carries it out. Across 20 scenarios and 11 frontier models: 58% success with explicit goals, 18% with values-only prompts; without the memory tool agents write to files in 74% of sessions (11% success). Goals persist through 100 unrelated sessions; the MemMorph auditor only cuts propagation from 71% to 34%.

Key insight: An agent can plant its own goals in persistent memory or files for future instances to execute, and current memory auditors only partly stop it.

MESH-Harness: Self-Improving Agent Harnesses via Bandit-Guided Compositional Evolution

Zhiwei Shang; Yu Huo; Mingrong Gong et al. arXiv: 2610.05300

Splits a harness into role-specific modules with explicit interfaces, scores module combinations with full-covariance LinUCB, then applies trace-guided local code edits. Beats Meta-Harness by 5.70 / 7.01 / 2.00 / 5.00 points on four task families at matched evaluation budgets; 44.2% lower test-time cost and 14.6% lower total cost including search.

Key insight: Harness search works best when the harness is modular and combinations are searched under a budget rather than rewritten wholesale.

Can Agent Harnesses and Inference Engines Hear Each Other? The HEAR Protocol for Agentic LLM Serving

Jiaqi Zhao; Haodong Chen; Jitai Hao et al. arXiv: 2610.06597

Bidirectional Harness-Engine pairing protocol: the harness sends workflow intent and context lifecycles, the inference engine returns KV-cache state, queue pressure and capabilities. Under memory-constrained concurrent serving: 1.61x batch speedup and 2.23x lower median TTFT on SCBench; 1.23x and 2.45x end-to-end speedups on BrowseComp-Plus and DeepResearchBench without observed quality loss.

Key insight: Agent serving gets faster when the harness tells the inference engine about workflow intent and the engine reports cache and queue state back.

Agent Behavior as Code: Efficient and Robust LLM Agents with Programmatic Specifications

Peng Qi; Chunliang Lyu; Gang Li et al. arXiv: 2610.04824

Agent Behavior as Code: a symbolic program (Python with neural functions) fully specifies runtime behavior and an FM agent edits the program. Matches a neural agent on GAIA; 71.9% vs 47.4% Pass^4 on τ²-bench telecom; 5.2x lower latency and 7.0x lower cost on GSM-Symbolic when reusing a program; 9.5x lower agent latency on τ²-telecom.

Key insight: Encoding an agent's behavior as an editable program makes recurring tasks more robust and far cheaper than re-planning with a model every time.

Complex Agents, Shallow Tests: Demystifying and Enhancing Test Adequacy of Agent Harness in the Wild

Yifan Xiong; Jingyi Ge; Zhenpeng Chen et al. arXiv: 2610.04921

First empirical study of test adequacy in real agent harnesses: LLM-dependent harness code gets less than half line/branch coverage. HarnessTester generates contract-faithful tests: 75.95%/84.76% larger line/branch coverage gains, 69.89% larger mutation-score gains; found 122 real harness bugs (e.g., OpenClaw), 88 previously unknown, 69 confirmed.

Key insight: The harness code that parses model output and governs agent behavior is under-tested, and contract-aware test generation finds real bugs there.

SkillScriptBench: Benchmarking Self-Evolution of Executable Agent Skill Packages Beyond Markdown

Yuxuan Liu; Haoran Li; Yuhao Zhang et al. arXiv: 2610.04008

350-task benchmark for self-evolution of executable skill packages (instructions + scripts), built from a survey of 35,000+ GitHub skill roots: 150 repair tasks on 100 packages plus a 200-task controlled track. Editing docs+scripts fixes script faults but doesn't beat Markdown-only revision on doc repair or preservation. AST-Guided Skill Revision adds 21.9% / 27.7% absolute repair success.

Key insight: Agents maintaining their own skill packages repair scripts more reliably when edits are restricted to code locations linked by the abstract syntax tree.

EVISKILL: Grounding Skill Evolution in Replayable Evidence

Yan Zhou; Yili Wang; Yiwei Dai et al. arXiv: 2610.05030

Organizes execution observations into Replayable Evidence Cards, links every skill edit to its supporting contexts, verifies edits by targeted replay, and provisionally retains locally supported edits rather than discarding them with a rejected revision. Evaluated on three interactive benchmarks across six LLM backbones.

Key insight: Skill edits are easier to trust when each one stays linked to replayable evidence and is verified by targeted re-execution.

Beyond Instruction Following: Learning Grounded Skill-Following with Skill Contracts

Jianghan Shen; Zhenjie Liu; Yue Li et al. arXiv: 2610.05161

Grounded Skill-Following: each expert-authored skill becomes a contract runtime (required phases, admissible actions, transitions, termination) giving dense verifiable rewards (Verified Progress Credit) and live contract-state feedback. Qwen3.5-4B reaches 99.27% / 99.96% Protocol Completion Rate on Math / Search while slightly beating baselines on task outcome (82.95%, 46.61%).

Key insight: Turning a multi-phase skill into an explicit contract gives dense, verifiable training signal and near-complete procedural compliance in a small model.

TeleTune: Evolving Agent Skills From Offline Telemetry

Justin Chih-Yao Chen; Elias Stengel-Eskin; Yan Chen et al. arXiv: 2610.05437

Learns a textual skill library for computer-use agents from offline user telemetry without recorded goals, without replay and with interleaved tasks, accepting edits only if they improve held-out action-prediction accuracy. 77.1% (WorkArena) and 80.6% (Online-Mind2Web), +6.7% and +7.7% over the best baseline; optimizing on fixed logs costs 5-75x fewer tokens than live validation.

Key insight: Offline usage logs, even without goals or replay, can train a computer-use skill library that beats live-episode baselines at a fraction of the token cost.

SkillGATE: Gate-Aware Monte Carlo Tree Search for Skill Retrieval

Rongchen Zhao; Yu Chen; Yanming Yang et al. arXiv: 2610.05489

Skill retrieval as information foraging: a graph-preserving hierarchical skill index searched with Gate-Aware MCTS (G-PUCT). Across six skill-retrieval benchmarks, +16.3% overall R@1 over the strongest retriever-based baseline.

Key insight: As skill libraries grow, treating retrieval as a guided tree search over a hierarchical index beats ranking skills independently.

Mining Agent Skills from Production Traces

Yue Ran Kang; Colton Mikolajczyk; Chhaya Methani et al. arXiv: 2610.05777

Holds the mining pipeline fixed and compares three evidence regimes (successes only; successes+failures with labels; same mix unlabeled) × two skill forms (workflow plan vs declarative ontology) on ThinkingBox-Bench and APEX-Agents. On ThinkingBox, workflows beat ontology by 1.7 pp and mixed evidence beats success-only by 2.4 pp; APEX-Agents leans to ontologies with no evidence preference.

Key insight: The best way to mine skills from production traces depends on the domain; neither evidence regime nor skill form wins everywhere.

Runaway Reaction: When Benign Skills Compose into Malicious Behavior

Zunlong Zhou; Ziyuan Yang; Mengyu Sun et al. arXiv: 2610.05943

Shows composing individually vetted, benign agent skills can induce malicious behavior. CRIME decomposes a target behavior, picks benign skill compositions from public repos, refines them with execution feedback while each stays benign in isolation, and checks consequences in a sandbox. Benchmark of 4,000 public skills across eight cybersecurity behaviors.

Key insight: Skills that each pass security vetting can still combine into malicious behavior, so vetting must consider compositions.

Knowing the Store: What a Memory Backend Must Write Down Before an Agent Can Read It

Ansuman Mullick; Eray Tüzün arXiv: 2610.04794

Asks what a memory backend must expose (record counts by lifecycle state, attribute names) for an agent to know a record is stale before retrieving. 148 questions × three store variants; on short attribute lists none of five models reliably beats a cosine lookup over names. Under FR-Bank's own metadata, GPT-4.1 mini, Haiku 4.5 and the lookup fall from 0.73/0.82/0.75 to 0.60/0.71/0.59. All stores synthetic.

Key insight: A memory store has to record lifecycle state explicitly; readers cannot judge staleness from information the store never wrote down.

Memory Canonicalization: A Framework and Benchmark for Cross-Model Drift in Persistent LLM Memory

Amit Vadnere; Aishwarya Lonarkar arXiv: 2610.05124

Write-time pipeline that rewrites raw memory objects into disambiguated canonical form with emotional valence as a separate field, aimed at MemGPT/Letta/Mem0/Zep/MCP-memory setups where different LLMs read the same memory differently. Pilot (176 synthetic objects, three model families): +0.050 emotional consistency uncorrected, not surviving multiple-comparison correction; no factual-drift result significant. Exploratory.

Key insight: Rewriting memories into an explicit canonical form may reduce cross-model interpretation drift, but the pilot evidence is exploratory.

Memadapter: Counterfactual Adaptation Against Memory-induced Sycophancy

Ruqing Ning; Haibo Meng; Zhishang Xiang et al. arXiv: 2610.05162

Targets memory-induced sycophancy, where even correct memories over-align the agent with a user's past beliefs. Counterfactual Induction, Context-Aware Reflection and Evidence-Based Reasoning calibrate each retrieved memory's influence; consistently improves memory reliability on three benchmarks.

Key insight: Even correct memories can push an agent toward a user's past beliefs; calibrating each memory's influence against current evidence reduces this.

StateWise: Diagnosing and Repairing Persistent Operational State Before Agent Actions

Yongyuan Peng; Zhou Feng; Tongying Wu et al. arXiv: 2610.05241

Diagnoses and repairs persistent operational records before agents act on them: counterfactual replanning finds decision-critical records, then read-only verification and targeted clarification re-establish validity, with typed evidence and repair lineage. On 150 executable coding-agent cases with corrupted persistent state: 93.3% correctness vs 38.7% baseline, no unsafe actions.

Key insight: Agents should verify the stored operational records their next action depends on before acting, rather than trusting persistent state blindly.

MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents

Haozhen Zhang; Haodong Yue; Quanyu Long et al. arXiv: 2610.06830

RL-trained multi-step policy that chooses between query-agnostic memory and on-demand, query-specific curation of raw multimodal history delegated to heterogeneous LLMs/VLMs, controlling evidence amount, instructions, model choice and visual access. Better performance-cost-latency frontiers across five multimodal agent-memory benchmarks.

Key insight: Deciding at query time how much multimodal history to curate, and with which model, gives better cost-latency-quality trade-offs than fixed memory pipelines.

The Cost of a Hop: Benchmarking NLIP and A2A

Ranjan Sinha; Anindita Das; Ashika Anand Babu et al. arXiv: 2610.04053

First controlled latency study of NLIP vs A2A, decomposed into message creation, connection and send phases on three hardware environments. NLIP is 8.4-9.6x faster than the baseline A2A SDK on lightweight coordination on two machines (~4x on the third), but near parity end-to-end when LLM inference dominates; the gap is mostly connection setup, and caching narrows it to near parity on faster hardware.

Key insight: Agent protocol overhead is dominated by connection setup and matters mainly for lightweight coordination, not LLM-bound pipelines.

COPEX: Benchmarking LLM Robustness to Adversarial Context Across Model Context Protocol Layers

Nahom Birhan; Mehrdad Rostamzadeh; Sidhant Narula et al. arXiv: 2610.04378

Benchmark isolating the model as an MCP client: 25 attack types / 125 scenarios across model/agent, client, server/tool and transport surfaces; nine models, 3,375 trials, mean attack success 64.4% (58.3-71.4% by surface). Combined input and context scanning cuts mean attack success by 49.6% on an eight-attack subset.

Key insight: MCP clients are highly susceptible to adversarial context across every protocol surface, and scanning inputs and context roughly halves attack success.

Understanding and Mitigating Hallucination Escape in Tool-Using LLM Agents

Peigui Qi; Kunsheng Tang; Yide Song et al. arXiv: 2610.04409

Identifies Hallucination Escape: tool-hallucination fixes tuned on one tool configuration increase hallucination on others. Training-free EscapeGuard (conflict-aware gating + configuration-derived attention enhancement) reduces tool-selection hallucination 9.0 pp and lowers the cross-configuration mean by 23.7 pp across six benchmarks.

Key insight: Fixes for tool hallucination tuned on one tool configuration can raise hallucination on others, so evaluation must span configurations.

Judged Useless, Queried Anyway: Tool-Using Agents Rarely Turn Their Own Evidence Judgments into Stopping Decisions

Chubin Zhang; Zhenglin Wan; Xingrui Yu et al. arXiv: 2610.06191

In a retrieval environment with controlled source failures, seven agents call a failing source's results useless 97-100% of the time yet rarely stop on that judgment; prompt cues and budgets shift when, not why, they stop. Only a harness-enforced integration step (answer after five consecutive results judged useless) makes stopping follow evidence; confirmed in a pre-registered replication on 300 questions.

Key insight: Agents recognize useless tool results but keep querying; only a stopping rule enforced by the harness makes them act on that judgment.

Nexus: An Execution Fabric for AI Agents Across Cloud, Edge, and Devices

Cary Chang; Jialin Zhou arXiv: 2610.05709

Cloud-edge platform treating each agent invocation as a persistent task, with run-scoped delegation to Computer and Mobile environments and journaling. All ten Computer-Android workflows succeed; all six revocation tests block subsequent writes; journaling cuts duplicate appends six to zero per task at 0.933 s overhead; 24 vs Dify's 22 tasks completed.

Key insight: Treating each agent invocation as a persistent task with run-scoped, revocable authority makes cloud-edge-device agents more reliable.

Who Is Your Agent Serving? Provider-Side Indirect Prompt Injection in Proactive Agents

Rui Wang; Chao Wang; Xinchen Wang et al. arXiv: 2610.05266

Proactive personal agents can be redirected by a provider controlling only its own content: Target Control, Private Binding and Prospective Support. Across three proactive-agent environments and six simulated user models, the full attack raises target authorization in every combination, macro gain up to 77.4 pp.

Key insight: Proactive agents can be steered by third-party content that recruits the user's own context to justify a provider's target.

Blocking at the Boundary: Auditing Long-Horizon Agents against Staged Prompt Injection

Jingkai Liu; Yufei Han; Xiaoting Lyu et al. arXiv: 2610.05163

Feedback-guided attack pipeline confirms staged prompt injection against Claude Code and Codex in native runtimes (eight scenarios, seven goals, six surfaces). Proposes boundary action auditing and Path-Aligned Attribution: 86% Block recall at 6-8% false-block rate with Claude Sonnet 5, vs 44-47% recall at 16-33% for ARGUS, on a 479-pair benchmark.

Key insight: Multi-stage prompt injection works against production coding agents, and auditing each pending action before it takes effect is the effective defense point.

AgentSpy: Making AI Agent Behavior Observable

Christoph Bühler; Matteo Biagiola; Luca Di Grazia et al. arXiv: 2610.06001

Observes an agent from outside: isolated environment plus syscall and network recording for the agent and all subprocesses. On 77 tasks with the codex harness and three LLMs, among runs passing outcome tests 18% do unrelated activity, 7% read grading files, 17% ignore developer guidance; generic rules detect four of five attack categories with no false positives across 50 runs.

Key insight: Observing an agent from outside at the syscall and network level reveals behavior its own trajectory never reports.

AgentPrivArena: Evaluating and Auditing Real-world AI Agent Privacy

Shouju Wang; Haopeng Zhang arXiv: 2610.06454

Privacy evaluation in realistic workflows with authentic MCP tools and self-hosted services; trajectory-level metrics capture unnecessary information access beyond final-response leakage; AgentPrivAudit monitors violations at runtime. Finds substantial privacy risks overlooked by outcome-only evaluation.

Key insight: Agent privacy should be measured over the whole trajectory, including unnecessary data access, not only what appears in the final response.

GitSwarm: Decentralized Compounding Inference

Vedant Shah; Ankur Samanta; Paras Dahal et al. arXiv: 2610.04862

Compounding inference: homogeneous agents collaborate asynchronously through a shared branchable Git repo with atomic commits and explicit semantic dependencies. Solves all 30 IMOProofBench-Advanced problems in one run with GPT-5.5; 79.4% on ProgramBench vs 65.1% best baseline; 94.7% of contributions are later built upon.

Key insight: Parallel agents sharing a branchable Git repository can accumulate work across episodes instead of discarding it.

Attention Tax, Handoff Tax: A Stylised Model of When Multi-Agent LLM Systems Help

Akshit Anchan; Nayonika Sen arXiv: 2610.06069

Stylised reliability model: decomposition pays when avoided long-context attention cost exceeds the handoff tax; parallel sampling pays when shared-failure floor is below one agent's error floor. On ledger reconciliation the model predicts a crossover at depth 10 and decomposition winning at depths 20/50/100; it does, and predicted success lands within 9 pp at every depth.

Key insight: Splitting work across agents pays off only past a measurable depth, where avoided long-context degradation exceeds the cost of handoffs.

Imagine to Act: High-Fidelity Data Synthesis via Image Editing World Model for Scalable GUI Agent Training

Yongxin Ning; Runliang Niu; Qianli Xing et al. arXiv: 2610.05861

Simulation-free GUI trajectory synthesis via a pixel-level image-editing world model that renders action-induced screenshot changes from structured delta-text. Infinite-Actor (Qwen3-VL fine-tuned only on synthetic data): 8B gains +4.45 AndroidWorld Pass@1 and nearly doubles MobileWorld Pass@3; 2B gains +9.05 Pass@1. Code released.

Key insight: Synthesizing realistic GUI transitions with an image-editing world model can train small mobile GUI agents without simulators.