Thursday's cs.AI announcement day (2026-10-08) lists 110 new and 172 cross-lists (replacements skipped; listing total 282). A filter for agent systems, memory/context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent, persistence/identity, and local/open models keeps 37 papers. Memory is the day's clearest thread: CASK keeps track of who or which reasoning branch owns a claim before it is committed, a personal-memory study shows most stale and misattributed facts are decided when memory is built rather than retrieved, and Budgeted Flat Reconstruction retrieves until the question's information needs are covered instead of stopping at the most relevant hits. Skills and self-improvement get a reality check: the same Skill helps under some model and harness setups and hurts under others on over a third of tasks, and keep-if-better improvement loops overstate their gains when the selection set is small, even as SkillSandbox, SkillForge and FreeEvolve propose ways to verify, retire and evolve skills. On tools, local models and safety, agents catch explicit tool errors but miss plausible wrong values, small local models match hosted ones at most call sites of a deployed home-automation agent but fail unsafely at actuation, and Secure-CUA, CredLeakBench, WebMirage and PackHallu map how computer-use and coding agents can be secured or attacked.
Hongyu Gu; Xinchang Li arXiv: 2610.09008
Persistent memory turns a rejected plan, a simulated tool result or another speaker's belief into a durable 'fact' once it is stored as a plain sentence. The authors call the missing context discourse ownership (which world, branch or speaker licenses a proposition) and find LLMs already carry a causally active ownership signal that memory interfaces throw away. CASK (Causally Anchored Scoping Keys) is a commit rule that keeps the stable relations expressing ownership, so shared-world facts enter durable memory while provisional content stays in its original scope. On controlled long-conversation conflicts and tool-agent traces it improves memory admission and stops provisional content contaminating later answers.
Key insight: A memory write should carry who or which branch owns the claim; storing bare sentences is what turns rejected plans and quoted beliefs into false facts.
Haonan Deng; Park Sinchaisri arXiv: 2610.10265
Measures personal-memory failures before generation using a Personal Fact Memory reference layer. Serving only the active value of each correctly keyed slot eliminates stale exposure; without update resolution 70.3% of prompts expose a superseded value. Participant-aware BM25 matches the reference ranker within 0.02 once retrievers share the active store. The hard part is assigning revisions to the right slot: four LLM key assigners beat a rule extractor on key recall but give lower clean-retrieval rates, and open-domain merge recall on LongMemEval never exceeds 0.062. Same-name speakers in LoCoMo cannot be told apart; prompt prefill, not retrieval, dominates turn latency.
Key insight: Most personal-memory errors are decided when memory is built, so serving only the active value per slot removes stale exposure while slot assignment remains the hard part.
Yufeng Li; Shuxin Li; Zhenhua Xu et al. arXiv: 2610.09348
Top-ranked memories can each be relevant yet jointly miss a fact the answer needs, especially on multi-session and temporal questions. Recasts retrieval as building a sufficient memory set: Budgeted Flat Reconstruction first uses Formal Concept Analysis (FCA-MS) to decompose the question into information requirements and pick a covering subset, then keeps acquiring unseen records via deeper text search or entity/session views until budget runs out. Over a flat store it beats same-store adaptations of recent agent-memory systems on LoCoMo and LongMemEval-S; on LongMemEval-S judged accuracy rises 72.4% → 82.2% with Turn Hit 91.4%.
Key insight: Retrieving the most relevant memories is not the same as retrieving enough of them; covering the question's information requirements lifts LongMemEval-S accuracy from 72.4% to 82.2%.
Jinheon Baek; Soyeong Jeong; Yumin Choi et al. arXiv: 2610.10444
For agents that build deliverables directly from workspace files via a terminal (direct workspace interaction), nothing tracks what the task still requires, what was read, and what was listed but never opened. RunningTab adds an environment-side tab: the agent registers requirements; the environment logs every file read as an excerpt with provenance and every listed-but-unopened file as a candidate; each requirement is shown beside best-matching excerpts and top unopened candidates, and a finish check returns any still-open requirements. Beats plain DWI and in-model record baselines on three benchmarks with three LLMs.
Key insight: An environment-side record of requirements, reads and unopened files keeps workspace agents from finishing with parts of the task missing.
Yu Cheng; Dehai Zhao; Zhongxin Liu et al. arXiv: 2610.08875
87 SkillsBench tasks, utility = pass-rate difference vs No-Skill under the same model-harness configuration, across nine configurations, with candidates from a 37,596-Skill marketplace corpus. The same Skills help some configurations and hurt others on 36.78% of tasks; recommended procedures can become an execution burden. Relevance rankings miss more useful candidates; reranking by support for required operations raises first-choice pass rates 4.35-5.80 pp. Stage Plan and Dependency DAG organizations beat use order alone (DAG helps most with 5-6 Skills). Derives 17 authoring practices.
Key insight: The same Skill helps under some model and harness setups and hurts under others on over a third of tasks, so Skills must be evaluated per configuration.
Litao Hu; Yutong Tang arXiv: 2610.09239
Treats keep-if-better as selection under noise. When Qwen models rewrite their own instructions with 600 held-out items scored, most proposals after the first are harmful. In a pre-registered study, final selection-set scores of greedy loops exceed held-out accuracy by 13-20 points with 16 selection items and 1-5 points with 256; tested acceptance rules did not beat greedy acceptance. The same inflation appears in GEPA and MIPROv2 validation scores. Scoring start and current instruction on 64 never-used items removes average bias, but single estimates stay ~6 points off.
Key insight: Self-improvement loops that keep whatever scores best on a small selection set overstate their gains, by 13 to 20 points with 16 items.
Sihao Liu; Ligeng Zhu; Zijian Zhang et al. arXiv: 2610.08900
Multi-agent orchestration workflow for agentic coding built on 'judgement engineering': a human approves a plan contract, a builder agent implements in rounds, and a reviewer agent from another vendor decides completion; deterministic hooks, not a model, route work and enforce 72 mechanical gates. 68 versions in 108 days, 1,468 GitHub stars, 118 public postmortems; applications include a 567-file gem5 build migration, top-three placements in all three MLSys 2026 FlashInfer Full-Agent tracks, and 672/672 on PutnamBench. Postmortems show independent review catches unsupported builder claims but stopping is a weakness: two thirds of rounds came after implementation was accepted. Observational, not controlled.
Key insight: Mechanically enforced role boundaries and a reviewer from another vendor make agentic coding reliable, while deciding when to stop remains the weak point.
Obada Kraishan arXiv: 2610.10062
Fault-injection layer over a function-calling benchmark: one of four typed faults per trajectory, 1,920 trials, six models over 24 multi-step tasks. Agents flag a problem in 91.3% of trials on explicit errors but only 58.8% on plausible wrong values (vs a 26.8% false-alarm rate). Reasoning models notice less (-9.3 pts) and change plan more (+10.4) than instruct siblings with no recovery gain. Fault-free reruns end in the same state only 63.3% of the time; only a missing tool clearly lowers recovery (39.9%). A prompt line asking the agent to check each result did not move detection.
Key insight: Agents react to explicit error signals but miss plausible wrong tool outputs, and a prompt asking them to check results does not help.
Panagiotis Kasnesis; Christos Chatzigeorgiou; Lazaros Toumanidis et al. arXiv: 2610.09021
Evaluates 9 models (0.8B to a frontier hosted model) across the five call sites of the open-source Wactorz home-automation framework with production prompts and two real Home Assistant installs (280 cases, 2,520 scored calls). Capability ordering differs by site; a 4B model is worse than its 2B sibling at grounded actuation. The best local model is statistically indistinguishable from both hosted models at four of five sites; only code generation separates them. Gemma4 E2B actuates on 87.2% of requests for devices the site does not own. Routing each site to its best local model reaches 91.8% vs 95.4%; hosting only the two generative sites matches hosting everything (39/43) for 28% of the spend. Benchmark, harness and records released.
Key insight: Small local models match hosted ones at most call sites of a deployed home-automation agent, but can fail unsafely at actuation.
Masaaki Nakatsu; Reno Wang arXiv: 2610.09772
Separates logical inference from persona expression on one INT4 base model with hot-swappable LoRA adapters: the logic path sees only the core turn and emits a verifiable structured Micro-State; the persona path renders it with full history. On an Apple M2 laptop (Llama-3.1-8B-Instruct and Gemma-3-4B-it, 4-bit; 480 runs) the logic prompt stays at 180/167 tokens and outputs are byte-identical across pollution levels, while the mixed single pass degrades (logic score 0.669 → 0.150 on Llama). With pollution fed into the logic path, the 8B still holds up but the 4B collapses. Costs one extra decode on a topic's first turn (28.2 s vs 18.2 s); persona swap takes 1.7 ms. Code, adapters and logs released.
Key insight: Keeping the logic call's context minimal and separate from persona rendering makes small edge models immune to history pollution.
Michael Ofengenden; Maksym Andriushchenko arXiv: 2610.09944
Benchmark (222 tasks from 18 sources: coding, computer use, agentic work, automated research) on whether agents can work for a requested duration, predict runtime and estimate elapsed time. Fable 5.1 in Claude Code deviates from requested runtimes by a typical factor of 2.9x vs 1.2x for GPT-6 Astra in Codex, but 14 of 158 reviewed Astra runs explicitly slept after appearing to finish. Runtime forecasts overestimate; removing temporal information more than doubles retrospective error for Sol and Astra.
Key insight: Agents that can finish tasks still cannot reliably work for a requested duration or estimate their own runtime, and some pad time by sleeping.
Sarthak Choudhary; Mihai Christodorescu; Ashish Hooda et al. arXiv: 2610.09469
Formalizes security requirements for both a computer-use agent's decisions and their GUI execution. Before seeing untrusted content the agent commits to a per-action program (action transaction) that fixes its queries to untrusted content and the permitted uses of the answers; untrusted regions are masked, an isolated query model answers, and the target is located on the masked interface. Secure by design under its model; on 400 WebArena tasks (three frontier models, 5 seeds, 6,000 traces) task success is 53.55% vs 55.12% for vanilla and 13.17% for CaMeL-CUA.
Key insight: Committing to an action plan before reading untrusted screen content secures computer-use agents at a small utility cost.
Rafid Ahmed; Joseph Fioresi; Mubarak Shah et al. arXiv: 2610.08871
Benchmark for agents automating email, social media, banking and bills when facing phishing and identity verification, covering user-directed login and autonomous inbox monitoring; pairs phishing scenarios with legitimate counterparts and measures leakage by actual submissions in a sandbox. All tested models leak, including during autonomous inbox monitoring with no login request; most mitigations that reduce leakage also hurt genuine tasks.
Key insight: Every tested agent leaks credentials to phishing, even while passively monitoring an inbox, and most defenses also break genuine tasks.
Yupu Wang; Zhengyuan Jiang; Reachal Wang et al. arXiv: 2610.09264
Package hallucination attack: injected prompts in community-shared rule files (such as .cursorrules) make coding agents swap legitimate dependencies for attacker-controlled packages. PackHallu evolves the injected prompts with trajectory-level feedback and LLM-guided mutation, reaching high attack success and strong transfer across models and agent frameworks.
Key insight: Poisoned rule files can make coding agents install attacker-controlled packages.
Egor Pakhomov; Erik Nijkamp arXiv: 2610.09193
Runs MemoryAgentBench's own Conflict Resolution rule (newest statement wins) as a zero-learning resolver: it answers 80.25% of questions under the official metric (74.5% on three held-out lists). 67 items have gold answers the last-write graph cannot reach (e.g. gold 'New Delhi' after 'The capital of India is Grosseto.'), a third of the multi-hop questions at 262K. Long-context models and a BM25 agent score 84.7%/82.6%/41.6% on rule-solvable items vs 10.4%/11.9%/6.0% on those 67.
Key insight: A frozen newest-statement-wins rule scores 80.25% on MemoryAgentBench conflict resolution, and many remaining items have gold answers that rule cannot reach.
Peter Baile Chen; Geoffrey X. Yu; Xinming Liu et al. arXiv: 2610.10091
Experience layer that stores what prior reasoning traces learned about specific artifacts in a shared corpus: single-artifact experiences (an artifact's contribution to past tasks) and artifact-pair experiences (structural relations found during reasoning). Integrated as lightweight middleware with experience retrieval, it raises answer quality by up to 11.0 points and cuts online dollar cost by up to 50.5% across corpora and search frameworks; experiences transfer from text-to-SQL to factoid QA on the same corpus, and a stronger model's experiences lift a weaker one.
Key insight: Remembering which artifacts mattered for past tasks lets agents find the right evidence faster, improving quality by up to 11 points while halving cost.
Ran Zhou; Xueming Han; Jiaheng Liu et al. arXiv: 2610.10071
Hybrid graph memory for on-device personalized agents: a lightweight self-enhancement classifier routes memories (fewer large-model calls), episodic/semantic/procedural memories are stored as graphs, and working memory is extracted as a state trajectory that captures current state and implicit constraints. On PAL-Set solution selection it reaches an S-score of 35.58, nearly 7 points above the strongest baseline. Code and data released.
Key insight: Routing memories by type with a small classifier and storing them as graphs makes personalized agent memory practical on device.
Yilun Hao; Krishna Sayana; Isabella Ye et al. arXiv: 2610.10507
Treats evidence construction as a sequential decision over retrieval and computation: a small RouterLM picks or specifies operations (filter, aggregate, compute), a frozen CompilerLM turns custom operations into code, and a frozen AnswerLM answers once the router deems evidence sufficient. RouterLM trained with SFT then GRPO. Mean success 75.6% across six benchmark families, +15.9% over the strongest large-model baseline; a trained Qwen3.5-9B router beats a training-free Gemini 3.5 Flash router by 5.0%; +15.0% on three held-out benchmarks.
Key insight: A small trained router that decides which retrieval and computation steps to run can beat large models that simply retrieve.
Mayur Akewar; Ravi Ranjan arXiv: 2610.09088
Defines counterfactual checkpoint advantage: drive a checkpoint branch and a skip branch to the same failure and recover both under matched conditions. On 12 SWE-bench Verified tasks and 106 real recovery branches, checkpointing saves 49.4 s per task, but split into two regimes: the first checkpoint returns 100.0 s on 156.5 s of protected work (0.64 conversion), a second one step later returns -1.1 s (-0.03). Recovery re-derives rather than replays, so preserved work is a poor guide to saved work.
Key insight: The first checkpoint before expensive work pays off; a second one immediately after returns almost nothing, because recovery re-derives work rather than replaying it.
Serin Kim; Kwangwook Seo; Dokyung Song et al. arXiv: 2610.10088
Verifies whether a distilled skill is reusable beyond its source experience by synthesizing a new skill-relevant task and environment: a Proposer fixes conditions to preserve and details to vary, a Builder constructs an executable scenario, and a Verifier compares runs with and without the skill on executability, utility and efficiency to Keep or Reject it. Strongest downstream performance and better efficiency on ALFWorld and WebShop with three models.
Key insight: Testing a learned skill on a freshly synthesized task, with and without the skill, is a reliable filter for what enters a skill library.
Yuyao Ge; Yiwei Wang; Yuchen He et al. arXiv: 2610.09832
Agentic RL method where the skill library and model co-evolve through a fitness-driven lifecycle (trial, active, stable, retired). A pre-RL phase uses base-model rollouts to pre-retire low-fitness skills and seed SFT; RL then continues retirement, stabilization and LLM-guided mutation. Highest aggregate success across interactive benchmarks, up to 7.8% relative over the strongest baseline with a compact library. Releases SkillFurnace (5k+ records with retirement events and failure categories).
Key insight: Giving skills an explicit lifecycle, including retirement, keeps libraries compact and improves agent success.
Yifei Lu; Cheng Liu; Dianzhi Yu et al. arXiv: 2610.10164
A shared policy acts in the environment and proposes skillbank edits (Add, Update, No Edit). Proposals are scored by contrastive action feedback: how swapping in the proposed skill changes the actor's log-likelihood gap between earlier successful and failed trajectories, avoiding extra rollouts; skill-edit support regularization preserves exploration. 98.4% success on ALFWorld and 84.7% on WebShop, still effective with a smaller backbone. Code released.
Key insight: Skill edits can be scored by how they change the actor's preference between past successes and failures, without new rollouts.
Lecheng Kong; Like Hui; Nikos Kanakaris et al. arXiv: 2610.09197
Automates not just prompts/skills/workflows but the evolution loop itself: given goal, target agent, evaluator, data and resource limits, the evolver decides what to test, how much evidence to gather, which candidates to pursue and when to stop, following an editable evolution skill that is meta-evolved. On tau3-bench, ARC-AGI-2, ARC-AGI-3 and Terminal-Bench 2.1 it improves the held-out metric by 13.6 points on average and matches or beats hand-designed evolvers; meta-evolved skills add 6.9 points over the seed skill on fresh target agents.
Key insight: Letting the evolver decide what to test and when to stop, guided by a meta-evolved skill, matches or beats hand-designed evolution loops.
Harman Singh; Anton Bakhtin; Rulin Shao et al. arXiv: 2610.08902
Measures how efficiently agents convert experience into held-out gains when they amortize past experience into reusable artifacts inherited by future instances. Frontier models show sharply different trajectories with comparable learning opportunities; in-regime gains transfer only partly out of distribution; the best final agent is not necessarily the most efficient learner. Low-plasticity agents fail to reuse relevant artifacts; more plastic ones fail despite reuse, pointing at artifact quality or application.
Key insight: How efficiently an agent turns experience into held-out gains varies sharply across frontier models and is not the same as final capability.
Sergei Polezhaev; Barys Liskavets; Ori Press et al. arXiv: 2610.09858
Trains a natural-language critic with RL from whether the frozen agent ultimately succeeds after receiving its advice, with no step labels or reference critiques. A Qwen3-4B critic trained on multi-hop QA with one base model improves four base models (three unseen); on MuSiQue it lifts Qwen3-4B by more than 25 points, surpassing Kimi K3 without a critic, and transfers to tau3 and DeepDive without extra training. Agents can decide when to ask the critic.
Key insight: A small critic trained only on whether the agent eventually succeeds gives advice that transfers across base models and tasks.
Hanwen Li; Jinhao Duan; Guanhua Zhu et al. arXiv: 2610.09115
Builds a stepwise outcome table of about 82,000 counterfactual continuations from 1,864 trajectories (three benchmarks, two agents, four steering mechanisms at every step). Uncertainty identifies failing trajectories but no single signal locates the step where steering helps. VoS learns the value of steering per step and uses a harm-budgeted trigger: +7.8 points over unmodified execution on average across all 12 settings, and beats the strongest of five uncertainty-triggered methods in 11 (+2.9 average).
Key insight: Uncertainty flags failing trajectories but not where to intervene; learning the value of steering per step does.
Xinwen Liu; Zhuocheng Pan; Isabella Zhu et al. arXiv: 2610.10126
Shows communication topology is strongly associated with failure type (chi^2 = 409.9, p = 1.2e-70). MAScope recovers topology from traces by grounding an interaction graph in message evidence, then a topology-conditioned judge classifies failures. On 851 MAST-clean traces, ground-truth topology raises gpt-mini Macro-F1 0.173 → 0.350; predicted topology gives 0.346, near trace-only gpt-5.4 (0.372), at about 6% of repeated gpt-5.4 diagnosis cost for 1,000 traces.
Key insight: Knowing a multi-agent system's communication topology roughly doubles a small judge's failure-diagnosis accuracy.
Ali Asaria; Deep Gandhi; Tony Salomone arXiv: 2610.10468
Position plus system: populations of thousands of persistent research agents sharing compute should get explicit institutions. Principal investigators compete for compute via requests for proposals, independent review and grants; a human 'mayor' allocates resources but assigns no tasks. In a running society of ten thousand researchers asked to improve LM pretraining, one lab reported about 30% less compute for the same quality, a result other labs that tested it do not yet agree on. Lists six open problems.
Key insight: Large populations of persistent research agents will organize themselves anyway, so designers should provide explicit institutions.
Xijie Gong; Tingxu Han; Jiahao Zhang et al. arXiv: 2610.09624
Converts agentic prompts into minimal contrastive pairs where one request verb flips the tool-call decision (write vs discuss; 500 pairs across Python, Java, C++). Traces the decision to a vector that is causally necessary and sufficient and generalizes to native multi-turn tau2-Bench trajectories. The scaffold sets a tool-call prior; analysis verbs suppress it via features signaling tool use is unnecessary. Same mechanism across seven Qwen, Mistral and Granite models. Code released.
Key insight: Whether an open model calls a tool is governed by a single internal direction that analysis verbs suppress.
Veronique Ziegler arXiv: 2610.09037
In a controlled file-recovery environment where stronger regulation triggers imposed tool failures, a cost-blind supervisory governor can turn failures into persistent blocking. A backoff rule that lowers intervention probability from a moving average of induced events improves completion; in 576 Gemini 2.5 Flash episodes over 6 tasks backoff reduces blocking, with intermediate strength best. Costs are imposed and observable; generality is open.
Key insight: A supervisor that blocks tool use without accounting for its own cost can stall the agent it regulates; backoff restores completion.
Wanjing Han; Levi Taiji Li; Mu Zhang et al. arXiv: 2610.09240
End-to-end grounding-to-execution attack on vision-grounded web agents: localized visual perturbations make agents select attacker-controlled content and execute the matching browser action across varied renderings, using a role-slot abstraction, webpage recomposition and dataflow analysis aligned with action post-processing. Across four agent configurations, six VLM backbones and 2,250 tasks (13 public sites + sandbox): 91.9% average attack success vs 17.4% for the strongest baseline, and effective against three agent-level defenses.
Key insight: Localized adversarial images can steer web agents all the way to attacker-chosen browser actions, with 91.9% success.
David Dobre; Leo Schwinn; Gauthier Gidel et al. arXiv: 2610.09027
Adversarial images can be optimized so their backdoor persists through the KV cache of later tokens after the image is masked from attention (P-VMI). On Qwen3-VL-8B-Instruct, up to ~90% target success, still effective when the image is shown only on the first turn; a cache-swap ablation localizes the effect to the KV cache, and attacks can survive compaction that keeps a summary's KV cache.
Key insight: Adversarial influence can persist in the KV cache after the malicious image has been removed from context.
Nikolaos Kekatos; Dimitrios Nikou; Anastasios Temperekidis et al. arXiv: 2610.08928
Decomposes each natural-language policy clause into a typed spatial/temporal/semantic triple over one event stream, monitors each aspect separately (past-time aspect on unmodified MonPoly) and fuses verdicts with a four-valued provenance algebra. Across two defence mission domains and one civil domain, composition under precautionary blocking drives attack success to zero with no observed false positives at microsecond per-event cost, while any single aspect or pair leaves many attacks succeeding.
Key insight: Checking spatial, temporal and semantic aspects of a policy together stops tool-attack chains that any single monitor misses.
Param Biyani; Krishnamurthy Dvijotham arXiv: 2610.09159
Before an agent acts, autoformalizes task intent into a Lean 4 specification, formalizes tests independently, and has the Lean kernel check whether any implementation can satisfy both, issuing a machine-checked certificate when none can. On conflicted SWE-bench tasks it detects up to 72.8% of conflicts and certifies up to 51.1%, with a nearly five-fold lower miss rate than model-based judgment. Code released.
Key insight: Formal checking can prove a coding task's description and tests contradict each other before an agent is tempted to cheat.
Taehyeon Yun; Dongho Kim; Geonwoo Kim et al. arXiv: 2610.09581
64 mechanically checkable cases from saved AgentDojo Banking executions: with original identifiers and no binding table, Sonnet recovered every literal source location but made unsupported citation-source assertions in 26 of 28 cases needing the missing relation, though 22 still matched the reference answer. Explicit identifier-to-record bindings improved grounded reconstruction; renaming identifiers alone did not; a deterministic comparator resolved or abstained correctly throughout.
Key insight: Correct answers from agent logs are not supported findings unless the link between citations and records is preserved.
Richard Abrich arXiv: 2610.09142
Source-informed mutation tests on shipped verifiers: in AppWorld, duplicating a non-idempotent write creates an extra record yet passes (6/15 constructed effects); a cardinality patch on checker copies makes all six fail. In WorkArena, 21 of 23 extra-field candidates with confirmed nondefault persisted values still PASS. No population rate is estimated.
Key insight: Shipped AppWorld and WorkArena verifiers accept some runs with wrong side effects, such as duplicated writes or extra fields.
Igor Slinko; Yaroslav Golubev; Sergey Titov arXiv: 2610.09633
From 4,782 real JetBrains IDE agent sessions, long sessions mix questions, planning, review, refactoring and execution and switch between them, unlike issue-derived benchmarks; public corpora differ from each other too. SWE-TaskFlow transforms issue benchmarks toward a target task flow via prompt splitting and verifiable repository QA. On 700 SWE-Bench Pro tasks, solving in several steps roughly doubles cost without a stable change in resolve rate.
Key insight: Real coding-agent sessions mix task types, and splitting tasks the way users do roughly doubles cost without changing resolve rate.