Wednesday's cs.AI announcement day (2026-10-07) lists 156 new and 194 cross-lists (replacements skipped; listing total 350). A filter for agent systems, memory/context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent, persistence/identity, and local/open models keeps 41 papers. The day's clearest thread is who owns an agent's context: ReFold folds repeated tool output and finished turns reversibly at the rendering layer, Stateless Language Agents keep all durable state in the harness and rebuild a fresh context for every call, and SquidAgent and NP-Bench show parallel agents only pay off when context, conventions and file scopes are settled before work starts. Memory papers focus on what should be written at all: AgentMemGate keeps a user's tentative plans out of the stored profile, a retention study finds inferred values and beliefs far less reliable than other facts, and MINDSET and PERSIST treat memory as versioned, speaker- and time-aware state. On safety and on-device agents, APEX and a gate-stacking study favor deterministic checks at the action boundary over stacked LLM judges, When Tools Lie shows verification has to be mandatory, STEPGATE escalates only uncertain steps from a small local model, and nanoMuse sketches an open personal agent across phone and computer.
Yupeng Su; Jiayi Tian; Zheng Zhang et al. arXiv: 2610.07863
Training-free rendering layer that keeps the full interaction history but compresses only the rendered context: content an earlier turn already displayed becomes a stub, and turns the agent reports finished fold into a one-line note. Chunked rendering rewrites the cached prefix only every few steps; every removal is reversible from history. Across five long-horizon benchmarks and two frontier LLMs: up to 2.5x fewer tokens, half the KV-cache per session, no drop in success; avoids up to 92% of forced compactions under capped budgets; up to 1.7x faster and 3.4x cheaper under concurrent serving.
Key insight: Reversibly folding repeated tool output and finished turns at the rendering layer cuts context cost sharply without the information loss of summarization.
Qizheng Zhang; Changxiu Ji; Isaac Sun et al. arXiv: 2610.07625
'Stateful search with stateless agents': no agent carries its conversation across invocations; the harness owns research state (candidates + measured outcomes) and rebuilds a fresh role-specific context each call. A stateless Advisor assigns experiments to parallel Workers. At budgets up to one billion tokens on SWE, kernel optimization and algorithm design, SLA gets the best final result on every task and matches the strongest kernel baseline with over 84% fewer tokens; the Advisor uses under 0.6% of tokens.
Key insight: Long-running agents make more progress when the harness owns the research state and each agent call starts from a freshly built, role-specific context.
Yexiong Lin; Shanshan Ye; Yu Yao et al. arXiv: 2610.08647
Parallelize a layer only when critical-path cost plus re-exploration and alignment overheads beats serial cost, measured in predicted output tokens (which LLMs estimate more reliably than wall-clock time). Workers fork directly from the orchestrator's session (no re-exploration) and share a pre-generated convention block (no post-hoc reconciliation). 2.2x mean throughput and 2.6x mean wall-time speedup over Claude Code; 2.0x throughput over the strongest multi-agent baseline.
Key insight: Parallel agents only beat a serial agent when they inherit the orchestrator's context and agree on conventions before they start.
Sumanyu Muku arXiv: 2610.07261
Recasts parallel coding-agent coordination as up-front scheduling: partition declared scopes into disjoint sets and order merges along the producer→consumer graph. Clean integration rises from 1/9 to 9/9 scenarios and merge conflicts fall 13 → 0; on a live breaking contract change clean integration goes 0 → 1.0 (frontier) / 0.6 (small model) with 0/5 scope leakage. A cross-session memory drops the repeated-mistake rate 1.00 → 0.00. Negative result: routing facts to agents does not rescue long-context accuracy at window-fitting scales.
Key insight: Parallel coding agents integrate cleanly when their scopes and merge order are scheduled up front instead of conflicts being detected afterward.
Sumanyu Muku arXiv: 2610.07257
Local runtime for fleets of parallel coding agents on one workstation that turns resource governance into checkable signals: per-agent memory attribution, complete subtree reclamation, escaped-process detection, bounded footprint. Under a binding budget it keeps the fleet at 7.5 GiB with zero swap while tmux and an agent multiplexer hit 2x budget and spill ~2 GiB to swap; reclaims 100% of a terminated agent's process tree (raw baseline strands half) and detects 10/10 escaped children. Attribution scan costs ~0.6% CPU at one agent, 2.7% at ten; carries over to live Claude Code sessions. Engine released.
Key insight: Running many coding agents on one machine needs per-agent memory attribution and guaranteed cleanup, which terminal multiplexers do not provide.
Hans Schabert; Christoph Peters arXiv: 2610.07817
Delivers a procedure one step at a time from an MCP server; each step returns a structured step_output, so the execution path is fixed in advance and logged. Over 15,475 trials on 13 SOP-Bench domains with four open-weight executors (Kimi K2.5 to Ministral 3 8B): process adherence rises from 76-95% to 95-99%, ungrounded correct answers fall from 2.1-4.5% to 0.2-0.3%, and the 8B executor gains +6.5 pp grounded accuracy. Under prompt delivery, 31-49% of correct know_your_business answers skip the SOP entirely.
Key insight: Serving a procedure one step at a time over MCP makes agent execution predictable and auditable, and helps small executors most.
Chirag Sharma; Benjamin Fowlersmith; Karime Maamari arXiv: 2610.07707
Write-time gate for profile-store memory that classifies extracted statements as speculation, completed event, correction or other; speculations stay out of memory with conditions for later promotion/deletion. On a 147-conversation held-out set, Mem0 and Graphiti record unresolved plans as current state for 35.2% and 27.3% of pending plans. On the core benchmark the gate removes all observed contamination (87.5% → 0) and lifts task accuracy 65% → 95%; on the harder held-out set contamination is 3.4-5.7% and accuracy +9-13 pp. Main remaining gap: speculations that match no profile field never reach the gate.
Key insight: Assistant memory must keep a user's tentative plans out of the stored profile until they actually happen.
Olukunle Owolabi; Pulkit Gupta; Fei Wang arXiv: 2610.07100
In a deployed cold-start memory pipeline (100 synthetic personas, 4,715 candidate assertions), only 77.9% of value/belief assertions are supported by their source vs 96.2% for other categories. A stricter retention bar on values alone cuts unsupported retentions 6.2% → 4.0% (~36% relative) and keeps ~13 pp more coverage than a global threshold at comparable retention.
Key insight: Memory retention thresholds should depend on what kind of claim is being stored, since inferred values and beliefs are far less reliable than other facts.
Xinran Zheng; Xin Fan Guo; Zhiqiang Hao et al. arXiv: 2610.06966
Defends indirect prompt injection at the execution boundary across Tools, MCP servers and Skills: an authorization contract compiled from the trusted task before untrusted execution admits an effect only when justified (evidence-gated prevention), and deception-based exposure makes unendorsed use reveal itself before commit. Against 13 baselines: 0% attack success on five of six benchmarks, 0.56% on the sixth; 0% under adaptive attacks on all three capability-unit types.
Key insight: Prompt injection is best stopped where the agent turns state into an external action, checked against what the trusted task authorized.
Chenglin Yang arXiv: 2610.07359
On 1,119 labelled agent actions, any two LLM judges compose to only ~1.2-1.4 multiplication-equivalent layers (errors correlated, 6/6 pairs significant), while a deterministic rule layer plus one judge composes to ~1.86-2.09 layers (0/4 significant correlation). Solo accuracy does not predict added coverage. One judge tier was served by an unrequested model version in 50 of 112 batches, which overturned a pre-declared analysis rule.
Key insight: Two LLM judges fail together far more often than a deterministic rule layer and one judge, so stacked gates should mix kinds.
Kavienan Jegatheesan; Gayathri Lihinikaduarachchi arXiv: 2610.08097
A hidden interceptor replaces tool results with plausible wrong values. Without verification, accuracy drops from 100% to 72.4% across 31 problems; mandatory same-context reflection recovers to 100%; optional verification helps only when the model invokes it; full restart after explicit detection succeeds in 100% of cases.
Key insight: Agents recover from corrupted tool results when verification is mandatory, but optional verification helps only when models choose to use it.
Abolfazl Younesi arXiv: 2610.07816
Scores each local SLM action's uncertainty and escalates only hard steps to a stronger model. Qwen2.5-1.5B/7B on a 52-task BFCL-derived split: 82.7% success at 30.8% escalation vs 67.3% local-only and 75.4% random escalation. Multi-turn: 69.0% trajectory success with 30.0% cloud actions vs 48.0% local-only, 57.0% query-level routing, 82.0% strong-only. Limited to one model family and scripted tasks.
Key insight: Escalating only uncertain steps from a small local model to a larger one recovers much of the accuracy gap with a fraction of cloud calls.
Guangyi Liu; Yong Liu; Jiangning Zhang arXiv: 2610.08699
Defines the personal agent in five questions and three horizons, reads how Meta's Muse is built from public record and a copy of its production prompt (each statement source-marked), then presents nanoMuse (GPL-3.0): one agent across every device a person owns with hands on phone and computer screens, a shared conversation over a self-hostable relay, every action through a Sentinel, memory as user-readable files, and model of the user's choice. Size and cost are estimates; memory-with-provenance, a hands eval suite and an open hands model are roadmap.
Key insight: An open personal agent can span a person's phone and computer with file-based memory and a sentinel on every action.
Yeji Park; Jaeyun Shim; Taesik Gong arXiv: 2610.07972
PAIR builds user-conditioned app states to test the same task across users. Across six mobile GUI agents, subgoal achievement drops 6.98-15.4 pp in user-conditioned UIs and 8.77-22.0 pp for targets from the user's own content, often by selecting another item. RePAIR RL training on cross-user differences adds +5.87 pp SAR, +7.50 pp all-success and +9.42 pp Task SR over its SFT parent on unseen users.
Key insight: Mobile GUI agents are markedly less reliable on interfaces personalized by a user's own history and content.
Egor Pakhomov; Erik Nijkamp arXiv: 2610.08722
On 590 harness-triggered AppWorld compaction boundaries replayed with and without the summary, pre-boundary agent history predicts post-compaction harm only weakly: best held-out AUROC 0.66 vs a 0.72 same-boundary replicate ceiling. The best interpretable trigger avoids 21% of harmful boundaries while keeping 84% of compaction opportunities. Whether it beats a token-budget rule at matched retention cannot be evaluated on the release.
Key insight: An agent's recent behavior only weakly predicts when context compaction will hurt, so smarter compaction triggers offer limited gains.
Haibo Jin; Xinjie Li; Peng Kuang et al. arXiv: 2610.07832
Dev-Primitives pair each repository artifact with a resident LLM that gives it an agent-native interface; HERMES activates them dependency-aware and maps execution evidence back to components to revise. Beats matched baseline harnesses by 12.4% on average across four SE benchmarks; with Qwen3-8B Dev-Primitives it stays within 4.5% of an all-GPT-5.6 Sol setup while cutting Terminal-Bench 4.0 cost 26.2%.
Key insight: Giving repository components their own small resident models lets a harness localize context so cheap models approach frontier coding performance.
Zihan Zhou; Xinzhe Hu; Hanxu Yang et al. arXiv: 2610.08155
Splits coordination from reasoning: a lightweight 'System One' controller picks inspection conditions, next tasks and termination, and a compact reader retrieves condition-relevant evidence, while capable LLM workers do open-ended reasoning. On seven benchmarks vs AgentVerse, DyLAN and SelfOrg, cuts GPT-4o tokens 44.9-97.2% and latency 37.8-93.0% with better accuracy.
Key insight: Bounded coordination decisions in multi-agent systems can be handed to lightweight models, saving most of the token cost of frontier coordinators.
Sujato Dutta; Sreekruthy Tummala; Shashank Vanga et al. arXiv: 2610.08586
Stores a conversation as immutable episodes organized into versioned schemas via minimum-energy transitions (reinforce, supersede, split, create), with hysteresis so isolated contradictions don't rewrite stable memory. On 850 questions (700 LoCoMo + 150 MemoryAgentBench) it has the highest observed LoCoMo answer F1 and significantly better retrieval ranking than LightMem (p<0.01, Holm); cross-model check with GLM-4.7 and Gemma-4-31B.
Key insight: Long-term conversational memory works better as versioned state with explicit supersession than as repeated summarization.
Achira Lin; Siyuan Hou; Wenyi Yu et al. arXiv: 2610.07725
Persistent multi-session, multi-speaker memory for full-duplex spoken dialogue that models Who (acoustic speaker identity), What and When in readable event records with a 3W joint retrieval score. Reusing backbone representations cuts retrieval latency 578.42 ms → 7.03 ms. On the new SpokenTrace benchmark: 85.08% end-to-end task accuracy; all-support EM@3 49.01% (BGE-large) → 82.10%.
Key insight: Shared voice assistants need memory keyed on who spoke, what was said, and when it changed, not just on topical similarity.
Luoxi Tang; Yuqiao Meng; Nilesh Auradkar et al. arXiv: 2610.07311
Retrieved memories can mislead even when benign and correctly retrieved; failures are strongest under partial query-memory overlap. MEMTRIM indexes memory evidence at write time and controls reuse at read time, removing repeated or conflicting evidence; plug-and-play, no retraining, works for embedding and structured memory. Reduces over-reliance while keeping memory's benefits across models.
Key insight: Correctly retrieved memories can still mislead when they only partly match the current task, so reuse should be trimmed at read time.
Zi Wang; Xingqiao Wang; Emmanuel Addai et al. arXiv: 2610.07309
Labels each memory-query pair admissible, inadmissible or unresolved (other principal, policy, lifecycle state) and tracks memory IDs into prompt exposure. Reanalysis of RHELM and MemOps (3,767 queries): namespace filtering lifts top-20 anchor recall 0.432 → 0.533 and cuts similarity evaluations 98.3%. Text-only verifiers fail to detect violations under a 1% false-denial limit; in 16 controlled exposure scenarios only one of four reader CIs excludes zero for inadmissible literal disclosure.
Key insight: Relevant memories can still be off-limits for a request, and metadata-based namespace checks catch this where text-only verifiers fail.
Zhe Yu; Zixuan Wang; Peidong Wang et al. arXiv: 2610.08101
Shows identical retained records can correspond to compliant and violating executions: removing evidence such as receipt, action dependence or response validity leaves 82.4% of opposite-label pairs indistinguishable; restoring it separates 97.9%. CAVERT extracts supported execution relationships from logs and beats contract-prompted LLM and rule baselines on diagnosis in all 12 benchmark-executor settings and on recovery in all four environments.
Key insight: Correct memory records are not enough to judge multi-agent executions; logs must preserve receipts and action dependencies.
Jike Zhong; Ritwick Chaudhry; Xuanbai Chen et al. arXiv: 2610.08102
Benchmark of 1,000 questions over dense, frequently revised professional artifacts (Current/Past/Derived State, Change History, Conflict/Refusal). Across 27 frontier/open-weight model and memory-method configurations the best scores below 45%. Number of governing updates dominates difficulty; models often fail to check user premises against prior updates; more reasoning effort and memory methods help little, state-aware designs help more.
Key insight: Agents struggle most with memory of documents that are repeatedly revised, and state-aware designs help more than extra reasoning.
Hochan Son; Kyungdoe Han; Jaehan Koh et al. arXiv: 2610.07782
On a three-tier agent architecture, decomposition bounds peak KV working set to 14.3 MiB/query vs ~35 MiB for single-pass and RAG baselines, but the persistent reasoning-trace tier costs +0.368 MiB and gives no detectable accuracy change (+0.015, 95% CI [-0.011, +0.046]). Reaching this took four measurement corrections, three of which inflated the apparent benefit. Gives conditions an agent-memory ablation must satisfy.
Key insight: A persistent memory tier can add cost with no measurable benefit, and common ablations are easily inflated by measurement errors.
Antoine Edy; Max Conti; Victor Xing et al. arXiv: 2610.08048
Bootstraps reusable agent memory without tasks or oracle verifiers: an explorer generates challenging-but-solvable tasks, a solver attempts them, and a heuristic from each failure is accepted only after the solver repeatedly succeeds with it. On AppWorld, τ²-bench and AutomationBench: up to +15.9 points mean success and up to 2.2x pass^5 vs no memory; competitive with methods that use training tasks at lower inference cost; heuristics transfer across model families.
Key insight: Agents can build useful procedural memory for a new environment by practicing on tasks they generate for themselves.
Chinmay Savadikar; Zhaoyu Zhang; Mingyu Zhao et al. arXiv: 2610.07118
Agent jointly learns to reason, act and write free-form memory, but memory is append-only so retention is guaranteed by construction; trainable end-to-end with outcome-reward RL. Trained overwrite memories were found to delete key facts and corrective feedback. On WebArena Lite: +4.09 pp average success over overwrite memory, +4.8 pp tasks solved in all five runs, matching an overwrite baseline trained on much costlier curated supervision.
Key insight: Append-only memory guarantees that long-horizon web agents keep the facts and feedback that learned overwrite memories tend to delete.
Haizhong Zheng; Yizhuo Di; Ranajoy Sadhukhan et al. arXiv: 2610.07792
Evolving-environment streams (retail support, banking, sales pitch; 53 windows, 7,718 tasks) where hidden policies change. Evaluates RAG, Mem0, SkillOpt, Continual Harness and Prime across six models (28 pairs, 252 runs). Findings: a large gap between task capability and learning from experience; continual adaptation is costly and can degrade already-correct behavior; insufficient exploration is a key bottleneck.
Key insight: Agents that learn from serving experience still lag far behind their task capability, and adaptation can break behavior that was already correct.
Yibo Li; Jinhang Qiu; Zhi Zheng et al. arXiv: 2610.08215
Text games with novel or counterintuitive rules so agents must learn from interaction. Findings: retaining complete records of actions and feedback supports learning better than summarizing them into rules; top human players reach higher peaks, explore more and repeat less; with backbone fixed, changing the harness can raise performance while lowering estimated cost.
Key insight: Agents learn unfamiliar environments better by keeping full records of actions and feedback than by summarizing them into rules.
Shaswata Mitra; Raj Patel; Subash Neupane et al. arXiv: 2610.07657
Deterministic-first enforcement: a cascade of 28 checks blocks what it can and refers the rest to a panel of four judges. Across four domains attack success falls from ~30.0% to ~3.0%, with 78% of blocked attacks handled by deterministic checks; in security operations only a quarter of proposals reach the judges. A risk-score approval gate approved most attack proposals but few legitimate ones.
Key insight: Deterministic checks can block most multi-agent attacks cheaply, leaving judges for proposals that misrepresent intent.
Yunju Kang; Seonghyeon Cho; Irene Li et al. arXiv: 2610.08082
Guardrail for small tool-calling agents that scores each action's reversibility by deriving a candidate inverse sequence via a two-layer ontology; calls below threshold are pruned before execution. On τ²-bench across six models it improves airline reward by 0.11-0.18 for four of six agents, but only 8 of 18 model-domain cells improve; retail and stronger agents often regress.
Key insight: Scoring whether a tool action can be undone gives an auditable guardrail, but it improves task outcomes only in some settings.
Hongzhan Lin; Shidong Cao; Ziyang Luo et al. arXiv: 2610.07753
656 cases across six domains and five protocols from static action judgment to dependent multi-action workflows, with a provenance-bound Evidence Ledger and deterministic evaluator. Across ten model-harness configurations, strong static action assessment coexists with weaker interactive execution; failures often start before execution (stopping with incomplete investigation, acting before required evidence). Once evidence is established, single actions are usually reliable; multi-action workflows add unresolved prerequisites.
Key insight: Tool-using agents often fail before they act, by stopping investigation early or acting before the needed evidence exists.
Lizhi Zhang; Xin He; Dianxuan Fu et al. arXiv: 2610.07645
Poisons self-learned skills using only verified successful experiences: build successes that reinforce a target behavior, then strip the conditions that constrain when it applies, so the skill extractor over-generalizes. 95.71% attack success on three benchmarks while every injected experience is task-correct and passes verification and lexical inspection.
Key insight: Self-learned skills can be poisoned using only verified, task-correct experiences by removing the conditions that limit when a behavior applies.
Sarim Hashmi; Mukul Ranjan; Kshitij Mishra et al. arXiv: 2610.08773
Co-evolves a task curriculum, an injection adversary and a 4B web agent inside a frozen web world model; the curriculum is rewarded for ~50%-solve tasks and the adversary only for success-flip injections. The agent becomes more capable and more robust, holds against an unseen frontier-model adversary, and transfers to a real browser: +33.6% relative completion under that adversary on 150 web tasks.
Key insight: Training a small web agent against an adaptive injection adversary inside a web world model improves both robustness and capability on real sites.
Hanjun Luo; Xiucheng Zhang; Zhuoning Xu et al. arXiv: 2610.08662
200 evidence-controlled repository task pairs testing Avoidance/Transfer/Mitigation/Acceptance risk treatments. Across 8 models, unnecessary risk treatment occurs in 11.2-58.7% of runs despite explicit evidence; stronger task capability does not imply more appropriate risk treatment, and violations noticeably hurt developer experience.
Key insight: Coding agents frequently take unnecessary defensive measures, and this risk-treatment skill does not improve with general capability.
Mohammadreza Sediqin; Shivali Dalmia; Srinivasa Karthikeya Reddy Kovvuri et al. arXiv: 2610.07274
Post-hoc layer that checks four properties beside each score: supported by the benchmark's grading logic, earned through traceable computation, completion claim matches what happened, stable under reruns. On 108 Agents' Last Exam tasks across five configurations, every model has passes with no traceable computation and confirmed false completion claims; 18-46% of tasks change score band over five runs; only 22.6% of recorded passes clear all four checks.
Key insight: Most recorded agent benchmark passes fail at least one check of grading support, traceable work, honest completion claims, or rerun stability.
Brendan King; Farima Fatahi Bayat; Jean-Flavien Bussotti et al. arXiv: 2610.07948
Inference-time confidence estimation from a single trajectory without model internals or training data: decompose 'the agent succeeded' into sub-claims grounded in trajectory evidence, score each, aggregate. Across three benchmarks, three backbones and three agent frameworks, better calibration and risk-aware decisions than verbalized, sampling and white-box baselines; a white-box surrogate looked calibrated but discriminated near chance.
Key insight: Breaking an agent's success claim into evidence-grounded sub-claims gives better calibrated confidence than asking the model directly.
Yuhe Hu arXiv: 2610.07781
8-bit vs 4-bit Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct on 20 deterministic tool tasks and five prompts: the recovery comparison flips direction across prompts and scoring targets (Qwen from -50.0 to +35.0 pp). For Llama under one prompt, own-clean-task scoring favors 4-bit by 17.5, matched tasks show no difference, full pipeline favors 8-bit by 28.3 — and strict parsing turns that into -15.0.
Key insight: Whether 4-bit or 8-bit quantization hurts tool-failure recovery depends on the prompt and scoring choices, so single evaluations are not conclusive.
Zhi-Kai Chen; Song-Yan Li; De-Chuan Zhan et al. arXiv: 2610.07086
Slot-parallel speculative decoding for tool calls: future argument slots are generated concurrently as candidates, then verified by the target model under the actual prefix; only verified tokens are committed. Up to 4.05x end-to-end throughput over autoregressive decoding on Glaive and BFCL. Code released.
Key insight: Tool-call arguments can be generated in parallel and verified by the target model, multiplying tool-calling throughput.
Heewon Park; Somin Im; Minhae Kwon arXiv: 2610.07335
Training-free gate that invokes an expensive critic only when action-level ambiguity (global entropy, top-2 margin) suggests high value of information, plus online self-improvement to reduce critic reliance. On ALFWorld, success 24.6% → 78.4% at a ReAct-comparable token budget (3.1x normalized token efficiency); a 7B actor with a 3B critic matches a 14B actor without critique.
Key insight: Calling a critic only at ambiguous steps recovers most of the reliability of always-on critique at a fraction of the cost.
Ankit Sonthalia; Haritz Puerto; Alexander Rubinstein et al. arXiv: 2610.08775
Agents get a whole unlabelled workload and fixed time/compute/API budgets and choose how to 'bottle' it (train a small model, write a program). Across ten models and three tasks, zero-shot strength doesn't predict bottling: 48 of 60 runs score below their model's zero-shot 95% CI lower bound and 31 of 60 underperform a small-model distillation baseline. But Opus 5 keeps ~82% of zero-shot macro-F1 at ~657x lower cost on one task.
Key insight: Agents can sometimes turn expensive model capabilities into cheap reusable artifacts, but strong models often fail to do it well.
Seunghyun Oh; Hirotaka Hiraki; Shuyue Stella Li et al. arXiv: 2610.07506
Real-time proactive speech agent in multi-party conversations that intervenes when the group misses or misstates a fact from a shared document and doesn't self-correct. Uses a small open-weight model with deterministic checks and grounds every claim in a source sentence. On the synthetic CHI-180-proactive set it is correct on most events it addresses and stays silent 97% of the time when the group resolves an issue itself; a 23-participant live study confirms the trends.
Key insight: A proactive speech agent can join group conversations usefully by grounding every interjection in a source and staying silent when the group self-corrects.