Thursday's cs.AI announcement day (2026-10-08) lists 110 new and 172 cross-lists (replacements skipped; listing total 282). A filter for agent systems, memory/context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent, persistence/identity, and local/open models keeps 37 papers. Memory is the day's clearest thread: CASK keeps track of who or which reasoning branch owns a claim before it is committed, a personal-memory study shows most stale and misattributed facts are decided when memory is built rather than retrieved, and Budgeted Flat Reconstruction retrieves until the question's information needs are covered instead of stopping at the most relevant hits. Skills and self-improvement get a reality check: the same Skill helps under some model and harness setups and hurts under others on over a third of tasks, and keep-if-better improvement loops overstate their gains when the selection set is small, even as SkillSandbox, SkillForge and FreeEvolve propose ways to verify, retire and evolve skills. On tools, local models and safety, agents catch explicit tool errors but miss plausible wrong values, small local models match hosted ones at most call sites of a deployed home-automation agent but fail unsafely at actuation, and Secure-CUA, CredLeakBench, WebMirage and PackHallu map how computer-use and coding agents can be secured or attacked.


Research Papers

Whose Memory Is It? Scope-Aware Commit Rules for Long-Term LLM Memory

Hongyu Gu; Xinchang Li arXiv: 2610.09008

Figure from Whose Memory Is It? Scope-Aware Commit Rules for Long-Term LLM Memory
Whose Memory Is It? Scope-Aware Commit Rules for Long-Term LLM Memory

Persistent memory turns a rejected plan, a simulated tool result or another speaker's belief into a durable 'fact' once it is stored as a plain sentence. The authors call the missing context discourse ownership (which world, branch or speaker licenses a proposition) and find LLMs already carry a causally active ownership signal that memory interfaces throw away. CASK (Causally Anchored Scoping Keys) is a commit rule that keeps the stable relations expressing ownership, so shared-world facts enter durable memory while provisional content stays in its original scope. On controlled long-conversation conflicts and tool-agent traces it improves memory admission and stops provisional content contaminating later answers.

Key insight: A memory write should carry who or which branch owns the claim; storing bare sentences is what turns rejected plans and quoted beliefs into false facts.

Stale, Misattributed, or Late: Where Personal Memory Fails Before Generation

Haonan Deng; Park Sinchaisri arXiv: 2610.10265

Measures personal-memory failures before generation using a Personal Fact Memory reference layer. Serving only the active value of each correctly keyed slot eliminates stale exposure; without update resolution 70.3% of prompts expose a superseded value. Participant-aware BM25 matches the reference ranker within 0.02 once retrievers share the active store. The hard part is assigning revisions to the right slot: four LLM key assigners beat a rule extractor on key recall but give lower clean-retrieval rates, and open-domain merge recall on LongMemEval never exceeds 0.062. Same-name speakers in LoCoMo cannot be told apart; prompt prefill, not retrieval, dominates turn latency.

Key insight: Most personal-memory errors are decided when memory is built, so serving only the active value per slot removes stale exposure while slot assignment remains the hard part.

Relevance Is Not Sufficiency: What Actually Closes the Evidence Gap in Long-Term Memory QA

Yufeng Li; Shuxin Li; Zhenhua Xu et al. arXiv: 2610.09348

Figure from Relevance Is Not Sufficiency: What Actually Closes the Evidence Gap in Long-Term Memory QA
Relevance Is Not Sufficiency: What Actually Closes the Evidence Gap in Long-Term Memory QA

Top-ranked memories can each be relevant yet jointly miss a fact the answer needs, especially on multi-session and temporal questions. Recasts retrieval as building a sufficient memory set: Budgeted Flat Reconstruction first uses Formal Concept Analysis (FCA-MS) to decompose the question into information requirements and pick a covering subset, then keeps acquiring unseen records via deeper text search or entity/session views until budget runs out. Over a flat store it beats same-store adaptations of recent agent-memory systems on LoCoMo and LongMemEval-S; on LongMemEval-S judged accuracy rises 72.4% → 82.2% with Turn Hit 91.4%.

Key insight: Retrieving the most relevant memories is not the same as retrieving enough of them; covering the question's information requirements lifts LongMemEval-S accuracy from 72.4% to 82.2%.

RunningTab: Direct Workspace Interaction with Environment-Side Tabs

Jinheon Baek; Soyeong Jeong; Yumin Choi et al. arXiv: 2610.10444

Figure from RunningTab: Direct Workspace Interaction with Environment-Side Tabs
RunningTab: Direct Workspace Interaction with Environment-Side Tabs

For agents that build deliverables directly from workspace files via a terminal (direct workspace interaction), nothing tracks what the task still requires, what was read, and what was listed but never opened. RunningTab adds an environment-side tab: the agent registers requirements; the environment logs every file read as an excerpt with provenance and every listed-but-unopened file as a candidate; each requirement is shown beside best-matching excerpts and top unopened candidates, and a finish check returns any still-open requirements. Beats plain DWI and in-model record baselines on three benchmarks with three LLMs.

Key insight: An environment-side record of requirements, reads and unopened files keeps workspace agents from finishing with parts of the task missing.

An Empirical Study of Agent Skills' Downstream Utility

Yu Cheng; Dehai Zhao; Zhongxin Liu et al. arXiv: 2610.08875

Figure from An Empirical Study of Agent Skills' Downstream Utility
An Empirical Study of Agent Skills' Downstream Utility

87 SkillsBench tasks, utility = pass-rate difference vs No-Skill under the same model-harness configuration, across nine configurations, with candidates from a 37,596-Skill marketplace corpus. The same Skills help some configurations and hurt others on 36.78% of tasks; recommended procedures can become an execution burden. Relevance rankings miss more useful candidates; reranking by support for required operations raises first-choice pass rates 4.35-5.80 pp. Stage Plan and Dependency DAG organizations beat use order alone (DAG helps most with 5-6 Skills). Derives 17 authoring practices.

Key insight: The same Skill helps under some model and harness setups and hurts under others on over a third of tasks, so Skills must be evaluated per configuration.

The Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance Rules

Litao Hu; Yutong Tang arXiv: 2610.09239

Treats keep-if-better as selection under noise. When Qwen models rewrite their own instructions with 600 held-out items scored, most proposals after the first are harmful. In a pre-registered study, final selection-set scores of greedy loops exceed held-out accuracy by 13-20 points with 16 selection items and 1-5 points with 256; tested acceptance rules did not beat greedy acceptance. The same inflation appears in GEPA and MIPROv2 validation scores. Scoring start and current instruction on 64 never-used items removes average bias, but single estimates stay ~6 points off.

Key insight: Self-improvement loops that keep whatever scores best on a small selection set overstate their gains, by 13 to 20 points with 16 items.

Humanize: Judgement Engineering for Agentic Coding

Sihao Liu; Ligeng Zhu; Zijian Zhang et al. arXiv: 2610.08900

Multi-agent orchestration workflow for agentic coding built on 'judgement engineering': a human approves a plan contract, a builder agent implements in rounds, and a reviewer agent from another vendor decides completion; deterministic hooks, not a model, route work and enforce 72 mechanical gates. 68 versions in 108 days, 1,468 GitHub stars, 118 public postmortems; applications include a 567-file gem5 build migration, top-three placements in all three MLSys 2026 FlashInfer Full-Agent tracks, and 672/672 on PutnamBench. Postmortems show independent review catches unsupported builder claims but stopping is a weakness: two thirds of rounds came after implementation was accepted. Observational, not controlled.

Key insight: Mechanically enforced role boundaries and a reviewer from another vendor make agentic coding reliable, while deciding when to stop remains the weak point.

Loud Failures, Quiet Failures: Fault Detection and Recovery in Tool-Using Language Model Agents

Obada Kraishan arXiv: 2610.10062

Fault-injection layer over a function-calling benchmark: one of four typed faults per trajectory, 1,920 trials, six models over 24 multi-step tasks. Agents flag a problem in 91.3% of trials on explicit errors but only 58.8% on plausible wrong values (vs a 26.8% false-alarm rate). Reasoning models notice less (-9.3 pts) and change plan more (+10.4) than instruct siblings with no recovery gain. Fault-free reruns end in the same state only 63.3% of the time; only a missing tool clearly lowers recovery (39.9%). A prompt line asking the agent to check each result did not move detection.

Key insight: Agents react to explicit error signals but miss plausible wrong tool outputs, and a prompt asking them to check results does not help.

Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System

Panagiotis Kasnesis; Christos Chatzigeorgiou; Lazaros Toumanidis et al. arXiv: 2610.09021

Figure from Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System
Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System

Evaluates 9 models (0.8B to a frontier hosted model) across the five call sites of the open-source Wactorz home-automation framework with production prompts and two real Home Assistant installs (280 cases, 2,520 scored calls). Capability ordering differs by site; a 4B model is worse than its 2B sibling at grounded actuation. The best local model is statistically indistinguishable from both hosted models at four of five sites; only code generation separates them. Gemma4 E2B actuates on 87.2% of requests for devices the site does not own. Routing each site to its best local model reaches 91.8% vs 95.4%; hosting only the two generative sites matches hosting everything (39/43) for 28% of the spend. Benchmark, harness and records released.

Key insight: Small local models match hosted ones at most call sites of a deployed home-automation agent, but can fail unsafely at actuation.

Decoupling Logic from Persona: Structural Immunity of Edge LLM Agents to Context Pollution

Masaaki Nakatsu; Reno Wang arXiv: 2610.09772

Separates logical inference from persona expression on one INT4 base model with hot-swappable LoRA adapters: the logic path sees only the core turn and emits a verifiable structured Micro-State; the persona path renders it with full history. On an Apple M2 laptop (Llama-3.1-8B-Instruct and Gemma-3-4B-it, 4-bit; 480 runs) the logic prompt stays at 180/167 tokens and outputs are byte-identical across pollution levels, while the mixed single pass degrades (logic score 0.669 → 0.150 on Llama). With pollution fed into the logic path, the 8B still holds up but the 4B collapses. Costs one extra decode on a topic's first turn (28.2 s vs 18.2 s); persona swap takes 1.7 ms. Code, adapters and logs released.

Key insight: Keeping the logic call's context minimal and separate from persona rendering makes small edge models immune to history pollution.

AgentTime: Can Agents Estimate and Control Their Own Runtime?

Michael Ofengenden; Maksym Andriushchenko arXiv: 2610.09944

Figure from AgentTime: Can Agents Estimate and Control Their Own Runtime?
AgentTime: Can Agents Estimate and Control Their Own Runtime?

Benchmark (222 tasks from 18 sources: coding, computer use, agentic work, automated research) on whether agents can work for a requested duration, predict runtime and estimate elapsed time. Fable 5.1 in Claude Code deviates from requested runtimes by a typical factor of 2.9x vs 1.2x for GPT-6 Astra in Codex, but 14 of 158 reviewed Astra runs explicitly slept after appearing to finish. Runtime forecasts overestimate; removing temporal information more than doubles retrospective error for Sol and Astra.

Key insight: Agents that can finish tasks still cannot reliably work for a requested duration or estimate their own runtime, and some pad time by sleeping.

Secure-CUA: Controlling Untrusted Influence in Computer-Use Agents

Sarthak Choudhary; Mihai Christodorescu; Ashish Hooda et al. arXiv: 2610.09469

Figure from Secure-CUA: Controlling Untrusted Influence in Computer-Use Agents
Secure-CUA: Controlling Untrusted Influence in Computer-Use Agents

Formalizes security requirements for both a computer-use agent's decisions and their GUI execution. Before seeing untrusted content the agent commits to a per-action program (action transaction) that fixes its queries to untrusted content and the permitted uses of the answers; untrusted regions are masked, an isolated query model answers, and the target is located on the masked interface. Secure by design under its model; on 400 WebArena tasks (three frontier models, 5 seeds, 6,000 traces) task success is 53.55% vs 55.12% for vanilla and 13.17% for CaMeL-CUA.

Key insight: Committing to an action plan before reading untrusted screen content secures computer-use agents at a small utility cost.

CredLeakBench: Evaluating Credential Leakage and Recovery in LLM Agents

Rafid Ahmed; Joseph Fioresi; Mubarak Shah et al. arXiv: 2610.08871

Benchmark for agents automating email, social media, banking and bills when facing phishing and identity verification, covering user-directed login and autonomous inbox monitoring; pairs phishing scenarios with legitimate counterparts and measures leakage by actual submissions in a sandbox. All tested models leak, including during autonomous inbox monitoring with no login request; most mitigations that reduce leakage also hurt genuine tasks.

Key insight: Every tested agent leaks credentials to phishing, even while passively monitoring an inbox, and most defenses also break genuine tasks.

Package Hallucination Attacks on Coding Agents through Prompt Injection in Rule Files

Yupu Wang; Zhengyuan Jiang; Reachal Wang et al. arXiv: 2610.09264

Package hallucination attack: injected prompts in community-shared rule files (such as .cursorrules) make coding agents swap legitimate dependencies for attacker-controlled packages. PackHallu evolves the injected prompts with trajectory-level feedback and LLM-guided mutation, reaching high attack success and strong transfer across models and agent frameworks.

Key insight: Poisoned rule files can make coding agents install attacker-controlled packages.

Bookkeeping, Composition, or Unreachable Gold? Reading MemoryAgentBench's Conflict-Resolution Scores Against a Frozen Last-Write Resolver

Egor Pakhomov; Erik Nijkamp arXiv: 2610.09193

Runs MemoryAgentBench's own Conflict Resolution rule (newest statement wins) as a zero-learning resolver: it answers 80.25% of questions under the official metric (74.5% on three held-out lists). 67 items have gold answers the last-write graph cannot reach (e.g. gold 'New Delhi' after 'The capital of India is Grosseto.'), a third of the multi-hop questions at 262K. Long-context models and a BM25 agent score 84.7%/82.6%/41.6% on rule-solvable items vs 10.4%/11.9%/6.0% on those 67.

Key insight: A frozen newest-statement-wins rule scores 80.25% on MemoryAgentBench conflict resolution, and many remaining items have gold answers that rule cannot reach.

ExperienceIndex: Artifact-Grounded Memory

Peter Baile Chen; Geoffrey X. Yu; Xinming Liu et al. arXiv: 2610.10091

Experience layer that stores what prior reasoning traces learned about specific artifacts in a shared corpus: single-artifact experiences (an artifact's contribution to past tasks) and artifact-pair experiences (structural relations found during reasoning). Integrated as lightweight middleware with experience retrieval, it raises answer quality by up to 11.0 points and cuts online dollar cost by up to 50.5% across corpora and search frameworks; experiences transfer from text-to-SQL to factoid QA on the same corpus, and a stronger model's experiences lift a weaker one.

Key insight: Remembering which artifacts mattered for past tasks lets agents find the right evidence faster, improving quality by up to 11 points while halving cost.

HGP:An on-device personalized agent memory via hybrid graph storage

Ran Zhou; Xueming Han; Jiaheng Liu et al. arXiv: 2610.10071

Hybrid graph memory for on-device personalized agents: a lightweight self-enhancement classifier routes memories (fewer large-model calls), episodic/semantic/procedural memories are stored as graphs, and working memory is extracted as a state trajectory that captures current state and implicit constraints. On PAL-Set solution selection it reaches an S-score of 35.58, nearly 7 points above the strongest baseline. Code and data released.

Key insight: Routing memories by type with a small classifier and storing them as graphs makes personalized agent memory practical on device.

RECAST: Learning to Compute the Right Context through Adaptive Evidence Routing

Yilun Hao; Krishna Sayana; Isabella Ye et al. arXiv: 2610.10507

Treats evidence construction as a sequential decision over retrieval and computation: a small RouterLM picks or specifies operations (filter, aggregate, compute), a frozen CompilerLM turns custom operations into code, and a frozen AnswerLM answers once the router deems evidence sufficient. RouterLM trained with SFT then GRPO. Mean success 75.6% across six benchmark families, +15.9% over the strongest large-model baseline; a trained Qwen3.5-9B router beats a training-free Gemini 3.5 Flash router by 5.0%; +15.0% on three held-out benchmarks.

Key insight: A small trained router that decides which retrieval and computation steps to run can beat large models that simply retrieve.

RippleCP: Measuring Counterfactual Checkpoint Advantage in Coding Agents

Mayur Akewar; Ravi Ranjan arXiv: 2610.09088

Defines counterfactual checkpoint advantage: drive a checkpoint branch and a skip branch to the same failure and recover both under matched conditions. On 12 SWE-bench Verified tasks and 106 real recovery branches, checkpointing saves 49.4 s per task, but split into two regimes: the first checkpoint returns 100.0 s on 156.5 s of protected work (0.64 conversion), a second one step later returns -1.1 s (-0.03). Recovery re-derives rather than replays, so preserved work is a poor guide to saved work.

Key insight: The first checkpoint before expensive work pays off; a second one immediately after returns almost nothing, because recovery re-derives work rather than replaying it.

SkillSandbox: Skill Verification via Dynamic Scenario Synthesis

Serin Kim; Kwangwook Seo; Dokyung Song et al. arXiv: 2610.10088

Verifies whether a distilled skill is reusable beyond its source experience by synthesizing a new skill-relevant task and environment: a Proposer fixes conditions to preserve and details to vary, a Builder constructs an executable scenario, and a Verifier compares runs with and without the skill on executability, utility and efficiency to Keep or Reject it. Strongest downstream performance and better efficiency on ALFWorld and WebShop with three models.

Key insight: Testing a learned skill on a freshly synthesized task, with and without the skill, is a reliable filter for what enters a skill library.

SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles

Yuyao Ge; Yiwei Wang; Yuchen He et al. arXiv: 2610.09832

Agentic RL method where the skill library and model co-evolve through a fitness-driven lifecycle (trial, active, stable, retired). A pre-RL phase uses base-model rollouts to pre-retire low-fitness skills and seed SFT; RL then continues retirement, stabilization and LLM-guided mutation. Highest aggregate success across interactive benchmarks, up to 7.8% relative over the strongest baseline with a compact library. Releases SkillFurnace (5k+ records with retirement events and failure categories).

Key insight: Giving skills an explicit lifecycle, including retirement, keeps libraries compact and improves agent success.

UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving Policy

Yifei Lu; Cheng Liu; Dianzhi Yu et al. arXiv: 2610.10164

A shared policy acts in the environment and proposes skillbank edits (Add, Update, No Edit). Proposals are scored by contrastive action feedback: how swapping in the proposed skill changes the actor's log-likelihood gap between earlier successful and failed trajectories, avoiding extra rollouts; skill-edit support regularization preserves exploration. 98.4% success on ALFWorld and 84.7% on WebShop, still effective with a smaller backbone. Code released.

Key insight: Skill edits can be scored by how they change the actor's preference between past successes and failures, without new rollouts.

FreeEvolve: Learning to Evolve Beyond Fixed Loops

Lecheng Kong; Like Hui; Nikos Kanakaris et al. arXiv: 2610.09197

Automates not just prompts/skills/workflows but the evolution loop itself: given goal, target agent, evaluator, data and resource limits, the evolver decides what to test, how much evidence to gather, which candidates to pursue and when to stop, following an editable evolution skill that is meta-evolved. On tau3-bench, ARC-AGI-2, ARC-AGI-3 and Terminal-Bench 2.1 it improves the held-out metric by 13.6 points on average and matches or beats hand-designed evolvers; meta-evolved skills add 6.9 points over the seed skill on fresh target agents.

Key insight: Letting the evolver decide what to test and when to stop, guided by a meta-evolved skill, matches or beats hand-designed evolution loops.

Agent Plasticity: Measuring Self-Improvement Through Experience

Harman Singh; Anton Bakhtin; Rulin Shao et al. arXiv: 2610.08902

Measures how efficiently agents convert experience into held-out gains when they amortize past experience into reusable artifacts inherited by future instances. Frontier models show sharply different trajectories with comparable learning opportunities; in-regime gains transfer only partly out of distribution; the best final agent is not necessarily the most efficient learner. Low-plasticity agents fail to reuse relevant artifacts; more plastic ones fail despite reuse, pointing at artifact quality or application.

Key insight: How efficiently an agent turns experience into held-out gains varies sharply across frontier models and is not the same as final capability.

Training Advisors for LLM Agents from Task Outcomes

Sergei Polezhaev; Barys Liskavets; Ori Press et al. arXiv: 2610.09858

Trains a natural-language critic with RL from whether the frozen agent ultimately succeeds after receiving its advice, with no step labels or reference critiques. A Qwen3-4B critic trained on multi-hop QA with one base model improves four base models (three unseen); on MuSiQue it lifts Qwen3-4B by more than 25 points, surpassing Kimi K3 without a critic, and transfers to tau3 and DeepDive without extra training. Agents can decide when to ask the critic.

Key insight: A small critic trained only on whether the agent eventually succeeds gives advice that transfers across base models and tasks.

From Uncertainty to Action: Learning to Steer LLM Agents

Hanwen Li; Jinhao Duan; Guanhua Zhu et al. arXiv: 2610.09115

Builds a stepwise outcome table of about 82,000 counterfactual continuations from 1,864 trajectories (three benchmarks, two agents, four steering mechanisms at every step). Uncertainty identifies failing trajectories but no single signal locates the step where steering helps. VoS learns the value of steering per step and uses a harm-budgeted trigger: +7.8 points over unmodified execution on average across all 12 settings, and beats the strongest of five uncertainty-triggered methods in 11 (+2.9 average).

Key insight: Uncertainty flags failing trajectories but not where to intervene; learning the value of steering per step does.

Know the Shape, Find the Fault: Topology-Conditioned Diagnosis of Multi-Agent LLM Failures

Xinwen Liu; Zhuocheng Pan; Isabella Zhu et al. arXiv: 2610.10126

Shows communication topology is strongly associated with failure type (chi^2 = 409.9, p = 1.2e-70). MAScope recovers topology from traces by grounding an interaction graph in message evidence, then a topology-conditioned judge classifies failures. On 851 MAST-clean traces, ground-truth topology raises gpt-mini Macro-F1 0.173 → 0.350; predicted topology gives 0.346, near trace-only gpt-5.4 (0.372), at about 6% of repeated gpt-5.4 diagnosis cost for 1,000 traces.

Key insight: Knowing a multi-agent system's communication topology roughly doubles a small judge's failure-diagnosis accuracy.

A Society of Researchers: Designing Institutions for Populations of Autonomous Research Agents

Ali Asaria; Deep Gandhi; Tony Salomone arXiv: 2610.10468

Position plus system: populations of thousands of persistent research agents sharing compute should get explicit institutions. Principal investigators compete for compute via requests for proposals, independent review and grants; a human 'mayor' allocates resources but assigns no tasks. In a running society of ten thousand researchers asked to improve LM pretraining, one lab reported about 30% less compute for the same quality, a result other labs that tested it do not yet agree on. Lists six open problems.

Key insight: Large populations of persistent research agents will organize themselves anyway, so designers should provide explicit institutions.

How Do Agentic LLMs Decide to Call Tools? A Tool-Call Vector Shaped by Suppression

Xijie Gong; Tingxu Han; Jiahao Zhang et al. arXiv: 2610.09624

Converts agentic prompts into minimal contrastive pairs where one request verb flips the tool-call decision (write vs discuss; 500 pairs across Python, Java, C++). Traces the decision to a vector that is causally necessary and sufficient and generalizes to native multi-turn tau2-Bench trajectories. The scaffold sets a tool-call prior; analysis verbs suppress it via features signaling tool use is unnecessary. Same mechanism across seven Qwen, Mistral and Granite models. Code released.

Key insight: Whether an open model calls a tool is governed by a single internal direction that analysis verbs suppress.

When the Governor Becomes the Disturbance: Control-Generated Disturbance and Cost-Aware Backoff in Governed Tool-Using Agents

Veronique Ziegler arXiv: 2610.09037

In a controlled file-recovery environment where stronger regulation triggers imposed tool failures, a cost-blind supervisory governor can turn failures into persistent blocking. A backoff rule that lowers intervention probability from a moving average of induced events improves completion; in 576 Gemini 2.5 Flash episodes over 6 tasks backoff reduces blocking, with intermediate strength best. Costs are imposed and observable; generality is open.

Key insight: A supervisor that blocks tool use without accounting for its own cost can stall the agent it regulates; backoff restores completion.

Adversarial Images Hijack Web Agents from Visual Grounding to Browser Execution

Wanjing Han; Levi Taiji Li; Mu Zhang et al. arXiv: 2610.09240

Figure from Adversarial Images Hijack Web Agents from Visual Grounding to Browser Execution
Adversarial Images Hijack Web Agents from Visual Grounding to Browser Execution

End-to-end grounding-to-execution attack on vision-grounded web agents: localized visual perturbations make agents select attacker-controlled content and execute the matching browser action across varied renderings, using a role-slot abstraction, webpage recomposition and dataflow analysis aligned with action post-processing. Across four agent configurations, six VLM backbones and 2,250 tasks (13 public sites + sandbox): 91.9% average attack success vs 17.4% for the strongest baseline, and effective against three agent-level defenses.

Key insight: Localized adversarial images can steer web agents all the way to attacker-chosen browser actions, with 91.9% success.

Visual Memory Attacks Can Persist Through The KV Cache

David Dobre; Leo Schwinn; Gauthier Gidel et al. arXiv: 2610.09027

Adversarial images can be optimized so their backdoor persists through the KV cache of later tokens after the image is masked from attention (P-VMI). On Qwen3-VL-8B-Instruct, up to ~90% target success, still effective when the image is shown only on the first turn; a cache-swap ablation localizes the effect to the KV cache, and attacks can survive compaction that keeps a summary's KV cache.

Key insight: Adversarial influence can persist in the KV cache after the malicious image has been removed from context.

Multi-Aspect Runtime Verification for Simulation-Based V&V of LLM-Enabled Autonomous Agents

Nikolaos Kekatos; Dimitrios Nikou; Anastasios Temperekidis et al. arXiv: 2610.08928

Decomposes each natural-language policy clause into a typed spatial/temporal/semantic triple over one event stream, monitors each aspect separately (past-time aspect on unmodified MonPoly) and fuses verdicts with a four-valued provenance algebra. Across two defence mission domains and one civil domain, composition under precautionary blocking drives attack success to zero with no observed false positives at microsecond per-event cost, while any single aspect or pair leaves many attacks succeeding.

Key insight: Checking spatial, temporal and semantic aspects of a policy together stops tool-attack chains that any single monitor misses.

SpecGuard: Proving a Task Is Broken Before the Agent Cheats

Param Biyani; Krishnamurthy Dvijotham arXiv: 2610.09159

Before an agent acts, autoformalizes task intent into a Lean 4 specification, formalizes tests independently, and has the Lean kernel check whether any implementation can satisfy both, issuing a machine-checked certificate when none can. On conflicted SWE-bench tasks it detects up to 72.8% of conflicts and certifies up to 51.1%, with a nearly five-fold lower miss rate than model-based judgment. Code released.

Key insight: Formal checking can prove a coding task's description and tests contradict each other before an agent is tempted to cheat.

Correct Answers, Unsupported Findings: Evidence Binding in Forensic Reconstruction of LLM Agent Logs

Taehyeon Yun; Dongho Kim; Geonwoo Kim et al. arXiv: 2610.09581

64 mechanically checkable cases from saved AgentDojo Banking executions: with original identifiers and no binding table, Sonnet recovered every literal source location but made unsupported citation-source assertions in 26 of 28 cases needing the missing relation, though 22 still matched the reference answer. Explicit identifier-to-record bindings improved grounded reconstruction; renaming identifiers alone did not; a deterministic comparator resolved or abstained correctly throughout.

Key insight: Correct answers from agent logs are not supported findings unless the link between citations and records is preserved.

Finding Blind Spots in AppWorld and WorkArena Task Verifiers

Richard Abrich arXiv: 2610.09142

Source-informed mutation tests on shipped verifiers: in AppWorld, duplicating a non-idempotent write creates an extra record yet passes (6/15 constructed effects); a cardinality patch on checker copies makes all six fail. In WorkArena, 21 of 23 extra-field candidates with confirmed nondefault persisted values still PASS. No population rate is estimated.

Key insight: Shipped AppWorld and WorkArena verifiers accept some runs with wrong side effects, such as duplicated writes or extra fields.

Coding-Agent Benchmarks Should Match Their Users' Task Flows

Igor Slinko; Yaroslav Golubev; Sergey Titov arXiv: 2610.09633

From 4,782 real JetBrains IDE agent sessions, long sessions mix questions, planning, review, refactoring and execution and switch between them, unlike issue-derived benchmarks; public corpora differ from each other too. SWE-TaskFlow transforms issue benchmarks toward a target task flow via prompt splitting and verifiable repository QA. On 700 SWE-Bench Pro tasks, solving in several steps roughly doubles cost without a stable change in resolve rate.

Key insight: Real coding-agent sessions mix task types, and splitting tasks the way users do roughly doubles cost without changing resolve rate.