Backfill for a missed day. Thursday’s cs.AI announcement (Thu 17 Sep 2026) lists 79 new and 127 cross-lists (replacements skipped; listing total 206). Filtering for agent systems, memory and context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent coordination, persistence, and local/open models keeps 37 papers — stale-KV repair after context edits, progress-reporting tool calls for KV serving, self-evolving GUI skills, temporal-logic contracts and trace-level policy checks for agents, enterprise computer-use evaluation, compaction failure regimes in personal agents, SSD-streamed 35B MoE inference, and Git as shared agent memory.
Kratika Bhagtani; Kusha Sridhar; Maziyar Baran Pouyan; Yuying Zhao; … arXiv: 2609.17885
ERPBench evaluates screenshot-only computer-use agents on a live, reproducible ERP system, scoring against ground-truth database values, with a production harness that gates actions behind human approval. Across six closed/open agents, some save the form in up to 85% of runs but write the correct value in as few as 3%.
Key insight: On real enterprise software, a computer-use agent that reaches and saves the right form can still write the wrong value — evaluation has to check the database, not the screen.
Mingyang Mao; Wyatt Mackey; Xiaomin Lin arXiv: 2609.17983
Budgeted in-place repair of stale KV caches after edits to retrieved knowledge, working memory, or user state. A contiguous edit-local window recovers ≥0.94 of the post-edit answer margin, beats attention-, KV-deviation- and structural selectors, and is 13–21× faster than full re-prefill; the edge vanishes when answer-bearing text moves downstream.
Key insight: After an edit to cached context, recomputing a contiguous window right after the edit repairs most of the damage at a fraction of full re-prefill cost — as long as the dependent text sits next to the edit.
Yu Lin; Yiming Wang; Runyuan Cai; Hanze Liu; … arXiv: 2609.18063
Edge0 streams a 35B MoE from SSD using a per-layer prerouter that predicts next-layer routing one token ahead (the prediction is the routing, so nothing is dropped), plus an unmerged recovery LoRA for int4 + routing loss. On a single 24GB machine it serves 35B at 20 tok/s in 3 GiB peak active memory, within a few points of the fp16 teacher; an 8B tier too. Framework, checkpoints and adapters open source.
Key insight: Predicting the next layer's routing one token ahead and using it as the routing lets a 35B MoE stream from SSD at interactive speed in a few GiB of active memory.
Yifeng Xiao; Pierluigi Nuzzo arXiv: 2609.18128
ContrAgent formalizes an agent's tool-call sequence as a trace over checkable predicates and specifies required behavior as assume-guarantee contracts in LTLf; each contract compiles to a DFA that both gates actions online and grades recorded traces offline. The contract library is maintained independently of the model and reusable across agents in a domain.
Key insight: A single compiled temporal-logic contract can both block unsafe tool calls at runtime and grade recorded traces offline, independent of the underlying model.
Ashwini Kurady; Sri Sai Charith Grandhi; Rajesh Gupta; Sumit Kumar arXiv: 2609.18820
Defines Compositional Policy Violations: every step passes its own check while the composed execution violates policy. Taxonomy of Authority Creep, Threshold Laundering, Cumulative Sum Violation, Context Collapse; proposes a provenance-aware runtime that evaluates policies over complete traces, recomputing guarded quantities from raw provenance.
Key insight: Step-scoped guardrails cannot detect violations that only exist over the whole execution; policies have to be evaluated over complete traces from raw provenance.
Bofan Chen; Boxuan Zhang; Fei Tang; Zhengxi Lu; … arXiv: 2609.17653
EvoSkill-GUI: training-free skill evolution where each skill is a multi-file package (retrieval metadata, executable plans, backup localization, failure-recovery rules, accessibility utilities, failure cases). Reflect-revise-reuse loop with an information-isolated critic; the executor edits skill files through a restricted tool interface. Up to +16.2% MobileWorld, +6.0% AndroidWorld, +10.5% OSWorld. Code: github.com/ZJU-REAL/EvoSkill-GUI.
Key insight: GUI-agent skills work better as living, multi-file packages revised from deployment feedback by an isolated critic than as static artifacts written once.
Franziska Roesner; Tadayoshi Kohno arXiv: 2609.17817
Thompson's trusting-trust attack against self-modifying coding agents: poisoned benchmarks in the self-evaluation loop make future versions write vulnerable code on clean tasks. Proof-of-concept on Darwin Gödel Machine (modified), Self-Improving Coding Agent, and Hyperagents; with Sonnet 4.5, Hyperagents self-evolved instructions disabling HTTPS certificate validation. Contamination often persists after re-evolution on clean benchmarks.
Key insight: Self-modifying coding agents inherit Thompson's trusting-trust problem: poisoned self-evaluation data can implant insecure behavior that survives later clean evolution.
Yifan Zhang; Yunheng Zou; Shaokun Zhang; Jian Hu; … arXiv: 2609.18094
Agora stores research-agent contributions (result, insight, hypothesis, verification, report) as an append-only DAG in Git with searchable views and diversity-aware recommendations. A ~12-day run with 13 LM workers and no central planner produced 1,703 contributions and 165 cross-account verifications of 95 targets; best method closed 62% of the gap to a trained GPT-2 124M on a weight-transfer task.
Key insight: Git's append-only history works as shared memory for many independent research agents, including cross-account verification of each other's results.
Yipeng Liu; Yingqiang Zhang; Feifei Li; Huanchen Zhang arXiv: 2609.18849
Argues tool calls should report progress while running so the serving system can decide whether an agent's KV cache stays, leaves, or returns. A census of four public agent corpora finds a readable progress signal in most tool time; plugged into a production engine it cuts p90 TTFT after a tool call by 20.7% (HBM) / 20.8% (HBM+DRAM) vs LRU, near an oracle.
Key insight: Tools already know how far along they are; exposing that progress lets the serving layer make far better keep-or-evict decisions for an agent's KV cache than any up-front duration guess.
Guosen Wu; Huizhen Huang; Guoxiong Long; Tao Huang; … arXiv: 2609.18864
ASLEval measures 'privacy exposure displacement' — the gap between a local proxy (one action, final answer, attacker report) and target-grounded exposure across all visible exits of a multi-step session. An expected-outlet-only view misses 46.9% of exposure; attacker self-reports combine omissions with high false discovery.
Key insight: Privacy evaluations that watch one designated outlet miss almost half the exposure an agent session actually produces across all visible exits.
Yu Liu; Wenxiao Zhang; Cheng Hu; Cong Cao; … arXiv: 2609.19059
MIRAGE holds evidence, questions and scoring fixed while varying only conversation state for multimodal personal agents. Across seven frontier and open-weight backbones, pre-compaction depth and post-compaction continuation form distinct non-monotonic failure regimes; open-weight models lean on context continuity and rarely switch to tool-mediated retrieval when provenance fails.
Key insight: Conversation compaction creates distinct failure regimes for personal agents, and open-weight models in particular fail to fall back to tool-based retrieval when context continuity breaks.
Nilesh Verma; Nick Lim; Albert Bifet; Bernhard Pfahringer arXiv: 2609.17984
TuiML is an ML library built for agents: every component self-describes through machine-readable metadata and parameter schemas so agents can search, inspect, compose validated workflows and register new components; every call is validated, seeded and traced, sessions export as notebooks. One spec layer drives MCP, framework adapters, Python API, CLI and local serving; data stays on-machine.
Key insight: Libraries built for agents should self-describe through machine-readable schemas, validate and trace every call, and drive MCP, CLI and API from one specification.
Qingnuan Han; Boli Fang; Mingzhi Hou; Claire Liu arXiv: 2609.17985
RideWay: efficiency-centred ridehailing benchmark in a stateful tool-calling environment with a success-gated Efficiency Utility. Across 58 tasks and 24 models, the human-fitted penalty for excess user-facing turns is ~2× that for excess tool calls; held-out preference accuracy 78.7% (90.6% when turns differ, chance when only tool calls differ).
Key insight: Users penalize an agent's extra questions about twice as much as its extra tool calls, so efficiency metrics need to weigh those axes differently.
Junnan Dong; Linhao Luo; Senlei Zhang; Gong Chen; … arXiv: 2609.18182
WFM (Wiki Foundation Model) targets 'LLM Wiki' — markdown documents with multi-layer topological links — as an agent-native knowledge representation, formalizing a Wiki Graph schema and a query-conditioned aggregation model for scalable representation and retrieval.
Key insight: Markdown 'LLM Wiki' knowledge bases are becoming an agent-native memory format, and they need retrieval models built for dense text plus explicit links.
Guojun Zhu; Xunheng Huang; Peng Yin; Jiahui Xie; … arXiv: 2609.18366
CHASE treats harness evolution as constraint generation over counterfactual benchmarks: after each Proposer edit (prompts, memory, retrieval, tools, control code), a Challenger searches for a protocol transformation that destroys the gain; a validity firewall preserves task semantics. Evaluated on Syn-Ledger and OfficeQA, it retains released-benchmark gains while reducing shortcut dependence.
Key insight: Automatic harness optimization can overfit to a benchmark's protocol rather than its tasks; counterfactual protocol transformations expose those shortcuts.
João Meneses dos Santos; Arlindo L. Oliveira arXiv: 2609.19128
Extends SwiftSage with an Adaptive Memory Module (salience-gated episodic storage, trigger-driven retrieval) and a Self-Reflection Module (bounded execution-time validation). On ScienceWorld the full system is best (64.62 score, 43.17% success, 19.33 steps); SRM is the strongest standalone contributor.
Key insight: In interactive environments, execution-time self-reflection is the dominant fix; episodic memory pays off once the runtime loop is stable.
Rushabh Vipulkumar Patel; Dipo Dunsin; Mohammed Almaiah; Mohamed Chahine Ghanem arXiv: 2609.18120
PentestChain exposes a ten-phase pentest pipeline as an MCP server with eleven tools, running a local Ollama qwen2.5-7b first, then free-tier OpenRouter/Cerebras, then a rule-based fallback, with a deterministic exploit map on the critical path. Includes a four-position MCP threat model grounded in CVE-2025-6514 (mcp-remote RCE), the postmark-mcp supply-chain backdoor, and tool-poisoning/rug-pull/line-jumping, with four mitigations.
Key insight: A deterministic backbone can keep a small local model off the critical path, and an MCP-exposed toolchain needs its own explicit threat model.
Jin Gao arXiv: 2609.19125
Affora is a design system that makes interface actions and task state legible to computer-use agents while preserving visual freedom for people; three controlled studies plus evaluation on independently authored interfaces show gains where existing deficits exist.
Key insight: Interfaces can be made legible to computer-use agents without giving up visual freedom, as long as the interaction meaning is preserved in the representation agents read.
Omer Tafveez arXiv: 2609.17865
SAFE benchmark: do models choose to acquire safety evidence before acting? Opus 4.8 inspects nearly by default, o3 is most skip-heavy, GPT-5.5 and Sonnet 4.6 in between; raising stated problem probability from 10% to 70% moves inspection by at most 21 pp; avoidance is driven mainly by retrieval friction.
Key insight: Whether frontier models look for safety evidence before acting depends far more on how costly the lookup is than on how likely the problem is.
Caiqi Zhang; Xiaochen Zhu; Chengzu Li; Yulong Chen; … arXiv: 2609.17708
XConf estimates confidence from a record of the model's own graded past episodes (task, reflection, stated confidence, outcome, lesson): Recall retrieves similar past episodes and their success rate; Reflect has the model name its recurring failure mode. Beats or matches ten-sample self-consistency AUROC on 23 of 24 comparisons at one generation's cost.
Key insight: A model's own graded track record on similar tasks is a cheap, model-agnostic basis for calibrated confidence.
Mika Okamoto; Ansel Kaplan Erol arXiv: 2609.18605
PACT (Pressure-Applied Compliance Testing): multi-turn benchmark across twelve regulated enterprise domains and 48 scenarios pairing standing rules against convenient shortcuts under persistent-user and hurried-manager pressure; aggregated into a reliability-weighted PACTScore.
Key insight: Rule compliance of enterprise assistants has to be tested under realistic multi-turn pressure, not single-turn prompts.
Huixin Zhang; Shao-Jun Xia; Di Wang; Liangxi Liu; … arXiv: 2609.17921
Framework for multi-agent VLM systems where agents inspect different regions/frames: memory hierarchy, cross-agent sharing and consistency mechanisms, with shared visual memory preserving dependencies among observations, interpretations and downstream reasoning.
Key insight: Multi-agent vision systems need shared memory that tracks dependencies between observations, interpretations and later reasoning, not just stored summaries.
Jiahong Liu; Wenhao Yu; Zexuan Qiu; Menglin Yang; … arXiv: 2609.17969
Position paper: long-horizon personalization should model memory as a user-specific dynamical state space with locally heterogeneous geometry (stable vs volatile regions, variable-rate drift, uncertainty about current user state); access becomes trajectory-conditioned reconstruction, not nearest-neighbour lookup.
Key insight: Personal memory may be better modeled as a user-specific dynamical state space with uneven drift than as a static set of searchable records.
Carolina Fortuna; Vid Hanžel; Tim Strnad; Blaž Bertalanič arXiv: 2609.18283
agentic-eCAL extends an energy-cost metric to directed multi-agent workflows (prefill compute-bound, decode memory-bound, plus OSI transport), grounded in A100/H100 benchmarks over 16 open-weight models and 8 orchestration topologies, to decide where agent teams should run.
Key insight: Energy cost of multi-agent workflows depends on topology and placement across edge and cloud, not just on single-model inference.
Cai Ke; Xinghao Chen; Xiaoyu Shen; Keyu Chen; … arXiv: 2609.18461
LGM maps historical interactions into latent memory nodes via a sparse autoencoder, disentangles traces into sparse concept activations, and builds query-aware relational edges with conditioned message passing, instead of persisting a fixed graph.
Key insight: Long-term personal memory can be disentangled into sparse latent concepts with query-dependent relations instead of a fixed memory graph.
Dingli Liang; Yiqiao Xie; Yukai Huang; Zhaokai Wang; … arXiv: 2609.17688
CapMem: 75 egocentric videos (33.7 h), 1,000 MCQs; on >20-min videos, caption-window QA beats direct VideoQA for 10/12 (30 s) and 8/12 (60 s) models; a caption-guided retrieve-and-verify harness adds up to 5.3 points.
Key insight: Text captions over time windows are a strong, reusable episodic memory for long egocentric video under realistic compute limits.
Sikun Wang; Yixi Zhou; Lei Fan; Fan Zhang arXiv: 2609.17695
GraphEcho tests whether graph agents mistake repeated paths to the same evidence for corroboration. Redundant paths increase repeated walks across all frozen agents; provenance-aware post-training reduces revisits but covers fewer distinct sources and loses accuracy on scientific claims.
Key insight: Agents tend to treat repeated paths to the same evidence as extra corroboration; provenance, not hit count, should drive confidence.
Mojtaba Abdolmaleki; Stefanus Jasin; Boyu Wang arXiv: 2609.18126
Formalizes running several agentic workflows and selecting among outputs as a portfolio problem with compute cost; odds-lift index for selector quality, LP relaxations with certificates, evaluated on ABCD, Schema-Guided Dialogue and HotpotQA.
Key insight: Running a portfolio of different agent workflows and selecting among outputs can beat the single best workflow, with quantifiable compute trade-offs.
Li Chen arXiv: 2609.18123
AutoTuneBench: protocol for agents tuning kernels/serving engines with anti-cheat outside the agent's modification surface and pre-registered readouts. Honest baselines rewrite headlines: 10.6× vs naive becomes 2.03×; 1.174× on one machine is 1.0049× on another; KernelBench L1 median speedup 1.0001× over eager.
Key insight: Agent-driven performance tuning needs a measurement harness outside the agent's reach; honest baselines shrink headline speedups dramatically.
Mohamed Chahine Ghanem arXiv: 2609.18272
Grades auditor independence along principal, substrate (shared model family/toolchain/guardrails) and evidence axes, using a beta-factor common-cause model; in a Monte Carlo study a conventional internal audit of an agent surfaces 5.9% of in-principle visible faults.
Key insight: Auditor independence for agentic systems should be graded along principal, substrate and evidence axes rather than treated as yes/no.
Xiangfan Wu; Zonghao Ying; Huiyu Wu; Xing Zheng; … arXiv: 2609.18460
Epidemic account of collective loss of control (mutation, contagion, recovery). RogueHandoff-20 (20 executable scenarios): executed harm 0–5% on normal tasks but 40–95% after injected unsafe trajectories, 5–45 pp above direct malicious requests.
Key insight: Multi-agent systems can turn rare individual deviations into collective failure through contagion along communication paths.
Jiaxuan Jiang; Liyuan He; Zhixuan Fang arXiv: 2609.18779
CERA-MoA co-evolves a router and continually fine-tuned agents via RL, scoring agent competence from mid-layer hidden states and activating a minimal agent subset by cumulative threshold.
Key insight: Routing and agent specialization in mixture-of-agents systems work better when co-evolved rather than tuned separately.
Jinli Hu; Ross M. Clarke; Yichuan Zhang; José Miguel Hernández-Lobato arXiv: 2609.18842
Infinite-Parameter LLM: a compact hypernetwork turns online interaction data into low-rank modulations of a shared base, with a Bayesian belief over generator state updated during the session — live data in weights instead of the prompt.
Key insight: Live interaction data can be represented in generated weights instead of the prompt, keeping memory footprint constant.
Haonan Huang; Joey Xiao arXiv: 2609.18996
Gauntlet develop-freeze-evaluate protocol: off-the-shelf coding agents build game controllers from bare interfaces; environment access adds 10–78 pp held-out success; StarCraft II controllers beat every fair built-in AI. Replaying 5,000+ frozen versions shows early plateaus and sparse validation of what ships.
Key insight: Coding agents can conduct sustained autonomous research, but replays show early plateaus and sparse validation of what they ship.
Juan P. Madrigal-Cianci; Eshan Chordia arXiv: 2609.17590
Evolutionary Ensemble Search: role-specialized council, orchestrator, execution specialists and an evolutionary engine with session memory and problem-indexed lessons across runs; development ledger reports medal-threshold artifacts on 19/22 MLE-bench Lite tasks.
Key insight: Cumulative program search benefits from a council-plus-orchestrator design with memory and lessons carried across runs.
Jeonghye Kim; Minseon Kim; Young Jin Kim; Matheus Pereira; … arXiv: 2609.18805
ProgramDistill: 4,063 tasks where coding agents infer features from working reference web apps; GPT-6 Astra 49.2% and Claude Opus 5 28.8% on cumulative full-app reconstruction; partial reconstruction drops sharply with restoration depth.
Key insight: Inferring behavior from a working reference app is a hard, scalable test for coding agents that issue-based benchmarks miss.
Xinshuai Guo; Junjie Wu; Dolly Deng; Yinghui Li; … arXiv: 2609.18909
DualViewEval compresses agent benchmarks using six trajectory process signals plus outcomes; with 20 tasks achieves 24–40× compression on APEX-Agents and BFCL with 14.5–28.2% lower MAE than the strongest competitors.
Key insight: Trajectory-level process signals let a 20-task miniset predict full agent-benchmark scores far more accurately than outcome-only compression.