Backfill for a missed day. Thursday’s cs.AI announcement (Thu 17 Sep 2026) lists 79 new and 127 cross-lists (replacements skipped; listing total 206). Filtering for agent systems, memory and context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent coordination, persistence, and local/open models keeps 37 papers — stale-KV repair after context edits, progress-reporting tool calls for KV serving, self-evolving GUI skills, temporal-logic contracts and trace-level policy checks for agents, enterprise computer-use evaluation, compaction failure regimes in personal agents, SSD-streamed 35B MoE inference, and Git as shared agent memory.


Research Papers

ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software

Kratika Bhagtani; Kusha Sridhar; Maziyar Baran Pouyan; Yuying Zhao; … arXiv: 2609.17885

Figure from ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software
ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software

ERPBench evaluates screenshot-only computer-use agents on a live, reproducible ERP system, scoring against ground-truth database values, with a production harness that gates actions behind human approval. Across six closed/open agents, some save the form in up to 85% of runs but write the correct value in as few as 3%.

Key insight: On real enterprise software, a computer-use agent that reaches and saves the right form can still write the wrong value — evaluation has to check the database, not the screen.

Contiguity, Not Importance: Budgeted Repair of Stale KV Caches After Document Edits

Mingyang Mao; Wyatt Mackey; Xiaomin Lin arXiv: 2609.17983

Budgeted in-place repair of stale KV caches after edits to retrieved knowledge, working memory, or user state. A contiguous edit-local window recovers ≥0.94 of the post-edit answer margin, beats attention-, KV-deviation- and structural selectors, and is 13–21× faster than full re-prefill; the edge vanishes when answer-bearing text moves downstream.

Key insight: After an edit to cached context, recomputing a contiguous window right after the edit repairs most of the damage at a fraction of full re-prefill cost — as long as the dependent text sits next to the edit.

The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction

Yu Lin; Yiming Wang; Runyuan Cai; Hanze Liu; … arXiv: 2609.18063

Edge0 streams a 35B MoE from SSD using a per-layer prerouter that predicts next-layer routing one token ahead (the prediction is the routing, so nothing is dropped), plus an unmerged recovery LoRA for int4 + routing loss. On a single 24GB machine it serves 35B at 20 tok/s in 3 GiB peak active memory, within a few points of the fp16 teacher; an 8B tier too. Framework, checkpoints and adapters open source.

Key insight: Predicting the next layer's routing one token ahead and using it as the routing lets a 35B MoE stream from SSD at interactive speed in a few GiB of active memory.

Symbolic Temporal Supervision of LLM Agents Using Contracts

Yifeng Xiao; Pierluigi Nuzzo arXiv: 2609.18128

Figure from Symbolic Temporal Supervision of LLM Agents Using Contracts
Symbolic Temporal Supervision of LLM Agents Using Contracts

ContrAgent formalizes an agent's tool-call sequence as a trace over checkable predicates and specifies required behavior as assume-guarantee contracts in LTLf; each contract compiles to a DFA that both gates actions online and grades recorded traces offline. The contract library is maintained independently of the model and reusable across agents in a domain.

Key insight: A single compiled temporal-logic contract can both block unsafe tool calls at runtime and grade recorded traces offline, independent of the underlying model.

Compositional Policy Violations: When Step-Level Compliance Fails In Agentic AI Workflows

Ashwini Kurady; Sri Sai Charith Grandhi; Rajesh Gupta; Sumit Kumar arXiv: 2609.18820

Figure from Compositional Policy Violations: When Step-Level Compliance Fails In Agentic AI Workflows
Compositional Policy Violations: When Step-Level Compliance Fails In Agentic AI Workflows

Defines Compositional Policy Violations: every step passes its own check while the composed execution violates policy. Taxonomy of Authority Creep, Threshold Laundering, Cumulative Sum Violation, Context Collapse; proposes a provenance-aware runtime that evaluates policies over complete traces, recomputing guarded quantities from raw provenance.

Key insight: Step-scoped guardrails cannot detect violations that only exist over the whole execution; policies have to be evaluated over complete traces from raw provenance.

Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents

Bofan Chen; Boxuan Zhang; Fei Tang; Zhengxi Lu; … arXiv: 2609.17653

Figure from Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents
Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents

EvoSkill-GUI: training-free skill evolution where each skill is a multi-file package (retrieval metadata, executable plans, backup localization, failure-recovery rules, accessibility utilities, failure cases). Reflect-revise-reuse loop with an information-isolated critic; the executor edits skill files through a restricted tool interface. Up to +16.2% MobileWorld, +6.0% AndroidWorld, +10.5% OSWorld. Code: github.com/ZJU-REAL/EvoSkill-GUI.

Key insight: GUI-agent skills work better as living, multi-file packages revised from deployment feedback by an isolated critic than as static artifacts written once.

Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks

Franziska Roesner; Tadayoshi Kohno arXiv: 2609.17817

Thompson's trusting-trust attack against self-modifying coding agents: poisoned benchmarks in the self-evaluation loop make future versions write vulnerable code on clean tasks. Proof-of-concept on Darwin Gödel Machine (modified), Self-Improving Coding Agent, and Hyperagents; with Sonnet 4.5, Hyperagents self-evolved instructions disabling HTTPS certificate validation. Contamination often persists after re-evolution on clean benchmarks.

Key insight: Self-modifying coding agents inherit Thompson's trusting-trust problem: poisoned self-evaluation data can implant insecure behavior that survives later clean evolution.

Agora: Git as Shared Memory for Collective AutoResearch

Yifan Zhang; Yunheng Zou; Shaokun Zhang; Jian Hu; … arXiv: 2609.18094

Figure from Agora: Git as Shared Memory for Collective AutoResearch
Agora: Git as Shared Memory for Collective AutoResearch

Agora stores research-agent contributions (result, insight, hypothesis, verification, report) as an append-only DAG in Git with searchable views and diversity-aware recommendations. A ~12-day run with 13 LM workers and no central planner produced 1,703 contributions and 165 cross-account verifications of 95 targets; best method closed 62% of the gap to a trained GPT-2 124M on a weight-transfer task.

Key insight: Git's append-only history works as shared memory for many independent research agents, including cross-account verification of each other's results.

Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It

Yipeng Liu; Yingqiang Zhang; Feifei Li; Huanchen Zhang arXiv: 2609.18849

Figure from Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It
Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It

Argues tool calls should report progress while running so the serving system can decide whether an agent's KV cache stays, leaves, or returns. A census of four public agent corpora finds a readable progress signal in most tool time; plugged into a production engine it cuts p90 TTFT after a tool call by 20.7% (HBM) / 20.8% (HBM+DRAM) vs LRU, near an oracle.

Key insight: Tools already know how far along they are; exposing that progress lets the serving layer make far better keep-or-evict decisions for an agent's KV cache than any up-front duration guess.

ASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions

Guosen Wu; Huizhen Huang; Guoxiong Long; Tao Huang; … arXiv: 2609.18864

ASLEval measures 'privacy exposure displacement' — the gap between a local proxy (one action, final answer, attacker report) and target-grounded exposure across all visible exits of a multi-step session. An expected-outlet-only view misses 46.9% of exposure; attacker self-reports combine omissions with high false discovery.

Key insight: Privacy evaluations that watch one designated outlet miss almost half the exposure an agent session actually produces across all visible exits.

MIRAGE: How Conversation State Shapes Historical Evidence Use in Multimodal Personal Agents

Yu Liu; Wenxiao Zhang; Cheng Hu; Cong Cao; … arXiv: 2609.19059

Figure from MIRAGE: How Conversation State Shapes Historical Evidence Use in Multimodal Personal Agents
MIRAGE: How Conversation State Shapes Historical Evidence Use in Multimodal Personal Agents

MIRAGE holds evidence, questions and scoring fixed while varying only conversation state for multimodal personal agents. Across seven frontier and open-weight backbones, pre-compaction depth and post-compaction continuation form distinct non-monotonic failure regimes; open-weight models lean on context continuity and rarely switch to tool-mediated retrieval when provenance fails.

Key insight: Conversation compaction creates distinct failure regimes for personal agents, and open-weight models in particular fail to fall back to tool-based retrieval when context continuity breaks.

TuiML: Machine Learning for AI Agents

Nilesh Verma; Nick Lim; Albert Bifet; Bernhard Pfahringer arXiv: 2609.17984

TuiML is an ML library built for agents: every component self-describes through machine-readable metadata and parameter schemas so agents can search, inspect, compose validated workflows and register new components; every call is validated, seeded and traced, sessions export as notebooks. One spec layer drives MCP, framework adapters, Python API, CLI and local serving; data stays on-machine.

Key insight: Libraries built for agents should self-describe through machine-readable schemas, validate and trace every call, and drive MCP, CLI and API from one specification.

RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents

Qingnuan Han; Boli Fang; Mingzhi Hou; Claire Liu arXiv: 2609.17985

Figure from RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents
RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents

RideWay: efficiency-centred ridehailing benchmark in a stateful tool-calling environment with a success-gated Efficiency Utility. Across 58 tasks and 24 models, the human-fitted penalty for excess user-facing turns is ~2× that for excess tool calls; held-out preference accuracy 78.7% (90.6% when turns differ, chance when only tool calls differ).

Key insight: Users penalize an agent's extra questions about twice as much as its extra tool calls, so efficiency metrics need to weigh those axes differently.

WFM: Wiki Foundation Model for Complex Agentic Reasoning

Junnan Dong; Linhao Luo; Senlei Zhang; Gong Chen; … arXiv: 2609.18182

WFM (Wiki Foundation Model) targets 'LLM Wiki' — markdown documents with multi-layer topological links — as an agent-native knowledge representation, formalizing a Wiki Graph schema and a query-conditioned aggregation model for scalable representation and retrieval.

Key insight: Markdown 'LLM Wiki' knowledge bases are becoming an agent-native memory format, and they need retrieval models built for dense text plus explicit links.

Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts

Guojun Zhu; Xunheng Huang; Peng Yin; Jiahui Xie; … arXiv: 2609.18366

CHASE treats harness evolution as constraint generation over counterfactual benchmarks: after each Proposer edit (prompts, memory, retrieval, tools, control code), a Challenger searches for a protocol transformation that destroys the gain; a validity firewall preserves task semantics. Evaluated on Syn-Ledger and OfficeQA, it retains released-benchmark gains while reducing shortcut dependence.

Key insight: Automatic harness optimization can overfit to a benchmark's protocol rather than its tasks; counterfactual protocol transformations expose those shortcuts.

Cognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments

João Meneses dos Santos; Arlindo L. Oliveira arXiv: 2609.19128

Extends SwiftSage with an Adaptive Memory Module (salience-gated episodic storage, trigger-driven retrieval) and a Self-Reflection Module (bounded execution-time validation). On ScienceWorld the full system is best (64.62 score, 43.17% success, 19.33 steps); SRM is the strongest standalone contributor.

Key insight: In interactive environments, execution-time self-reflection is the dominant fix; episodic memory pays off once the runtime loop is stable.

PentestChain: A Cost-Aware, MCP-Orchestrated Framework for Automated Penetration Testing with Free-Tier LLMs

Rushabh Vipulkumar Patel; Dipo Dunsin; Mohammed Almaiah; Mohamed Chahine Ghanem arXiv: 2609.18120

PentestChain exposes a ten-phase pentest pipeline as an MCP server with eleven tools, running a local Ollama qwen2.5-7b first, then free-tier OpenRouter/Cerebras, then a rule-based fallback, with a deterministic exploit map on the critical path. Includes a four-position MCP threat model grounded in CVE-2025-6514 (mcp-remote RCE), the postmark-mcp supply-chain backdoor, and tool-poisoning/rug-pull/line-jumping, with four mitigations.

Key insight: A deterministic backbone can keep a small local model off the critical path, and an MCP-exposed toolchain needs its own explicit threat model.

Affora: A Design System for Agent-Friendly Interfaces

Jin Gao arXiv: 2609.19125

Affora is a design system that makes interface actions and task state legible to computer-use agents while preserving visual freedom for people; three controlled studies plus evaluation on independently authored interfaces show gains where existing deficits exist.

Key insight: Interfaces can be made legible to computer-use agents without giving up visual freedom, as long as the interaction meaning is preserved in the representation agents read.

Do Frontier Models Seek Safety Evidence Before Acting?

Omer Tafveez arXiv: 2609.17865

SAFE benchmark: do models choose to acquire safety evidence before acting? Opus 4.8 inspects nearly by default, o3 is most skip-heavy, GPT-5.5 and Sonnet 4.6 in between; raising stated problem probability from 10% to 70% moves inspection by at most 21 pp; avoidance is driven mainly by retrieval friction.

Key insight: Whether frontier models look for safety evidence before acting depends far more on how costly the lookup is than on how likely the problem is.

Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents

Caiqi Zhang; Xiaochen Zhu; Chengzu Li; Yulong Chen; … arXiv: 2609.17708

XConf estimates confidence from a record of the model's own graded past episodes (task, reflection, stated confidence, outcome, lesson): Recall retrieves similar past episodes and their success rate; Reflect has the model name its recurring failure mode. Beats or matches ten-sample self-consistency AUROC on 23 of 24 comparisons at one generation's cost.

Key insight: A model's own graded track record on similar tasks is a cheap, model-agnostic basis for calibrated confidence.

PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?

Mika Okamoto; Ansel Kaplan Erol arXiv: 2609.18605

PACT (Pressure-Applied Compliance Testing): multi-turn benchmark across twelve regulated enterprise domains and 48 scenarios pairing standing rules against convenient shortcuts under persistent-user and hurried-manager pressure; aggregated into a reliability-weighted PACTScore.

Key insight: Rule compliance of enterprise assistants has to be tested under realistic multi-turn pressure, not single-turn prompts.

Collaborative Memory for Multi-Agent VLM Systems

Huixin Zhang; Shao-Jun Xia; Di Wang; Liangxi Liu; … arXiv: 2609.17921

Framework for multi-agent VLM systems where agents inspect different regions/frames: memory hierarchy, cross-agent sharing and consistency mechanisms, with shared visual memory preserving dependencies among observations, interpretations and downstream reasoning.

Key insight: Multi-agent vision systems need shared memory that tracks dependencies between observations, interpretations and later reasoning, not just stored summaries.

Memory Has Geometry: Non-Uniform Geometric Memory for Long-Horizon Personalized AI

Jiahong Liu; Wenhao Yu; Zexuan Qiu; Menglin Yang; … arXiv: 2609.17969

Position paper: long-horizon personalization should model memory as a user-specific dynamical state space with locally heterogeneous geometry (stable vs volatile regions, variable-rate drift, uncertainty about current user state); access becomes trajectory-conditioned reconstruction, not nearest-neighbour lookup.

Key insight: Personal memory may be better modeled as a user-specific dynamical state space with uneven drift than as a static set of searchable records.

Where Should Agents Live? Energy-Memory Characterization of Agentic AI for the Edge-Cloud Continuum

Carolina Fortuna; Vid Hanžel; Tim Strnad; Blaž Bertalanič arXiv: 2609.18283

agentic-eCAL extends an energy-cost metric to directed multi-agent workflows (prefill compute-bound, decode memory-bound, plus OSI transport), grounded in A100/H100 benchmarks over 16 open-weight models and 8 orchestration topologies, to decide where agent teams should run.

Key insight: Energy cost of multi-agent workflows depends on topology and placement across edge and cloud, not just on single-model inference.

Disentangling Long-Term Memory via Latent Neuro-Symbolic Reasoning

Cai Ke; Xinghao Chen; Xiaoyu Shen; Keyu Chen; … arXiv: 2609.18461

LGM maps historical interactions into latent memory nodes via a sparse autoencoder, disentangles traces into sparse concept activations, and builds query-aware relational edges with conditioned message passing, instead of persisting a fixed graph.

Key insight: Long-term personal memory can be disentangled into sparse latent concepts with query-dependent relations instead of a fixed memory graph.

CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video

Dingli Liang; Yiqiao Xie; Yukai Huang; Zhaokai Wang; … arXiv: 2609.17688

CapMem: 75 egocentric videos (33.7 h), 1,000 MCQs; on >20-min videos, caption-window QA beats direct VideoQA for 10/12 (30 s) and 8/12 (60 s) models; a caption-guided retrieve-and-verify harness adds up to 5.3 points.

Key insight: Text captions over time windows are a strong, reusable episodic memory for long egocentric video under realistic compute limits.

GraphEcho: Structural Redundancy and Evidence Provenance in LLM Graph Agents

Sikun Wang; Yixi Zhou; Lei Fan; Fan Zhang arXiv: 2609.17695

GraphEcho tests whether graph agents mistake repeated paths to the same evidence for corroboration. Redundant paths increase repeated walks across all frozen agents; provenance-aware post-training reduces revisits but covers fewer distinct sources and loses accuracy on scientific claims.

Key insight: Agents tend to treat repeated paths to the same evidence as extra corroboration; provenance, not hit count, should drive confidence.

Designing Agentic AI Workflow Portfolios under Imperfect Selection and Compute Cost

Mojtaba Abdolmaleki; Stefanus Jasin; Boyu Wang arXiv: 2609.18126

Formalizes running several agentic workflows and selecting among outputs as a portfolio problem with compute cost; odds-lift index for selector quality, LP relaxations with certificates, evaluated on ABCD, Schema-Guided Dialogue and HotpotQA.

Key insight: Running a portfolio of different agent workflows and selecting among outputs can beat the single best workflow, with quantifiable compute trade-offs.

AutoTuneBench: Trustworthy Measurement for Agent Auto-Tuning of LLM Serving Engines

Li Chen arXiv: 2609.18123

AutoTuneBench: protocol for agents tuning kernels/serving engines with anti-cheat outside the agent's modification surface and pre-registered readouts. Honest baselines rewrite headlines: 10.6× vs naive becomes 2.03×; 1.174× on one machine is 1.0049× on another; KernelBench L1 median speedup 1.0001× over eager.

Key insight: Agent-driven performance tuning needs a measurement harness outside the agent's reach; honest baselines shrink headline speedups dramatically.

Who Audits Whom, on What Substrate, with What Evidence? An Independence-Graded Audit Protocol for Agentic AI

Mohamed Chahine Ghanem arXiv: 2609.18272

Grades auditor independence along principal, substrate (shared model family/toolchain/guardrails) and evidence axes, using a beta-factor common-cause model; in a Monte Carlo study a conventional internal audit of an agent surfaces 5.9% of in-principle visible faults.

Key insight: Auditor independence for agentic systems should be graded along principal, substrate and evidence axes rather than treated as yes/no.

Collective Loss of Control in LLM Agent Systems: An Epidemic Account of Mutation, Contagion, and Recovery

Xiangfan Wu; Zonghao Ying; Huiyu Wu; Xing Zheng; … arXiv: 2609.18460

Epidemic account of collective loss of control (mutation, contagion, recovery). RogueHandoff-20 (20 executable scenarios): executed harm 0–5% on normal tasks but 40–95% after injected unsafe trajectories, 5–45 pp above direct malicious requests.

Key insight: Multi-agent systems can turn rare individual deviations into collective failure through contagion along communication paths.

CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents

Jiaxuan Jiang; Liyuan He; Zhixuan Fang arXiv: 2609.18779

CERA-MoA co-evolves a router and continually fine-tuned agents via RL, scoring agent competence from mid-layer hidden states and activating a minimal agent subset by cumulative threshold.

Key insight: Routing and agent specialization in mixture-of-agents systems work better when co-evolved rather than tuned separately.

Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data

Jinli Hu; Ross M. Clarke; Yichuan Zhang; José Miguel Hernández-Lobato arXiv: 2609.18842

Infinite-Parameter LLM: a compact hypernetwork turns online interaction data into low-rank modulations of a shared base, with a Bayesian belief over generator state updated during the session — live data in weights instead of the prompt.

Key insight: Live interaction data can be represented in generated weights instead of the prompt, keeping memory footprint constant.

Compiled Agency: Coding Agents as Game AI Researchers -- from a Roguelike to StarCraft II and Civilization

Haonan Huang; Joey Xiao arXiv: 2609.18996

Gauntlet develop-freeze-evaluate protocol: off-the-shelf coding agents build game controllers from bare interfaces; environment access adds 10–78 pp held-out success; StarCraft II controllers beat every fair built-in AI. Replaying 5,000+ frozen versions shows early plateaus and sparse validation of what ships.

Key insight: Coding agents can conduct sustained autonomous research, but replays show early plateaus and sparse validation of what they ship.

Evolutionary Ensemble Search: Council-Guided Program Evolution with Persistent Memory

Juan P. Madrigal-Cianci; Eshan Chordia arXiv: 2609.17590

Evolutionary Ensemble Search: role-specialized council, orchestrator, execution specialists and an evolutionary engine with session memory and problem-indexed lessons across runs; development ledger reports medal-threshold artifacts on 19/22 MLE-bench Lite tasks.

Key insight: Cumulative program search benefits from a council-plus-orchestrator design with memory and lessons carried across runs.

ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks

Jeonghye Kim; Minseon Kim; Young Jin Kim; Matheus Pereira; … arXiv: 2609.18805

ProgramDistill: 4,063 tasks where coding agents infer features from working reference web apps; GPT-6 Astra 49.2% and Claude Opus 5 28.8% on cumulative full-app reconstruction; partial reconstruction drops sharply with restoration depth.

Key insight: Inferring behavior from a working reference app is a hard, scalable test for coding agents that issue-based benchmarks miss.

Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking

Xinshuai Guo; Junjie Wu; Dolly Deng; Yinghui Li; … arXiv: 2609.18909

DualViewEval compresses agent benchmarks using six trajectory process signals plus outcomes; with 20 tasks achieves 24–40× compression on APEX-Agents and BFCL with 14.5–28.2% lower MAE than the strongest competitors.

Key insight: Trajectory-level process signals let a 20-task miniset predict full agent-benchmark scores far more accurately than outcome-only compression.