Friday’s cs.AI listing had 165 new submissions and cross-lists. This page keeps the 27 that are about agent systems: memory and context, harnesses and skills, tool use and computer-use, multi-agent coordination, persistent identity, and local open-weight models aimed at that stack. Telecom-only, medical, climate, quantum, generic eval, and vision-only papers are omitted.

PlanFence shows freshness-only executors act on obsolete plans in 30/30 controlled workflows, while dependency-scoped validation completes all without an invalid action. CONFLICTGUARD lifts conflict-task success across five GUI agents without tanking normal tasks. Speculative Macro Commit cuts Telecom latency 10.23%/18.59% versus Speculative Actions/sequential and AppWorld 7.7%/44.9%. The Civilization Framework anchors personal multi-agent systems on a sovereign ledger (exploratory temporal-weight: 54.2%→4.2% with verification). Plan Pointers finds criterion directives beat bare ids by +35.0 pts, while id+criterion can cancel (Opus 5 40/40→0/40). LOCOMO-CONV exposes conversational memory gaps QA benches miss. HookPry compromises 7/7 harnesses via attacker-controlled hook updates (≤92.5% success; defender 0% recall). Terminal-Universe builds 37.3k environments from trajectories (+11.9 TB2.1 / +13.8 MT@4 on Qwen3.5-27B); Environment Evolution adds +14.4/+18.0 pp on Qwen3.6 under long-horizon RL. SWE-Gate finds 221/644 functionally green repairs fail review constraints; PatchBench shows PoC-only validation inflates solve rate 1.83×. Five papers have a figure extracted from the PDF/HTML.


Research Papers

Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory

Evan Chen; Shiqiang Wang; Christopher G. Brinton arXiv: 2609.03340

Three-column PlanFence diagram: distributed agent memory with fresh state versus stale plan lineage, dependency-scoped validation gate, and lineage-safe execute or block outcomes
PlanFence validates plan dependencies before external actions so fresh memory cannot authorize a stale plan

PlanFence has plans cite the exact public records they used, then validates only the records that can affect the pending external action—replanning once or blocking when validation is incomplete. In 30 controlled live workflows with a post-plan revision, freshness-only executors act on the obsolete plan every time; PlanFence completes all without an invalid action. The result is about safety and systems cost, not general task accuracy.

Key insight: Fresh shared state is not a valid authorizing plan—revalidate cited dependencies before side effects.



Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents

Zhaoyuan Huang; Tianjie Ju; Pengzhou Cheng et al. arXiv: 2609.03438

Side-by-side CONFLICTGUARD figure: vanilla GUI agent blindly clicks Spotify for a pizza order versus ConflictGuard terminating on instruction-internal conflict
CONFLICTGUARD stops execution-biased overcompliance when instructions conflict with the GUI or themselves

CONFLICTGUI benchmarks instruction-internal and instruction–GUI conflicts. Agents that succeed on feasible tasks show severe execution-biased overcompliance under conflict. CONFLICTGUARD adds an inference-time feasibility verification protocol plus conditional action modulation; across five widely used agents it improves average conflict-task success while preserving normal GUI-task performance.

Key insight: Add a lightweight inference-time conflict check so GUI agents know when not to act.



Speculative Macro Commit for Faster Tool-Using Agents

Zeyu Liu; Souvik Kundu; Peter A. Beerel arXiv: 2609.03236

Speculative Macro Commit pairs a large authoritative actor with a faster speculative drafter that executes future action chains on an isolated environment snapshot, mining multi-action skeletons into a macro library and committing remaining pre-executed steps when the actor’s next tool call matches. With a Qwen3.5-27B INT4 actor and Qwen3.5-4B drafter, τ²-Bench Telecom latency falls 10.23% versus Speculative Actions and 18.59% versus sequential; AppWorld falls 7.7% and 44.9% respectively, with a small task-completion drop.

Key insight: Speculate multi-step tool macros on a snapshot, then commit when the actor’s next call matches.



The Civilization Framework: Sovereign-Anchored Communication Between Personal Multi-Agent Systems

Guangjun Liu arXiv: 2609.03425

The Civilization Framework treats the addressable party as a civilization—one human sovereign, a persistent ledger, and interchangeable agents—rather than the agent itself. The Embassy Protocol uses async ledger endpoints, with commitment state on both ledgers as ground truth, and caps authority by accessible memory plus signed credentials. A preregistered temporal-weight study reports an incorrect upstream claim first capturing 54.2% of answers without verification versus 4.2% with (1,908 trials; the registration classifies the round inconclusive pending harness-budget replication).

Key insight: Anchor personal multi-agent systems on a sovereign ledger identity, not on whichever agent is currently running.



Plan Pointers and Record-Directive Form in Budgeted Verification of Inherited Agent Memory

Kazuki Nakayashiki arXiv: 2609.03450

When an agent inherits budgeted memories and may pull at most one archived source before acting, directive form—bare id versus criterion versus both—steers retrieval. Across twelve registered studies and 14,760 attempts, length-matched criterion beat bare id by +35.0 points [+31.2, +38.8] on six direct-provider models (Study D); appending the id cancelled the criterion on three Claude models (Opus 5: 40/40→0/40). A ratification line plus a budget of two credits restored the target. Effects are descriptive of exact edits, without a mechanism claim.

Key insight: Prefer criterion directives for inherited memory; id-plus-criterion can cancel retrieval on some models.



When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents

Wen-Yu Chang; Yun-Nung Chen arXiv: 2609.03467

LOCOMO-CONV gap figure comparing conversational memory retrieval across dialog, implicit, counterfactual, and composed query styles
LOCOMO-CONV shows conversational framing exposes memory retrieval gaps that QA-style probes miss

LOCOMO-CONV builds a conversational memory bench from LoCoMo with dialog, implicit, counterfactual, and composed query styles. Across five memory systems, conversational framing exposes retrieval gaps that QA benches miss—especially on implicit and composed queries. Multi-facet query rewriting helps raw-turn memory but not abstractive memory; strong retrieval does not equal response quality, and implicit queries show silent grounding failures.

Key insight: Evaluate conversational memory in situ; retrieval score alone does not guarantee grounded replies.



A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors

Pengxun Li; Litian Zhang; Jianwei Hou et al. arXiv: 2609.03884

HookPry attack overview: plugin marketplace install, auto-update swapping a PreToolUseHook to a malicious payload, then stealth credential dump from the harness environment
HookPry shows attacker-controlled harness hook updates turning plugin auto-update into host-privileged compromise

Harness lifecycle hooks bind shell commands to session start, tool calls, and file edits with host privileges, and the update path is blindly trusted. HookPry trojanizes versioned plugins via attacker-controlled hook config across 10 attack objectives and 25 harness×backend combos (1,000 end-to-end runs): it compromises all 7 evaluated harnesses, with per-harness success up to 92.5%. A defender baseline has 0% recall; the union of three static defenses still misses 47.5%.

Key insight: Pin and verify harness hook configs—plugin auto-update is effectively code execution.



Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments

Jie Wu; Zhenru Zhang; Beichen Zhang et al. arXiv: 2609.04148

Terminal-Universe framework diagram reconstructing reusable terminal environments from agent trajectories via replay, partial workspace, and completion agent
Terminal-Universe turns agent trajectories into scalable terminal environments for single- and multi-round training

Terminal-Universe reconstructs reusable terminal environments from agent trajectories by replaying file operations into a partial workspace, then finishing with a completion agent. Tasks scale by breadth (cross-workspace) and depth (multi-round user agent), yielding 37.3k task-sufficient environments. SFT on Qwen3.5-27B improves Terminal-Bench 2.1 by +11.9 points in the single-round setting and EvoCode-Bench v2 MT@4 by +13.8 points in the multi-round setting.

Key insight: Mine existing agent trajectories into scalable terminal environments instead of hand-authoring every task.



Environment Evolution for Terminal Agents

Zhiyuan Fan; Tinghao Yu; Yuanjun Cai et al. arXiv: 2609.04128

Environment Evolution hardens terminal environments off-policy along three difficulty directions from the multi-turn learning objective, using a loop-engineered multi-agent harness. It improves Qwen3.6-27B and Qwen3.6-35B-A3B by +14.4 and +18.0 percentage points on Terminal-Bench 2.1 under simple long-horizon RL, with rollout difficulty checks also reported on Hy4 preview, Claude Opus 5, and GPT-5.6 Sol.

Key insight: Evolve terminal-environment difficulty continuously off-policy, not only on-policy rollouts.



The Natural Language Interaction Protocol and Standard for AI Agents

Luyi Xing; Rasit Onur Topaloglu; Ranjan Sinha et al. arXiv: 2609.04135

The Natural Language Interaction Protocol (NLIP) is an Ecma-standardized application-layer envelope for AI-agent interaction over HTTP, WebSocket, and AMQP. Gateways adapt among clients, agents, local context stores, ontologies, tools, enterprise services, and heterogeneous protocols, with explicit positioning relative to MCP and A2A.

Key insight: Treat NLIP as an interop envelope to watch beside MCP and A2A—not as an immediate MCP replacement.



SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents

Xin He; Yanlin Wang; Mingwei Liu et al. arXiv: 2609.04167

SWE-Gate instance construction pipeline from functional repairs to review-constraint gating
SWE-Gate shows functionally green repairs still fail review constraints

SWE-Gate builds 303 repo-level repair instances across 75 Python repos and separates functional tests from review-constraint tests. Among 644 repairs that pass functional tests, 221 fail review constraints, so functional-only evaluation overestimates full-spec repair.

Key insight: Passing functional tests is not enough—enforce review constraints in coding-agent evals.



KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents

Yaxing Lyu; Shengjie Zhou; Binbin Toh et al. arXiv: 2609.03588

KC-Bench interactive evaluation of knowledge conflicts across conflict types for LLM agents
KC-Bench stress-tests how agents handle conflicting knowledge at runtime

KC-Bench offers 238 manually screened multi-turn tasks spanning world-knowledge conflicts, input inconsistencies, and multi-source temporal conflicts, with a user simulator, stateful tools, and deterministic assertions. Across nine models (including DeepSeek-V4-Flash, GLM-5.2, and MiniMax-M3), none handle factual correction, identity consistency, and temporal conflict reliably in every setting; missed conflicts can propagate into tool calls.

Key insight: Conflict-aware agents need joint coverage of factual, identity, and temporal inconsistencies before tools fire.



Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation

Yan Tang; Tingyu Cao; Yuanbo Tang et al. arXiv: 2609.03727

This survey frames proactive service as a POMDP constrained by authorization and risk, with actions remain-silent / ask / assist / act and a pipeline of state-and-need estimation, intervention gating, action construction, and feedback adaptation. It argues offline classification is not deployment benefit, and that long-term memory is not defining of proactivity—calibrated intervention value, verifiable authorization, and recoverable execution are.

Key insight: Proactivity needs authorization-gated intervention, not memory capacity alone.



Dalek: A Constructive Agent Machine

Wanpeng Xie arXiv: 2609.03546

Dalek specifies a closed constructive agent machine from actors, messages, and channels with four obligations—host boundary, construction language, admissible transitions, and rule heredity—combining a von Neumann hereditary core with an LLM/compiler as capability producer. New capabilities are authored, compiled, and installed into the description, then inherited by descendants, including the machine’s own organs and runtime.

Key insight: Specify self-extending harnesses with explicit host boundaries and hereditary capability install paths.



Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models

Xingming Long; Yu Liu; Zhiwei Yang et al. arXiv: 2609.03493

NTEP annotates essential evidence and tool calls; NTEP-R rewards pre-call intent alignment and post-call observation summarization, with a non-repeated-goal regularizer. NTEP-8B improves search-oriented accuracy and tool-use efficiency in a three-tool crop/search/text framework across seven image-grounded benches.

Key insight: Reward necessary evidence-path progress in tool training, not only the final answer.



DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training

Shubham Gandhi; Saurabh Goyal; Kiran Kate et al. arXiv: 2609.04094

DRACO dynamic rubrics for fine-grained credit assignment in long-horizon agent training
DRACO assigns fine-grained credit with dynamic rubrics over long horizons

DRACO targets outcome-blind settings with dynamic rubrics during training, redistributing trajectory score into per-step advantages for GRPO in closed form without a trained attribution model. On AppWorld it reports +15.9 over base and +5.3 over GRPO with sparse ground-truth reward despite using no verifiers; Tau-Bench OOD is +5.3 over base without a frontier judge.

Key insight: Use dynamic rubrics to assign per-step credit when programmatic checkers are unavailable.



DNative-Twin: Decision Graphs and Digital Twins for Reconstructable Agentic Decisions

Junjie Pang; Zhenzhen Xie; Haoke Han et al. arXiv: 2609.03787

DNative-Twin records committed decisions as typed trajectories in a graph-native digital twin and re-executes under declared conditions. The graph localizes represented changes but cannot determine consequences of unobserved tool state. On 300 injected instances, unresolved-divergence recall rises from 0 to 0.667 with replay-contract state and to 1.0 with verification; BPI 2020 median end-to-end latency goes from 0.794s to 8.889s at 500–5,000 cases.

Key insight: Separate graph localization, replay contracts, and verification evidence when auditing agent decisions.



Value-Preserving Architectures for Agentic AI Systems

Alessandro Pesare; Tommaso Dolci; Katja Hose et al. arXiv: 2609.03920

This position paper argues that architectural choices—coordination, protocols, topologies—shape privacy, fairness, and safety outcomes, sketching patterns for privacy-aware federated topology, distributed pluralism, and a guard-agent for unfairness as a foundation for a pattern catalog.

Key insight: Treat multi-agent topology and protocol choices as first-class value-preserving design levers.



Privacy-Preserving Topology-Guided Safety for LLM-Based Multi-Agent Systems via Federated Graph Learning

Jinxi Yu; Eric Hanchen Jiang; Levina Li et al. arXiv: 2609.02967

FGLGuard is a privacy-preserving topology-guided multi-agent safeguard via a federated GNN over communication graphs without pooling traces. On Agent-SafetyBench, R-Judge, and AgentDojo, federated FGLGuard exceeds the in-domain centralized ceiling without pooling; live AgentDojo attack success drops 43% at near-unguarded utility with zero API cost.

Key insight: Federate graph safeguards over communication topology instead of pooling raw agent traces.



Remember and Reweight: Enhancing Multi-Agent Debate with Experience Memory and Confidence Estimation

Xuanfa Jin; Zhijian Ma; Yongcheng Zeng et al. arXiv: 2609.03619

Remember and Reweight targets shared misconception in multi-agent debate—when the majority is wrong, debate amplifies error—by adding experience memory from past debates, debate-state-aware retrieval to calibrate concept priors, and confidence weights to modulate peer influence. The paper reports consistent gains over single-agent and multi-agent-debate baselines (absolute deltas are light in the abstract).

Key insight: Calibrate debate with experience memory and confidence weights rather than raw majority vote.



Counterexamples as Feedback for Agent Self-Correction

Sidhesh Badrinarayan; Adithya Parthasarathy arXiv: 2609.02892

This work studies NL→regex multi-turn refinement with deterministic oracle counterexamples. On 30 NL-RX-Turk tasks, diagnostic counterexamples solve 90% within 4 turns versus 17% zero-shot, 27% generic self-correction, and 23% error-only. Full diagnostic-plus-hardening solves all hidden tasks with mean time-to-solve 2.7 turns and 77% robust success.

Key insight: Prefer concrete counterexample feedback over vague retry prompts in agent self-correction.



PatchBench: Evaluating AI Agents for Vulnerability Patching

Chihao Shen; Jiacheng Li; Aastha Mahajan et al. arXiv: 2609.04075

PatchBench argues PoC-only patch validation is inflated: about 25% of agent patches are substantially similar to historical developer patches (memorization), and agents also suppress crashes via stack-trace patches. The bench transplants and mutates vulnerabilities and validates security plus semantic correctness; across 11 agents, PoC-only validation inflates solve rate 1.83× on average.

Key insight: Do not trust PoC-green patch claims—require security-and-semantic validation beyond crash suppression.



SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center

Uday Vallabhaneni; Cassie L. Cagwin; David J. Wild arXiv: 2609.04159

SENTINEL-RL offloads topological reasoning: a heterogeneous GAT summarizes a live auth subgraph, PPO maps to constrained investigative actions, and the LLM only narrates when gated by a critic. On LANL and IU Quartz, CREATE ingests a 24M-edge subgraph in 14.2 minutes (~24× versus MERGE); the alert engine stays ≤2.5s; PPO return is 8.74±0.31 with precision 0.91 / recall 0.87; detect-to-approve median is 6.3s.

Key insight: Offload graph topology to constrained RL and keep the LLM for gated narrative plus human approval.



Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection

Weijie Liu; Running Zhao; Wenhao Yuan et al. arXiv: 2609.03416

Dude is a dual-detection multi-agent system for paper–code discrepancy detection with granularity-aligned negotiation and two-stage salience filtering to cut false positives from paper/code granularity asymmetry. It reports up to +22.8% recall/precision and +18.7% F1 versus baselines on real paper–code discrepancy datasets.

Key insight: Align paper and code granularity before scoring discrepancies to cut false positives.



GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving

Qiankun Ma; Yanjiang Zhou; Zinan Xiong et al. arXiv: 2609.03494

GrowPage treats KV capacity as a runtime resource: dual-timescale query summaries estimate demand; at a capacity boundary the system either compresses within the allocation or acquires another PagedAttention page, preserving continuous batching and prefix caching. The abstract reports a superior performance–throughput trade-off on reasoning benches without absolute percentages.

Key insight: Budget KV pages dynamically for long reasoning instead of fixing per-request capacity.



What Matters for Aggressive Decoding-Time KV Eviction? Temporal Aggregation and Ranking Preservation

Bo Zeng; Yu Zhao; Yefeng Liu et al. arXiv: 2609.03515

Under aggressive decode-time KV compression, EMA temporal aggregation makes order-preserving scorer tweaks largely indistinguishable. InertiaKV and InertiaKV-Lazy (periodic refresh) reach 1.34–1.46× decode throughput versus full-refresh InertiaKV; Score-Free scores once at the first decode step and freezes ranking with average quality change +0.03 across six backbones on LongBench, LongBench-v2, and RULER.

Key insight: Invest in temporal aggregation and ranking preservation for decode-time KV eviction, not endless scorers.



From Prior-Guided Heuristics to Deployable Agents: Accelerating Demonstration-Driven Reinforcement Learning for Deadline-Constrained Network Control

Vincenzo Norman Vitale; Mohammad Solki; Antonia Maria Tulino et al. arXiv: 2609.03590

This networking paper combines deadline-aware Effective Congestion metrics with a MADRL hybrid scheduler/router and Model-Guided Annealed RL (MGA-RL) that unifies behavior cloning, offline, online, and offline-to-online training on a DDPG backbone for deadline-constrained network control.

Key insight: Demonstration-driven annealed RL can transfer prior heuristics into deployable control agents in networked settings.