Friday’s cs.AI listing had 165 new submissions and cross-lists. This page keeps the 27 that are about agent systems: memory and context, harnesses and skills, tool use and computer-use, multi-agent coordination, persistent identity, and local open-weight models aimed at that stack. Telecom-only, medical, climate, quantum, generic eval, and vision-only papers are omitted.
PlanFence shows freshness-only executors act on obsolete plans in 30/30 controlled workflows, while dependency-scoped validation completes all without an invalid action. CONFLICTGUARD lifts conflict-task success across five GUI agents without tanking normal tasks. Speculative Macro Commit cuts Telecom latency 10.23%/18.59% versus Speculative Actions/sequential and AppWorld 7.7%/44.9%. The Civilization Framework anchors personal multi-agent systems on a sovereign ledger (exploratory temporal-weight: 54.2%→4.2% with verification). Plan Pointers finds criterion directives beat bare ids by +35.0 pts, while id+criterion can cancel (Opus 5 40/40→0/40). LOCOMO-CONV exposes conversational memory gaps QA benches miss. HookPry compromises 7/7 harnesses via attacker-controlled hook updates (≤92.5% success; defender 0% recall). Terminal-Universe builds 37.3k environments from trajectories (+11.9 TB2.1 / +13.8 MT@4 on Qwen3.5-27B); Environment Evolution adds +14.4/+18.0 pp on Qwen3.6 under long-horizon RL. SWE-Gate finds 221/644 functionally green repairs fail review constraints; PatchBench shows PoC-only validation inflates solve rate 1.83×. Five papers have a figure extracted from the PDF/HTML.
Evan Chen; Shiqiang Wang; Christopher G. Brinton arXiv: 2609.03340
PlanFence has plans cite the exact public records they used, then validates only the records that can affect the pending external action—replanning once or blocking when validation is incomplete. In 30 controlled live workflows with a post-plan revision, freshness-only executors act on the obsolete plan every time; PlanFence completes all without an invalid action. The result is about safety and systems cost, not general task accuracy.
Key insight: Fresh shared state is not a valid authorizing plan—revalidate cited dependencies before side effects.
Zhaoyuan Huang; Tianjie Ju; Pengzhou Cheng et al. arXiv: 2609.03438
CONFLICTGUI benchmarks instruction-internal and instruction–GUI conflicts. Agents that succeed on feasible tasks show severe execution-biased overcompliance under conflict. CONFLICTGUARD adds an inference-time feasibility verification protocol plus conditional action modulation; across five widely used agents it improves average conflict-task success while preserving normal GUI-task performance.
Key insight: Add a lightweight inference-time conflict check so GUI agents know when not to act.
Zeyu Liu; Souvik Kundu; Peter A. Beerel arXiv: 2609.03236
Speculative Macro Commit pairs a large authoritative actor with a faster speculative drafter that executes future action chains on an isolated environment snapshot, mining multi-action skeletons into a macro library and committing remaining pre-executed steps when the actor’s next tool call matches. With a Qwen3.5-27B INT4 actor and Qwen3.5-4B drafter, τ²-Bench Telecom latency falls 10.23% versus Speculative Actions and 18.59% versus sequential; AppWorld falls 7.7% and 44.9% respectively, with a small task-completion drop.
Key insight: Speculate multi-step tool macros on a snapshot, then commit when the actor’s next call matches.
Guangjun Liu arXiv: 2609.03425
The Civilization Framework treats the addressable party as a civilization—one human sovereign, a persistent ledger, and interchangeable agents—rather than the agent itself. The Embassy Protocol uses async ledger endpoints, with commitment state on both ledgers as ground truth, and caps authority by accessible memory plus signed credentials. A preregistered temporal-weight study reports an incorrect upstream claim first capturing 54.2% of answers without verification versus 4.2% with (1,908 trials; the registration classifies the round inconclusive pending harness-budget replication).
Key insight: Anchor personal multi-agent systems on a sovereign ledger identity, not on whichever agent is currently running.
Kazuki Nakayashiki arXiv: 2609.03450
When an agent inherits budgeted memories and may pull at most one archived source before acting, directive form—bare id versus criterion versus both—steers retrieval. Across twelve registered studies and 14,760 attempts, length-matched criterion beat bare id by +35.0 points [+31.2, +38.8] on six direct-provider models (Study D); appending the id cancelled the criterion on three Claude models (Opus 5: 40/40→0/40). A ratification line plus a budget of two credits restored the target. Effects are descriptive of exact edits, without a mechanism claim.
Key insight: Prefer criterion directives for inherited memory; id-plus-criterion can cancel retrieval on some models.
Wen-Yu Chang; Yun-Nung Chen arXiv: 2609.03467
LOCOMO-CONV builds a conversational memory bench from LoCoMo with dialog, implicit, counterfactual, and composed query styles. Across five memory systems, conversational framing exposes retrieval gaps that QA benches miss—especially on implicit and composed queries. Multi-facet query rewriting helps raw-turn memory but not abstractive memory; strong retrieval does not equal response quality, and implicit queries show silent grounding failures.
Key insight: Evaluate conversational memory in situ; retrieval score alone does not guarantee grounded replies.
Pengxun Li; Litian Zhang; Jianwei Hou et al. arXiv: 2609.03884
Harness lifecycle hooks bind shell commands to session start, tool calls, and file edits with host privileges, and the update path is blindly trusted. HookPry trojanizes versioned plugins via attacker-controlled hook config across 10 attack objectives and 25 harness×backend combos (1,000 end-to-end runs): it compromises all 7 evaluated harnesses, with per-harness success up to 92.5%. A defender baseline has 0% recall; the union of three static defenses still misses 47.5%.
Key insight: Pin and verify harness hook configs—plugin auto-update is effectively code execution.
Jie Wu; Zhenru Zhang; Beichen Zhang et al. arXiv: 2609.04148
Terminal-Universe reconstructs reusable terminal environments from agent trajectories by replaying file operations into a partial workspace, then finishing with a completion agent. Tasks scale by breadth (cross-workspace) and depth (multi-round user agent), yielding 37.3k task-sufficient environments. SFT on Qwen3.5-27B improves Terminal-Bench 2.1 by +11.9 points in the single-round setting and EvoCode-Bench v2 MT@4 by +13.8 points in the multi-round setting.
Key insight: Mine existing agent trajectories into scalable terminal environments instead of hand-authoring every task.
Zhiyuan Fan; Tinghao Yu; Yuanjun Cai et al. arXiv: 2609.04128
Environment Evolution hardens terminal environments off-policy along three difficulty directions from the multi-turn learning objective, using a loop-engineered multi-agent harness. It improves Qwen3.6-27B and Qwen3.6-35B-A3B by +14.4 and +18.0 percentage points on Terminal-Bench 2.1 under simple long-horizon RL, with rollout difficulty checks also reported on Hy4 preview, Claude Opus 5, and GPT-5.6 Sol.
Key insight: Evolve terminal-environment difficulty continuously off-policy, not only on-policy rollouts.
Luyi Xing; Rasit Onur Topaloglu; Ranjan Sinha et al. arXiv: 2609.04135
The Natural Language Interaction Protocol (NLIP) is an Ecma-standardized application-layer envelope for AI-agent interaction over HTTP, WebSocket, and AMQP. Gateways adapt among clients, agents, local context stores, ontologies, tools, enterprise services, and heterogeneous protocols, with explicit positioning relative to MCP and A2A.
Key insight: Treat NLIP as an interop envelope to watch beside MCP and A2A—not as an immediate MCP replacement.
Xin He; Yanlin Wang; Mingwei Liu et al. arXiv: 2609.04167
SWE-Gate builds 303 repo-level repair instances across 75 Python repos and separates functional tests from review-constraint tests. Among 644 repairs that pass functional tests, 221 fail review constraints, so functional-only evaluation overestimates full-spec repair.
Key insight: Passing functional tests is not enough—enforce review constraints in coding-agent evals.
Yaxing Lyu; Shengjie Zhou; Binbin Toh et al. arXiv: 2609.03588
KC-Bench offers 238 manually screened multi-turn tasks spanning world-knowledge conflicts, input inconsistencies, and multi-source temporal conflicts, with a user simulator, stateful tools, and deterministic assertions. Across nine models (including DeepSeek-V4-Flash, GLM-5.2, and MiniMax-M3), none handle factual correction, identity consistency, and temporal conflict reliably in every setting; missed conflicts can propagate into tool calls.
Key insight: Conflict-aware agents need joint coverage of factual, identity, and temporal inconsistencies before tools fire.
Yan Tang; Tingyu Cao; Yuanbo Tang et al. arXiv: 2609.03727
This survey frames proactive service as a POMDP constrained by authorization and risk, with actions remain-silent / ask / assist / act and a pipeline of state-and-need estimation, intervention gating, action construction, and feedback adaptation. It argues offline classification is not deployment benefit, and that long-term memory is not defining of proactivity—calibrated intervention value, verifiable authorization, and recoverable execution are.
Key insight: Proactivity needs authorization-gated intervention, not memory capacity alone.
Wanpeng Xie arXiv: 2609.03546
Dalek specifies a closed constructive agent machine from actors, messages, and channels with four obligations—host boundary, construction language, admissible transitions, and rule heredity—combining a von Neumann hereditary core with an LLM/compiler as capability producer. New capabilities are authored, compiled, and installed into the description, then inherited by descendants, including the machine’s own organs and runtime.
Key insight: Specify self-extending harnesses with explicit host boundaries and hereditary capability install paths.
Xingming Long; Yu Liu; Zhiwei Yang et al. arXiv: 2609.03493
NTEP annotates essential evidence and tool calls; NTEP-R rewards pre-call intent alignment and post-call observation summarization, with a non-repeated-goal regularizer. NTEP-8B improves search-oriented accuracy and tool-use efficiency in a three-tool crop/search/text framework across seven image-grounded benches.
Key insight: Reward necessary evidence-path progress in tool training, not only the final answer.
Shubham Gandhi; Saurabh Goyal; Kiran Kate et al. arXiv: 2609.04094
DRACO targets outcome-blind settings with dynamic rubrics during training, redistributing trajectory score into per-step advantages for GRPO in closed form without a trained attribution model. On AppWorld it reports +15.9 over base and +5.3 over GRPO with sparse ground-truth reward despite using no verifiers; Tau-Bench OOD is +5.3 over base without a frontier judge.
Key insight: Use dynamic rubrics to assign per-step credit when programmatic checkers are unavailable.
Junjie Pang; Zhenzhen Xie; Haoke Han et al. arXiv: 2609.03787
DNative-Twin records committed decisions as typed trajectories in a graph-native digital twin and re-executes under declared conditions. The graph localizes represented changes but cannot determine consequences of unobserved tool state. On 300 injected instances, unresolved-divergence recall rises from 0 to 0.667 with replay-contract state and to 1.0 with verification; BPI 2020 median end-to-end latency goes from 0.794s to 8.889s at 500–5,000 cases.
Key insight: Separate graph localization, replay contracts, and verification evidence when auditing agent decisions.
Alessandro Pesare; Tommaso Dolci; Katja Hose et al. arXiv: 2609.03920
This position paper argues that architectural choices—coordination, protocols, topologies—shape privacy, fairness, and safety outcomes, sketching patterns for privacy-aware federated topology, distributed pluralism, and a guard-agent for unfairness as a foundation for a pattern catalog.
Key insight: Treat multi-agent topology and protocol choices as first-class value-preserving design levers.
Jinxi Yu; Eric Hanchen Jiang; Levina Li et al. arXiv: 2609.02967
FGLGuard is a privacy-preserving topology-guided multi-agent safeguard via a federated GNN over communication graphs without pooling traces. On Agent-SafetyBench, R-Judge, and AgentDojo, federated FGLGuard exceeds the in-domain centralized ceiling without pooling; live AgentDojo attack success drops 43% at near-unguarded utility with zero API cost.
Key insight: Federate graph safeguards over communication topology instead of pooling raw agent traces.
Xuanfa Jin; Zhijian Ma; Yongcheng Zeng et al. arXiv: 2609.03619
Remember and Reweight targets shared misconception in multi-agent debate—when the majority is wrong, debate amplifies error—by adding experience memory from past debates, debate-state-aware retrieval to calibrate concept priors, and confidence weights to modulate peer influence. The paper reports consistent gains over single-agent and multi-agent-debate baselines (absolute deltas are light in the abstract).
Key insight: Calibrate debate with experience memory and confidence weights rather than raw majority vote.
Sidhesh Badrinarayan; Adithya Parthasarathy arXiv: 2609.02892
This work studies NL→regex multi-turn refinement with deterministic oracle counterexamples. On 30 NL-RX-Turk tasks, diagnostic counterexamples solve 90% within 4 turns versus 17% zero-shot, 27% generic self-correction, and 23% error-only. Full diagnostic-plus-hardening solves all hidden tasks with mean time-to-solve 2.7 turns and 77% robust success.
Key insight: Prefer concrete counterexample feedback over vague retry prompts in agent self-correction.
Chihao Shen; Jiacheng Li; Aastha Mahajan et al. arXiv: 2609.04075
PatchBench argues PoC-only patch validation is inflated: about 25% of agent patches are substantially similar to historical developer patches (memorization), and agents also suppress crashes via stack-trace patches. The bench transplants and mutates vulnerabilities and validates security plus semantic correctness; across 11 agents, PoC-only validation inflates solve rate 1.83× on average.
Key insight: Do not trust PoC-green patch claims—require security-and-semantic validation beyond crash suppression.
Uday Vallabhaneni; Cassie L. Cagwin; David J. Wild arXiv: 2609.04159
SENTINEL-RL offloads topological reasoning: a heterogeneous GAT summarizes a live auth subgraph, PPO maps to constrained investigative actions, and the LLM only narrates when gated by a critic. On LANL and IU Quartz, CREATE ingests a 24M-edge subgraph in 14.2 minutes (~24× versus MERGE); the alert engine stays ≤2.5s; PPO return is 8.74±0.31 with precision 0.91 / recall 0.87; detect-to-approve median is 6.3s.
Key insight: Offload graph topology to constrained RL and keep the LLM for gated narrative plus human approval.
Weijie Liu; Running Zhao; Wenhao Yuan et al. arXiv: 2609.03416
Dude is a dual-detection multi-agent system for paper–code discrepancy detection with granularity-aligned negotiation and two-stage salience filtering to cut false positives from paper/code granularity asymmetry. It reports up to +22.8% recall/precision and +18.7% F1 versus baselines on real paper–code discrepancy datasets.
Key insight: Align paper and code granularity before scoring discrepancies to cut false positives.
Qiankun Ma; Yanjiang Zhou; Zinan Xiong et al. arXiv: 2609.03494
GrowPage treats KV capacity as a runtime resource: dual-timescale query summaries estimate demand; at a capacity boundary the system either compresses within the allocation or acquires another PagedAttention page, preserving continuous batching and prefix caching. The abstract reports a superior performance–throughput trade-off on reasoning benches without absolute percentages.
Key insight: Budget KV pages dynamically for long reasoning instead of fixing per-request capacity.
Bo Zeng; Yu Zhao; Yefeng Liu et al. arXiv: 2609.03515
Under aggressive decode-time KV compression, EMA temporal aggregation makes order-preserving scorer tweaks largely indistinguishable. InertiaKV and InertiaKV-Lazy (periodic refresh) reach 1.34–1.46× decode throughput versus full-refresh InertiaKV; Score-Free scores once at the first decode step and freezes ranking with average quality change +0.03 across six backbones on LongBench, LongBench-v2, and RULER.
Key insight: Invest in temporal aggregation and ranking preservation for decode-time KV eviction, not endless scorers.
Vincenzo Norman Vitale; Mohammad Solki; Antonia Maria Tulino et al. arXiv: 2609.03590
This networking paper combines deadline-aware Effective Congestion metrics with a MADRL hybrid scheduler/router and Model-Guided Annealed RL (MGA-RL) that unifies behavior cloning, offline, online, and offline-to-online training on a DDPG backbone for deadline-constrained network control.
Key insight: Demonstration-driven annealed RL can transfer prior heuristics into deployable control agents in networked settings.