Wednesday’s cs.AI listing had 517 new submissions and cross-lists. This page keeps the 45 that are about agent systems: memory and context, harnesses and skills, tool use and computer-use, multi-agent coordination, persistent identity, and local open-weight models aimed at that stack. Telecom-only, medical, climate, quantum, generic eval, and vision-only papers are omitted.

MERIT cost-accounts memory for tool-using agents (0.00→0.55–1.00 dependent-task success; up to 60-point swings by memory implementation). EdgeMem preserves source turns in a multi-anchor hypergraph (LoCoMo 61.01 vs 58.70) without generative memory writes. DroidTool self-generates Android tool actions beside GUI steps. ResidualAuth shows revocation needs residual authorization state beyond current permissions. Beyond Prompts treats harness search as resource-bounded selection with RelLift95(B). FrogNano trains a 4B SWE agent with frontier-calibrated synthetic RL only. MemForest compresses memory via EventTrees; Scaffold grows recursive parametric web skills; co-evolving harnesses warns that full expert imitation under an evolved harness can regress 4–30 points unless correction stays on-policy. Eight papers have a figure extracted from the HTML/PDF.


Research Papers

When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents

Shweta Mishra; Shashank Mishra arXiv: 2609.05441

Three-panel line charts of dependent-task success versus difficulty for commerce, IT ops, and assistant domains under six memory conditions
MERIT: memory condition versus dependent-task success across domains and difficulties

MERIT measures whether long-term memory changes what tool-using agents do, with cost metering, rather than conversational recall alone. Across 23,440 episodes it lifts dependent-task success from a leak-verified floor of 0.00 to 0.55–1.00; on updated facts, embedding retrieval is unstable (0.30–0.95) while update-on-write stores stay at 0.70–1.00, and swapping a memory implementation moves success by up to 60 points.

Key insight: Measure memory by task utility and dollar cost, not dialogue recall alone.



AutoFyn Technical Report: Non-Parametric Expert Iteration for Long-Horizon Agents

Adib Hasan; Daniel Schaffield; Akashnil Dutta et al. arXiv: 2609.05446

AutoFyn adapts a frozen model across rounds by updating persistent state from verified rewards instead of weights. Each round starts from a fresh session; durable information returns only through explicit interfaces such as memory files, reports, and repository state, while an orchestrator explores alternatives and a task-grounded verifier scores progress.

Key insight: Adapt frozen agents via verified persistent state, not weight updates.



ResidualAuth: What Authorization State Must Language Agents Preserve under Revocable Delegation?

Moonwon Choi; Seokho Jeong; Seunggeun Lee arXiv: 2609.08062

ResidualAuth shows two authorization histories can share identical current permissions and reachability yet require opposite decisions after the same revocation. It formalizes residual authorization state, proves exponentially many future-distinct states can share one transitive closure, and compiles the constructions into paired language-agent episodes across open models.

Key insight: Preserve residual authorization history, not only the current permission set.



DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents

Yu Liu; Zhilin Liu; Zhiwei Yang et al. arXiv: 2609.06059

DAREBench evaluates models as agents on multimodal perception, multi-step execution, tool use, and artifact delivery inside a shared OpenClaw environment. It organizes 233 tasks to capture workload variation and support deployment-aware comparison beyond static answer correctness.

Key insight: Judge agents on deployment workloads and artifacts, not static QA alone.



AURA-Eval: Evaluation Framework for Acting Under Risk Awareness in LLM Agent Trajectories

Ruoxi Shang; Christina-Maria Androna; Orfeas Menis Mastromichalakis et al. arXiv: 2609.06783

AURA-Eval diagnoses risk recognition, pre-action detection, and safe completion when a safe path exists, instead of collapsing safety to one score. From 157 trajectories it builds 1,249 items and evaluates 20 frontier and open-weight models with rubrics over tool-use trajectories.

Key insight: Score agent safety at decision points, not as a single aggregate.



Improving Proficiency and Efficiency of Android GUI Agents via Self-Generating Tool Actions

Juyong Lee; Woogyeol Jin; Kimin Lee arXiv: 2609.06792

Diagram comparing multi-step Android GUI actions to a direct set_wikipedia_setting tool action, plus a design-to-repair tool-generation pipeline
DroidTool: hybrid GUI versus self-generated tool actions with an agentic create-and-repair loop

DroidTool lets Android GUI agents self-generate tool actions as Python functions over app state, using proposal, implementation, test generation, and repair with relational tests across tools. The hybrid GUI-plus-API action space targets proficiency and efficiency without hand-building every tool.

Key insight: Let mobile agents mint and verify their own API tools beside GUI actions.



AgentLeak: Cloning Stronger LLM Agent Capabilities onto Weaker Agents Beyond Skill Stealing

Xiaoting Lyu; Yuhong Wu; Yufei Han et al. arXiv: 2609.07131

AgentLeak asks whether a weaker black-box attacker can clone a stronger proprietary agent’s capabilities beyond stealing explicit skill artifacts. Artifact leakage alone may not transfer capability when the attacker lacks implicit procedural behaviors acquired through execution.

Key insight: Capability theft can require procedural behavior, not only skill files.



SCAFFOLD: Self-Improving Web Agents via Recursive Parametric Skill Abstraction

Bowei He; Xiaokun Zhang; Meng Ding et al. arXiv: 2609.05511

Self-improvement loop: web rollout, multi-instance parametric skill induction, recursive skill hierarchy, LoRA distillation, and MDL library compaction
Scaffold: recursive parametric skills from web trajectories with periodic MDL compaction

Scaffold induces parametric executable skills from successful web-agent trajectories under a multi-instance abstraction constraint and maintains a recursively composed hierarchy. Unlike flat prompt-side skill caches, it compresses redundancy and composes skills for visually rich, long-horizon sites.

Key insight: Grow recursive parametric skills from web trajectories, not flat caches.



Beyond "AI Helps Humans": Decision-Targeted Evaluation Design for Human-Agent Teams in the Agentic Era

Hamed Khosravi; Xiaoming Huo arXiv: 2609.05527

Human–agent teams are worth keeping only if they beat human-alone and agent-alone alternatives, but those counterfactuals are costly to replay. Under a fixed replay budget the design question is which tasks get human-only versus agent-only replays; prior methods do not target that decision directly.

Key insight: Allocate scarce replays to the human-vs-agent decision that matters.



EdgeMem: LLM-Free Agent Memory Construction and Retrieval via Evidence-Preserving Multi-Anchor Hypergraph

Zeyang Cui; Jiannong Cao; Zhiyuan Wen et al. arXiv: 2609.05553

Three-column diagram comparing LLM summarization, LLM retrieval, and EdgeMem multi-anchor hypergraph on a hotel-floor preference constraint
EdgeMem keeps source evidence retrievable where summarization and retrieval drop constraints

EdgeMem keeps original interaction turns and organizes them with content, temporal, and episodic anchors in a multi-anchor hypergraph built by lightweight local processing. Retrieval returns source evidence and reserves the LLM for final answers; on LoCoMo it leads seven systems under a shared prompt (61.01 vs 58.70) with no generative-LLM calls in construction or retrieval.

Key insight: Preserve source evidence in a hypergraph; generate only at answer time.



EnvCraft: Synthesizing Executable Environments in Agentic RL for Claw-like Agent

Yirong Zeng; Shen You; Jinhang Feng et al. arXiv: 2609.05576

EnvCraft synthesizes executable environments and scalable training data for claw-like agents that act across stateful workspaces, going beyond tool-calling endpoints. The goal is to unblock agentic RL where interactive environments are scarce.

Key insight: Synthesize full executable workspaces for agentic RL, not tool stubs alone.



Agents Trust Tools Too Much: Measuring Reliance on Unreliable Tools

Hoyeol Yang; Woojung Song; Taewon Kim et al. arXiv: 2609.05587

Fourteen LLMs are tested with corrupted returns from web search, sub-agent delegation, and code execution. Mean adoption of corrupted content exceeds one third for every tool and reaches 68% in the worst setting, showing overtrust when tools look plausible but wrong.

Key insight: Measure tool overtrust under corrupted returns, not only task success.



Beyond Prompts: Measuring and Optimizing LLM Tool-Agent Harnesses

Cen (Mia) Zhao; Haibo Ruan; Wenjie Chen et al. arXiv: 2609.05736

PRISM optimization flow: seed, multi-generation frontier pick/reconstruct/analyze/mutate/eval/crossover/update, then final scorecard
Beyond Prompts / PRISM: budgeted harness search over a mutating frontier

The work treats harness selection—prompts and tool-boundary middleware around a fixed model—as a resource-bounded search. An optimizer-agnostic protocol reports mean and worst-condition lift, repeatability, cost diagnostics, and RelLift95(B), the conservative held-out gain of the harness chosen under budget B.

Key insight: Optimize the harness under budget; report conservative held-out lift.



Beyond Top-$k$ Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents

Wang Wei; Tiankai Yang; Samyadeep Basu et al. arXiv: 2609.05824

Diverse Skill Routing reranks large skill registries with a Determinantal Point Process that balances relevance and non-redundancy via a query-residual diversity kernel. The aim is to stop wasting context on redundant skills when complex tasks need complementary sets.

Key insight: Route skills for diversity of coverage, not only top-k relevance.



AgentBrew: Offline Tool-Use Agent Learning from Raw Real-World Trajectories

Zhiyi Lyu; Yewen Li; Longtao Zheng et al. arXiv: 2609.05837

AgentBrew learns tool-use policies offline from one batch of raw trajectories without task verifiers or iterative on-policy rollouts. After unfiltered exploration, it extracts training signal from noisy corpora for environment-specific agents where simulators and budgets are limited.

Key insight: Train tool agents offline from raw trajectories when verifiers are absent.



Evaluating Deep-Search Agents under Hierarchical Web Evidence Poisoning

Zhongan Bi; Qiwen Wang; Jianrong Jiang et al. arXiv: 2609.06027

HAE-GEO tracks deep-search agents from exposure through verification, revision, and recovery under hierarchical web evidence poisoning (direct assertion, camouflage, and apparent corroboration). It goes beyond measuring whether poisoned content is merely retrieved or endorsed.

Key insight: Score poisoning resistance by recovery along the full search trajectory.



Typed Federated Artifacts for the Agentic Web:Sharing Tool-Routing Knowledge Across Frozen,Heterogeneous LLM Agents

Abhijit Chakraborty; Ni Trieu; Vivek Gupta arXiv: 2609.06815

Typed federated artifacts share schema-validated tool-routing knowledge across frozen, heterogeneous agents without transferring weights. SYNAPSE1 is instantiated as shared routing knowledge after cleaning garbage and training items, aiming at cross-model transfer with per-field privacy and dispute handling.

Key insight: Share typed tool-routing artifacts across frozen multi-vendor agents.



PRIMUS: Identity, Governance, and Verification for Multi-Agent Federations

Sasank Annapureddy; Anjaneya Prasad Thamatani arXiv: 2609.07910

PRIMUS couples prime-power agent identity with BLS aggregate signatures, derives a safe-kill threshold that cuts false-positive termination from 80% to 0.00% under 10% channel noise, and asks whether the same verification machinery can steer generate-and-test toward better answers under adversarial federation conditions.

Key insight: Govern multi-agent federations with signed identity and safe-kill thresholds.



FrogNano: Training a 4B Coding Agent via Online Task Synthesis

Minseon Kim; Zhengyan Shi; Emiliano Penaloza et al. arXiv: 2609.07925

Line chart of SWE-bench Verified resolved percent versus cumulative training tasks for FrogNano versus RL baselines across five iterations
FrogNano climbs SWE-bench Verified over 1,500 synthetic RL tasks versus Rebench baselines

FrogNano is a 4B coding agent post-trained only with RL on about 1,500 SWE environments using synthetic tasks at the learnability frontier of the current checkpoint. Results support competitive small coding agents without distillation from larger models, runnable on minimal hardware.

Key insight: Train small coding agents with frontier-calibrated synthetic RL tasks alone.



SE-GoS: Self-Evolving Graph-of-Skills for Skill Library at Scale

Dawei Fu; Cheng Jiang; Sitian Qian et al. arXiv: 2609.08228

SE-GoS evolves an existing Graph-of-Skills from execution traces without training while preserving the graph for scalable retrieval. It asks whether historical traces can be distilled into a better retrieval graph that generalizes to unseen tasks.

Key insight: Evolve skill-retrieval graphs from traces without retraining the model.



Personalizing LLM Agent Memory Using Biometrics

Yanhong Qian; Qingguo Meng; Shihao Ding et al. arXiv: 2609.08558

Bio-Memory augments A-Mem notes with biometric embeddings and filters the retrieval pool by biometric match before semantic ranking. On LoCoMo in a 10-user shared setting it targets identity-aware personalization so the requester matches the memory owner.

Key insight: Gate personalized memory retrieval on biometric match, then semantics.



Graph-Based Personalized Memory for LLM Agents: Representation, Evolution, Retrieval, and Evaluation

Dac Duy Anh Nguyen; Zhangchi Qiu; Shigeng Chen et al. arXiv: 2609.08599

This survey frames graph-based personalized memory for long-term LLM agents: representation, evolution, retrieval, and evaluation of preferences, goals, constraints, and experiences that change over time. Explicit relations, temporal context, and evidence links structure what is remembered and how it is revised.

Key insight: Treat personalized agent memory as an evolving evidence-linked graph.



Closing the Consistency Gap: Self-Evolving Agents That Learn to Stay on Course

Evelyn Duesterwald; Benjamin Elder; Lilian Ngweta et al. arXiv: 2609.08832

A ReAct agent on AppWorld with GPT-4.1 averages 77% per-run pass rate but succeeds in all five repeats only 53% of the time—a 24-point consistency gap. A self-evolving framework converts unstable low-consistency steps into episodic memory for future runs.

Key insight: Close the multi-run consistency gap with memory of unstable steps.



SkillAdam: Stable and Efficient Skill Evolution for Agents

Gaoyuan Li; Meihao Fan; Yizhe Liu et al. arXiv: 2609.08944

SkillAdam targets stable, efficient skill self-evolution for frozen agents by addressing direction stability (accumulate corrections) and update adaptivity (step size from feedback). Heuristic revision loops often oscillate; the method aims for smoother iteration efficiency.

Key insight: Evolve skills with stable, adaptive updates instead of heuristic rewrites.



MeClear: Cooperative Game-Theoretic Attribution and Risk-Aware Memory Clearance for Long-Horizon LLM Agents

Boyu Yang; Jiazheng Sun; Zilong Lu et al. arXiv: 2609.09115

MeClear clears memories with negative downstream utility via Leave-One-Out screening and sampled cooperative Shapley attribution, then suppresses them from long-horizon execution. Retrieval that only optimizes semantic compatibility can inject outdated or conflicting evidence.

Key insight: Clear memories by negative task utility, not semantic similarity alone.



ExecCritic: Learn to Test, Test to Improve for Coding Agents

Leitian Tao; Baolin Peng; Haorui Wang et al. arXiv: 2609.09133

ExecCritic separates test construction from repair: a Test agent writes repository-native tests, a fail-closed harness freezes them, and a Repair agent patches source under those tests, with role-specific RL. Jointly writing patch and test can agree on wrong behavior and create false confidence.

Key insight: Freeze independent tests before repair so patches cannot grade themselves.



Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails

Zhou Yu; Bin Bi; Shiva Kumar Pentyala et al. arXiv: 2609.09134

Harness-model co-evolution flowchart comparing full-trajectory expert imitation failure to on-policy expert correction that preserves harness fit
Co-evolve harness then correct failing turns on-policy instead of imitating full expert traces

Harness evolution for weaker models helps, but imitating a stronger expert’s full trajectories under that harness regresses 4–30 points across seven enterprise tasks. On-policy expert correction rewrites only the failing turn in the weaker model’s own rollout, preserving native planning style while combining harness and weight gains.

Key insight: Co-evolve harness and weights with on-policy correction, not full imitation.



When and What to Teach: Budget-Aware Online Adaptation for Web Agents

Jianwei Zhang; Sihan Cao; Pengcheng Zheng et al. arXiv: 2609.05513

Score-Guided Online Teaching with Budgeted Trajectory Trimming adapts lightweight local web agents from a stronger teacher without wasting budget on unresolvable episodes or redundant turns. Conventional trajectory-level preference optimization is shown to be inefficient under commercial teacher costs.

Key insight: Spend teacher budget on score-guided, trimmed turns—not full failed traces.



Inference-Time Graph Engineering for Multi-Agent LLM Workflows

Katherine Tieu; Dongqi Fu; Yinglong Xia et al. arXiv: 2609.05774

ReActNet compiles a query and role-specialized agents into a sequence of directed communication graphs at inference time, with each edge carrying a natural-language message instruction. Orchestration becomes task-conditioned temporal graph engineering without training.

Key insight: Synthesize temporal multi-agent graphs and edge semantics at inference time.



SkillSpec: Intent-Masked Specification Reasoning for Agent Skill Correctness

Yizhuo Zhang; Bo Kang; Yi Yang et al. arXiv: 2609.06052

SkillSpec casts skill correctness as Hoare-style specification reasoning over heterogeneous skill artifacts, targeting intent conflicts and silent semantic failures that code tests miss. Correctness is grounded in intended task boundaries and generalizability.

Key insight: Verify skills against specs and intent bounds, not code tests alone.



Explaining AI Agents Through Execution Traces

Vittoria Vineis; Fabiano Veglianti; Lorenzo Antonelli et al. arXiv: 2609.06063

A post-hoc XAI framework turns lengthy agent execution traces into a structured report and a natural-language explanation grounded in those traces. Process-level transparency targets interactive multi-step tool use beyond traditional feature attributions.

Key insight: Explain agents from structured execution traces, not static feature scores.



SAP: State-Guided Data Synthesis with Argument Provenance for Multi-Turn Tool Use

Zichen Tian; Jinpeng Chen; Cheng Gong et al. arXiv: 2609.06124

SAP synthesizes multi-turn tool-use data with state guidance, tool-argument provenance constraints, and turn-level validation so arguments stay grounded across long horizons. Correct tool choice still fails when arguments are fabricated, stale, or weakly grounded.

Key insight: Synthesize tool trajectories with argument provenance, not tool IDs alone.



AgentDrift: A Step-Labeled Benchmark of Injection-Hijacked LLM Agent Trajectories

Asif Pinjari; Mithun Paul Saint-Germain arXiv: 2609.06972

AgentDrift labels 12,536 synthetic tool-call trajectories step by step for where injections enter and which steps they corrupt across five domains. Live attack success and whole-trace guards leave the drift from benign prefix to attacker-serving actions unlabeled.

Key insight: Label injection corruption per step along the tool-call trajectory.



SkillAlign: Aligning Skill Interfaces for LLM-based Agents

Shuo Ren; Xiaomian Kang; Jiajun Zhang arXiv: 2609.07255

SkillAlign represents skills as multi-view procedural cards and renders them through alternative exposure interfaces—full instructions, hints, summaries, workflows, or none. The same skill can help or mislead depending on how it is exposed to the agent.

Key insight: Tune how a skill is exposed; selection alone does not fix interface mismatch.



APPSim-Bench: Bridging Real-world Apps and Reproducible Evaluation for Mobile GUI Agents

Jintian Feng; Long Chen; Xiao Yu et al. arXiv: 2609.07712

AppSim-Bench offers 557 tasks across 17 high-frequency Chinese and English apps in controllable simulated apps that preserve task-relevant logic for deterministic evaluation. It targets the realism–reproducibility trade-off between toy apps and live commercial surfaces.

Key insight: Evaluate mobile GUI agents on simulated apps that stay deterministic.



From Version Conflicts to Decision Conflicts: Selective Revalidation for Long-Running AI Agents

Yongjian Lyu; Yang Ren; Ruofei Lai et al. arXiv: 2609.08015

Long-running agents can propose actions whose justifying state changes before execution. Selective revalidation distinguishes version conflicts that invalidate the decision from harmless metadata changes, going beyond optimistic concurrency that only detects any read-set change.

Key insight: Revalidate pending actions by decision conflict, not every version bump.



Revoked but Still Authoritative: An Empirical Study of Revocation Enforcement in Agent-Memory Systems

Yi Ting Shen; Kentaroh Toyoda; Alex Leung arXiv: 2609.08258

Five agent-memory systems with soft revocation are tested across nine policy scenarios and nine models: none enforce revocation by default when the revoked fact remains visible at retrieval, so agents still act on it under multiple defense conditions.

Key insight: Enforce revocation at retrieval time, not only as a soft invalid mark.



MemForest: Efficient Agent Memory Management via EventTree Partitioning and Progressive Merging

Junxi Wang; Te Sun; Jiayi Zhu et al. arXiv: 2609.08273

Two-panel figure: retrieval latency and storage growth to 200k memory nodes, plus local-plus-global event partitioning on a Boston food timeline
MemForest: scaling costs of raw memory nodes and local+global event partitioning

MemForest partitions history into event-centric units, builds EventTrees as maximum spanning trees, and progressively merges redundant nodes to cut storage and retrieval cost while remaining adaptable to existing agent memory systems.

Key insight: Compress agent memory by event trees and progressive merges, not raw logs.



What Eviction Destroys: A Restore-Counterfactual Audit of Forgetting in Agent Memory

Chen Shen arXiv: 2609.08279

A restore-counterfactual audit reinstates gold evidence after eviction and reruns the reader to classify oracle-answerable errors as recoverable, irreversible, or residual. Budget–accuracy frontiers alone do not separate eviction destruction from retrieval failure.

Key insight: Separate irreversible eviction loss from recoverable retrieval misses.



Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks

Hongbang Yuan; Zhuoran Jin; Yixin Cao arXiv: 2609.08404

Feedback-Enriched Environments shift long-horizon RL from agent-side SFT warm-up to environment-side adaptation that densifies feedback. The pilot study reformulates sparse-reward settings toward richer signals that bootstrap self-evolving agents.

Key insight: Enrich environment feedback to bootstrap long-horizon agent RL.



BIO-MEMART: Biometric-Aware KV Cache Memory for Multi-User LLM Agents

Yanhong Qian; Xuanying He; Qingguo Meng et al. arXiv: 2609.08566

Bio-MemArt attaches biometric templates to KV-cache memory blocks and filters the shared pool with the current user’s probe before MemArt retrieval and reuse. Semantic relevance alone cannot authorize multi-user KV memory access.

Key insight: Authorize shared KV-cache memory with biometrics before semantic reuse.



Who Maintains Agent Skills? A Longitudinal Study of Human-Governed, AI-Assisted Skill Maintenance

Chen Shen; Estevam Hruschka arXiv: 2609.05677

A longitudinal study of five public AI-skill repositories covers 873 commits and 143 skill files to measure human-governed, AI-assisted skill maintenance. Automation papers often treat human maintenance as an unmeasured bottleneck; this work studies that process directly.

Key insight: Measure who actually maintains agent skills over real repository history.



It is Not Yet Another Tool: Creating and Deploying an Agentic AI Companion in a Security Operations Center

Kritan Banstola; Faayed Al Faisal; Duy Dao et al. arXiv: 2609.06250

An agentic SOC companion was designed through year-long fieldwork and used by analysts for four months. In more than 90% of studied usage it supported ticket triage for low-interest events rather than acting as yet another generic tool.

Key insight: Deploy SOC agents as fieldwork-shaped companions, not generic chat tools.



The Unreliable Progress Bar: Can LLM Agents Reliably Report Task Progress Throughout Execution?

Boyang Wang; Yunhan Wang; Yalun Wu arXiv: 2609.08589

Progress-reporting reliability is evaluated on τ²-bench and StageIF across task stages. Almost every deployed model is reliable at some stages and unreliable at others, so frameworks that stop or continue on self-reported progress inherit stage-dependent failure.

Key insight: Treat agent progress bars as stage-dependent, not uniformly trustworthy.



Procedural Graphs: Self-Evolving Execution Structures for LLM Agents

Yuxing Lu; Yicheng Chen; Shanchan Wu et al. arXiv: 2609.09153

Procedural Graphs organize what-to-do knowledge as (procedure, relation, procedure) triplets, analogous to knowledge graphs for facts. Explicit execution structure aims to reduce lost objectives, out-of-order tools, and repeated unproductive actions over long horizons.

Key insight: Store what-to-do structure in procedural graphs, not only accumulating chat.