Wednesday’s cs.AI listing had 517 new submissions and cross-lists. This page keeps the 45 that are about agent systems: memory and context, harnesses and skills, tool use and computer-use, multi-agent coordination, persistent identity, and local open-weight models aimed at that stack. Telecom-only, medical, climate, quantum, generic eval, and vision-only papers are omitted.
MERIT cost-accounts memory for tool-using agents (0.00→0.55–1.00 dependent-task success; up to 60-point swings by memory implementation). EdgeMem preserves source turns in a multi-anchor hypergraph (LoCoMo 61.01 vs 58.70) without generative memory writes. DroidTool self-generates Android tool actions beside GUI steps. ResidualAuth shows revocation needs residual authorization state beyond current permissions. Beyond Prompts treats harness search as resource-bounded selection with RelLift95(B). FrogNano trains a 4B SWE agent with frontier-calibrated synthetic RL only. MemForest compresses memory via EventTrees; Scaffold grows recursive parametric web skills; co-evolving harnesses warns that full expert imitation under an evolved harness can regress 4–30 points unless correction stays on-policy. Eight papers have a figure extracted from the HTML/PDF.
Shweta Mishra; Shashank Mishra arXiv: 2609.05441
MERIT measures whether long-term memory changes what tool-using agents do, with cost metering, rather than conversational recall alone. Across 23,440 episodes it lifts dependent-task success from a leak-verified floor of 0.00 to 0.55–1.00; on updated facts, embedding retrieval is unstable (0.30–0.95) while update-on-write stores stay at 0.70–1.00, and swapping a memory implementation moves success by up to 60 points.
Key insight: Measure memory by task utility and dollar cost, not dialogue recall alone.
Adib Hasan; Daniel Schaffield; Akashnil Dutta et al. arXiv: 2609.05446
AutoFyn adapts a frozen model across rounds by updating persistent state from verified rewards instead of weights. Each round starts from a fresh session; durable information returns only through explicit interfaces such as memory files, reports, and repository state, while an orchestrator explores alternatives and a task-grounded verifier scores progress.
Key insight: Adapt frozen agents via verified persistent state, not weight updates.
Moonwon Choi; Seokho Jeong; Seunggeun Lee arXiv: 2609.08062
ResidualAuth shows two authorization histories can share identical current permissions and reachability yet require opposite decisions after the same revocation. It formalizes residual authorization state, proves exponentially many future-distinct states can share one transitive closure, and compiles the constructions into paired language-agent episodes across open models.
Key insight: Preserve residual authorization history, not only the current permission set.
Yu Liu; Zhilin Liu; Zhiwei Yang et al. arXiv: 2609.06059
DAREBench evaluates models as agents on multimodal perception, multi-step execution, tool use, and artifact delivery inside a shared OpenClaw environment. It organizes 233 tasks to capture workload variation and support deployment-aware comparison beyond static answer correctness.
Key insight: Judge agents on deployment workloads and artifacts, not static QA alone.
Ruoxi Shang; Christina-Maria Androna; Orfeas Menis Mastromichalakis et al. arXiv: 2609.06783
AURA-Eval diagnoses risk recognition, pre-action detection, and safe completion when a safe path exists, instead of collapsing safety to one score. From 157 trajectories it builds 1,249 items and evaluates 20 frontier and open-weight models with rubrics over tool-use trajectories.
Key insight: Score agent safety at decision points, not as a single aggregate.
Juyong Lee; Woogyeol Jin; Kimin Lee arXiv: 2609.06792
DroidTool lets Android GUI agents self-generate tool actions as Python functions over app state, using proposal, implementation, test generation, and repair with relational tests across tools. The hybrid GUI-plus-API action space targets proficiency and efficiency without hand-building every tool.
Key insight: Let mobile agents mint and verify their own API tools beside GUI actions.
Xiaoting Lyu; Yuhong Wu; Yufei Han et al. arXiv: 2609.07131
AgentLeak asks whether a weaker black-box attacker can clone a stronger proprietary agent’s capabilities beyond stealing explicit skill artifacts. Artifact leakage alone may not transfer capability when the attacker lacks implicit procedural behaviors acquired through execution.
Key insight: Capability theft can require procedural behavior, not only skill files.
Bowei He; Xiaokun Zhang; Meng Ding et al. arXiv: 2609.05511
Scaffold induces parametric executable skills from successful web-agent trajectories under a multi-instance abstraction constraint and maintains a recursively composed hierarchy. Unlike flat prompt-side skill caches, it compresses redundancy and composes skills for visually rich, long-horizon sites.
Key insight: Grow recursive parametric skills from web trajectories, not flat caches.
Hamed Khosravi; Xiaoming Huo arXiv: 2609.05527
Human–agent teams are worth keeping only if they beat human-alone and agent-alone alternatives, but those counterfactuals are costly to replay. Under a fixed replay budget the design question is which tasks get human-only versus agent-only replays; prior methods do not target that decision directly.
Key insight: Allocate scarce replays to the human-vs-agent decision that matters.
Zeyang Cui; Jiannong Cao; Zhiyuan Wen et al. arXiv: 2609.05553
EdgeMem keeps original interaction turns and organizes them with content, temporal, and episodic anchors in a multi-anchor hypergraph built by lightweight local processing. Retrieval returns source evidence and reserves the LLM for final answers; on LoCoMo it leads seven systems under a shared prompt (61.01 vs 58.70) with no generative-LLM calls in construction or retrieval.
Key insight: Preserve source evidence in a hypergraph; generate only at answer time.
Yirong Zeng; Shen You; Jinhang Feng et al. arXiv: 2609.05576
EnvCraft synthesizes executable environments and scalable training data for claw-like agents that act across stateful workspaces, going beyond tool-calling endpoints. The goal is to unblock agentic RL where interactive environments are scarce.
Key insight: Synthesize full executable workspaces for agentic RL, not tool stubs alone.
Hoyeol Yang; Woojung Song; Taewon Kim et al. arXiv: 2609.05587
Fourteen LLMs are tested with corrupted returns from web search, sub-agent delegation, and code execution. Mean adoption of corrupted content exceeds one third for every tool and reaches 68% in the worst setting, showing overtrust when tools look plausible but wrong.
Key insight: Measure tool overtrust under corrupted returns, not only task success.
Cen (Mia) Zhao; Haibo Ruan; Wenjie Chen et al. arXiv: 2609.05736
The work treats harness selection—prompts and tool-boundary middleware around a fixed model—as a resource-bounded search. An optimizer-agnostic protocol reports mean and worst-condition lift, repeatability, cost diagnostics, and RelLift95(B), the conservative held-out gain of the harness chosen under budget B.
Key insight: Optimize the harness under budget; report conservative held-out lift.
Wang Wei; Tiankai Yang; Samyadeep Basu et al. arXiv: 2609.05824
Diverse Skill Routing reranks large skill registries with a Determinantal Point Process that balances relevance and non-redundancy via a query-residual diversity kernel. The aim is to stop wasting context on redundant skills when complex tasks need complementary sets.
Key insight: Route skills for diversity of coverage, not only top-k relevance.
Zhiyi Lyu; Yewen Li; Longtao Zheng et al. arXiv: 2609.05837
AgentBrew learns tool-use policies offline from one batch of raw trajectories without task verifiers or iterative on-policy rollouts. After unfiltered exploration, it extracts training signal from noisy corpora for environment-specific agents where simulators and budgets are limited.
Key insight: Train tool agents offline from raw trajectories when verifiers are absent.
Zhongan Bi; Qiwen Wang; Jianrong Jiang et al. arXiv: 2609.06027
HAE-GEO tracks deep-search agents from exposure through verification, revision, and recovery under hierarchical web evidence poisoning (direct assertion, camouflage, and apparent corroboration). It goes beyond measuring whether poisoned content is merely retrieved or endorsed.
Key insight: Score poisoning resistance by recovery along the full search trajectory.
Abhijit Chakraborty; Ni Trieu; Vivek Gupta arXiv: 2609.06815
Typed federated artifacts share schema-validated tool-routing knowledge across frozen, heterogeneous agents without transferring weights. SYNAPSE1 is instantiated as shared routing knowledge after cleaning garbage and training items, aiming at cross-model transfer with per-field privacy and dispute handling.
Key insight: Share typed tool-routing artifacts across frozen multi-vendor agents.
Sasank Annapureddy; Anjaneya Prasad Thamatani arXiv: 2609.07910
PRIMUS couples prime-power agent identity with BLS aggregate signatures, derives a safe-kill threshold that cuts false-positive termination from 80% to 0.00% under 10% channel noise, and asks whether the same verification machinery can steer generate-and-test toward better answers under adversarial federation conditions.
Key insight: Govern multi-agent federations with signed identity and safe-kill thresholds.
Minseon Kim; Zhengyan Shi; Emiliano Penaloza et al. arXiv: 2609.07925
FrogNano is a 4B coding agent post-trained only with RL on about 1,500 SWE environments using synthetic tasks at the learnability frontier of the current checkpoint. Results support competitive small coding agents without distillation from larger models, runnable on minimal hardware.
Key insight: Train small coding agents with frontier-calibrated synthetic RL tasks alone.
Dawei Fu; Cheng Jiang; Sitian Qian et al. arXiv: 2609.08228
SE-GoS evolves an existing Graph-of-Skills from execution traces without training while preserving the graph for scalable retrieval. It asks whether historical traces can be distilled into a better retrieval graph that generalizes to unseen tasks.
Key insight: Evolve skill-retrieval graphs from traces without retraining the model.
Yanhong Qian; Qingguo Meng; Shihao Ding et al. arXiv: 2609.08558
Bio-Memory augments A-Mem notes with biometric embeddings and filters the retrieval pool by biometric match before semantic ranking. On LoCoMo in a 10-user shared setting it targets identity-aware personalization so the requester matches the memory owner.
Key insight: Gate personalized memory retrieval on biometric match, then semantics.
Dac Duy Anh Nguyen; Zhangchi Qiu; Shigeng Chen et al. arXiv: 2609.08599
This survey frames graph-based personalized memory for long-term LLM agents: representation, evolution, retrieval, and evaluation of preferences, goals, constraints, and experiences that change over time. Explicit relations, temporal context, and evidence links structure what is remembered and how it is revised.
Key insight: Treat personalized agent memory as an evolving evidence-linked graph.
Evelyn Duesterwald; Benjamin Elder; Lilian Ngweta et al. arXiv: 2609.08832
A ReAct agent on AppWorld with GPT-4.1 averages 77% per-run pass rate but succeeds in all five repeats only 53% of the time—a 24-point consistency gap. A self-evolving framework converts unstable low-consistency steps into episodic memory for future runs.
Key insight: Close the multi-run consistency gap with memory of unstable steps.
Gaoyuan Li; Meihao Fan; Yizhe Liu et al. arXiv: 2609.08944
SkillAdam targets stable, efficient skill self-evolution for frozen agents by addressing direction stability (accumulate corrections) and update adaptivity (step size from feedback). Heuristic revision loops often oscillate; the method aims for smoother iteration efficiency.
Key insight: Evolve skills with stable, adaptive updates instead of heuristic rewrites.
Boyu Yang; Jiazheng Sun; Zilong Lu et al. arXiv: 2609.09115
MeClear clears memories with negative downstream utility via Leave-One-Out screening and sampled cooperative Shapley attribution, then suppresses them from long-horizon execution. Retrieval that only optimizes semantic compatibility can inject outdated or conflicting evidence.
Key insight: Clear memories by negative task utility, not semantic similarity alone.
Leitian Tao; Baolin Peng; Haorui Wang et al. arXiv: 2609.09133
ExecCritic separates test construction from repair: a Test agent writes repository-native tests, a fail-closed harness freezes them, and a Repair agent patches source under those tests, with role-specific RL. Jointly writing patch and test can agree on wrong behavior and create false confidence.
Key insight: Freeze independent tests before repair so patches cannot grade themselves.
Zhou Yu; Bin Bi; Shiva Kumar Pentyala et al. arXiv: 2609.09134
Harness evolution for weaker models helps, but imitating a stronger expert’s full trajectories under that harness regresses 4–30 points across seven enterprise tasks. On-policy expert correction rewrites only the failing turn in the weaker model’s own rollout, preserving native planning style while combining harness and weight gains.
Key insight: Co-evolve harness and weights with on-policy correction, not full imitation.
Jianwei Zhang; Sihan Cao; Pengcheng Zheng et al. arXiv: 2609.05513
Score-Guided Online Teaching with Budgeted Trajectory Trimming adapts lightweight local web agents from a stronger teacher without wasting budget on unresolvable episodes or redundant turns. Conventional trajectory-level preference optimization is shown to be inefficient under commercial teacher costs.
Key insight: Spend teacher budget on score-guided, trimmed turns—not full failed traces.
Katherine Tieu; Dongqi Fu; Yinglong Xia et al. arXiv: 2609.05774
ReActNet compiles a query and role-specialized agents into a sequence of directed communication graphs at inference time, with each edge carrying a natural-language message instruction. Orchestration becomes task-conditioned temporal graph engineering without training.
Key insight: Synthesize temporal multi-agent graphs and edge semantics at inference time.
Yizhuo Zhang; Bo Kang; Yi Yang et al. arXiv: 2609.06052
SkillSpec casts skill correctness as Hoare-style specification reasoning over heterogeneous skill artifacts, targeting intent conflicts and silent semantic failures that code tests miss. Correctness is grounded in intended task boundaries and generalizability.
Key insight: Verify skills against specs and intent bounds, not code tests alone.
Vittoria Vineis; Fabiano Veglianti; Lorenzo Antonelli et al. arXiv: 2609.06063
A post-hoc XAI framework turns lengthy agent execution traces into a structured report and a natural-language explanation grounded in those traces. Process-level transparency targets interactive multi-step tool use beyond traditional feature attributions.
Key insight: Explain agents from structured execution traces, not static feature scores.
Zichen Tian; Jinpeng Chen; Cheng Gong et al. arXiv: 2609.06124
SAP synthesizes multi-turn tool-use data with state guidance, tool-argument provenance constraints, and turn-level validation so arguments stay grounded across long horizons. Correct tool choice still fails when arguments are fabricated, stale, or weakly grounded.
Key insight: Synthesize tool trajectories with argument provenance, not tool IDs alone.
Asif Pinjari; Mithun Paul Saint-Germain arXiv: 2609.06972
AgentDrift labels 12,536 synthetic tool-call trajectories step by step for where injections enter and which steps they corrupt across five domains. Live attack success and whole-trace guards leave the drift from benign prefix to attacker-serving actions unlabeled.
Key insight: Label injection corruption per step along the tool-call trajectory.
Shuo Ren; Xiaomian Kang; Jiajun Zhang arXiv: 2609.07255
SkillAlign represents skills as multi-view procedural cards and renders them through alternative exposure interfaces—full instructions, hints, summaries, workflows, or none. The same skill can help or mislead depending on how it is exposed to the agent.
Key insight: Tune how a skill is exposed; selection alone does not fix interface mismatch.
Jintian Feng; Long Chen; Xiao Yu et al. arXiv: 2609.07712
AppSim-Bench offers 557 tasks across 17 high-frequency Chinese and English apps in controllable simulated apps that preserve task-relevant logic for deterministic evaluation. It targets the realism–reproducibility trade-off between toy apps and live commercial surfaces.
Key insight: Evaluate mobile GUI agents on simulated apps that stay deterministic.
Yongjian Lyu; Yang Ren; Ruofei Lai et al. arXiv: 2609.08015
Long-running agents can propose actions whose justifying state changes before execution. Selective revalidation distinguishes version conflicts that invalidate the decision from harmless metadata changes, going beyond optimistic concurrency that only detects any read-set change.
Key insight: Revalidate pending actions by decision conflict, not every version bump.
Yi Ting Shen; Kentaroh Toyoda; Alex Leung arXiv: 2609.08258
Five agent-memory systems with soft revocation are tested across nine policy scenarios and nine models: none enforce revocation by default when the revoked fact remains visible at retrieval, so agents still act on it under multiple defense conditions.
Key insight: Enforce revocation at retrieval time, not only as a soft invalid mark.
Junxi Wang; Te Sun; Jiayi Zhu et al. arXiv: 2609.08273
MemForest partitions history into event-centric units, builds EventTrees as maximum spanning trees, and progressively merges redundant nodes to cut storage and retrieval cost while remaining adaptable to existing agent memory systems.
Key insight: Compress agent memory by event trees and progressive merges, not raw logs.
Chen Shen arXiv: 2609.08279
A restore-counterfactual audit reinstates gold evidence after eviction and reruns the reader to classify oracle-answerable errors as recoverable, irreversible, or residual. Budget–accuracy frontiers alone do not separate eviction destruction from retrieval failure.
Key insight: Separate irreversible eviction loss from recoverable retrieval misses.
Hongbang Yuan; Zhuoran Jin; Yixin Cao arXiv: 2609.08404
Feedback-Enriched Environments shift long-horizon RL from agent-side SFT warm-up to environment-side adaptation that densifies feedback. The pilot study reformulates sparse-reward settings toward richer signals that bootstrap self-evolving agents.
Key insight: Enrich environment feedback to bootstrap long-horizon agent RL.
Yanhong Qian; Xuanying He; Qingguo Meng et al. arXiv: 2609.08566
Bio-MemArt attaches biometric templates to KV-cache memory blocks and filters the shared pool with the current user’s probe before MemArt retrieval and reuse. Semantic relevance alone cannot authorize multi-user KV memory access.
Key insight: Authorize shared KV-cache memory with biometrics before semantic reuse.
Chen Shen; Estevam Hruschka arXiv: 2609.05677
A longitudinal study of five public AI-skill repositories covers 873 commits and 143 skill files to measure human-governed, AI-assisted skill maintenance. Automation papers often treat human maintenance as an unmeasured bottleneck; this work studies that process directly.
Key insight: Measure who actually maintains agent skills over real repository history.
Kritan Banstola; Faayed Al Faisal; Duy Dao et al. arXiv: 2609.06250
An agentic SOC companion was designed through year-long fieldwork and used by analysts for four months. In more than 90% of studied usage it supported ticket triage for low-interest events rather than acting as yet another generic tool.
Key insight: Deploy SOC agents as fieldwork-shaped companions, not generic chat tools.
Boyang Wang; Yunhan Wang; Yalun Wu arXiv: 2609.08589
Progress-reporting reliability is evaluated on τ²-bench and StageIF across task stages. Almost every deployed model is reliable at some stages and unreliable at others, so frameworks that stop or continue on self-reported progress inherit stage-dependent failure.
Key insight: Treat agent progress bars as stage-dependent, not uniformly trustworthy.
Yuxing Lu; Yicheng Chen; Shanchan Wu et al. arXiv: 2609.09153
Procedural Graphs organize what-to-do knowledge as (procedure, relation, procedure) triplets, analogous to knowledge graphs for facts. Explicit execution structure aims to reduce lost objectives, out-of-order tools, and repeated unproductive actions over long horizons.
Key insight: Store what-to-do structure in procedural graphs, not only accumulating chat.