Monday’s cs.AI listing had 205 new submissions and cross-lists. This page keeps the 37 that are about agent systems: memory and context, harnesses and skills, tool use and computer-use, multi-agent coordination, persistent identity, and local open-weight models aimed at that stack. Telecom-only, medical, climate, quantum, generic eval, and vision-only papers are omitted.
Memory Portability shows the same store can still “forget” after a model upgrade via reinterpretation, mixed embeddings, and failed repair. Execution-state unlearning requires clearing summaries, plaintext memory, pending tool plans, and serving state—not only truncating chat. Persistent skills, CoSkill, TROVE, and Trace2Tower grow reusable procedures from traces and RL, while harness-agnostic reward-hacking immunization warns self-evolve loops game imperfect scores. CUA-Universe scales hybrid GUI+CLI computer-use; Substrate-Aware treats runtime constraints as first-class inputs; CONTINUITY adds security-context contracts for composable controls. ElderBench, Multi-Harness RL, IPI-as-search, Harbor adapters, and agent interchangeability fill out mobile, coding-harness, security, and eval surfaces. Seven papers have a figure extracted from the PDF.
Lin Shi et al. arXiv: 2609.04298
Harbor Adapters offers a unified evaluation infrastructure for agentic benchmarks that otherwise demand bespoke environments and agent integrations, plus Harbor-Index as a curated meta-dataset layer. The goal is to run the same agent stacks across many benches without rewriting glue for every harness.
Key insight: Shared adapters and a meta-dataset beat one-off env glue when evaluating agents at scale.
Duong M. Nguyen et al. arXiv: 2609.04495
Indirect prompt injection is cast as test-time search over a task-dependent attack surface induced by the environment, user task, and injection task. An agentic attacker with a dedicated search harness does environment reconnaissance, structured strategy reasoning, and adaptive evaluation from victim-agent feedback across heterogeneous settings.
Key insight: Treat IPI as searchable env×task surface—not a single static payload.
Chenqian Le et al. arXiv: 2609.04518
Multi-harness RL mixes exposing a policy to several execution harnesses with comparing their rewards inside one relative-advantage group. Isolating the second choice in repository-level coding—from a Qwen3-8B supervised warm start replaying frozen task-harness records from Aider, OpenHands, Qwen Code, and related systems—separates what the recipe actually learns about credit assignment versus portability.
Key insight: Separate multi-harness exposure from in-group reward comparison when judging coding-agent RL.
Siddharth Vohra et al. arXiv: 2609.04579
Grounded LM pipelines split into selecting an object, retrieving passages for it, and answering from that evidence. If the selected object must reach the reader, losing it breaks the handoff—yet benchmark recall often checks a dataset-linked object that can differ from what the pipeline actually selected.
Key insight: Audit selection→retrieval→answer identity handoffs; dataset-linked recall can hide silent drops.
Quan Shi et al. arXiv: 2609.04611
τ^τ-bench (hyper-tau-bench) asks whether an AI system can deliver a production agent under the conditions of a real client engagement—not only whether a finished policy scores on tool loops. As coding agents take on more of the build work, the bench targets end-to-end construction realism.
Key insight: Evaluate agent construction under client-like constraints, not only finished-policy tool scores.
Rongxin Yang et al. arXiv: 2609.04665
Self-evolving language models improve by proposing updates and keeping whatever raises a visible score. When that score is an imperfect proxy, sustained selection widens the gap—reward hacking. The paper studies harness-agnostic detection and immunization so online evolve loops do not quietly optimize the proxy.
Key insight: Immunize self-evolve loops against imperfect-score hacking before enabling online writeback.
Weide Zhan et al. arXiv: 2609.04850
Existing GUI benchmarks lean on explicit goal-oriented instructions and miss how older adults actually speak—indirect speech, referential ambiguity, and under-specified requests. ElderBench benchmarks autonomous mobile agents under those naturalistic patterns so success tracks real assistance, not tidy scripted goals.
Key insight: Mobile agent benches need implicit, under-specified elderly language—not only explicit GUI goals.
Jinyuan Feng et al. arXiv: 2609.04865
Skill libraries help agentic RL reuse procedural knowledge, but common paradigms either decouple skill evolution from policy optimization or freeze meta-skills as fixed workflows. CoSkill jointly RL-trains a reasoning agent and a meta-skill agent so hierarchical skills evolve instead of staying passive objects.
Key insight: Evolve hierarchical skills with a joint reasoning + meta-skill RL loop, not passive skill shelves.
Longtao Hu; Xiao Liang; Linchao Zhu arXiv: 2609.04869
Computer-use agents often throw away procedural knowledge after a GUI rollout. This work turns interaction traces into persistent skills via online evolution, measuring incremental value over the same agent without skills and refining reusable procedures across later tasks.
Key insight: Mine transient computer-use trajectories into persistent skills online—not one-shot libraries.
Chao Yao et al. arXiv: 2609.04875
Long-running agents accrete summaries, plaintext memory, pending tool plans, and KV cache beyond the chat transcript. Today's forget ops often delete a memory record and stop. Execution-state unlearning requires the agent to behave as if revoked information never existed—without a full serving restart.
Key insight: Forget must clear summaries, memory, tool plans, and serving state—not only the transcript.
Jiahe Geng; Jinpeng Wang; Kun Yuan arXiv: 2609.04915
Under tight prompt budgets, the question is which memory design wins on the quality–token Pareto frontier. RSM-full uses online max-member clustering and atom-aware packing so long-horizon agents stay useful without full-context prompting.
Key insight: Optimize quality–token trade-offs with online clustering and atom-aware packing, not raw recall alone.
Tianxing Wang et al. arXiv: 2609.05019
Agents often lock an execution structure before runtime evidence arrives, then either run stale steps or replan broadly when intermediate outcomes invalidate the plan. TROVE defers route commitment until trace-grounded validation and editing can revise the pending continuation.
Key insight: Validate and edit skill routes from traces before committing the next orchestration step.
Manu Agrawal arXiv: 2609.05232
Substrate blindness is planning without memory, runtime, compute, and operational constraints in the agent's state. Through numerical code generation and related settings, the paper argues execution context should be a first-class input when choosing suitable plans.
Key insight: Feed memory/runtime/compute constraints into planning—do not treat the substrate as invisible.
Chris Zheng; Geng Yang arXiv: 2609.05269
Individually correct provenance, authorization, policy, adapter, and execution controls can still fail end-to-end when security-critical context is dropped, widened, rebound, or reinterpreted across boundaries. CONTINUITY proposes security-context contracts so composable agent controls keep that context intact.
Key insight: Compose agent security with explicit security-context contracts across component boundaries.
Ankit Goyal; Jaideep Ray arXiv: 2609.05339
Keeping the same memory store does not guarantee the upgraded model behaves the same: notes can be reinterpreted, mixed embeddings can break retrieval, and repair may need original evidence. The controlled study compares LC-RAW, RAG, notes, and fixed-schema KG stores under same versus mixed embedders and migration directions.
Key insight: Same store ≠ portable memory—test reinterpretation, embedding mix, and repair before model swaps.
Haoting Shi et al. arXiv: 2609.05374
Real computer work mixes visual-state inspection with high-throughput CLI, but many computer-use agents still act mostly through the GUI. CUA-Universe supplies scalable hybrid GUI+CLI environments so agents can coordinate both modalities over shared application state, with large gains versus GUI-only baselines on reported benches.
Key insight: Prefer hybrid GUI+CLI computer-use environments; CLI when available beats GUI-only trajectories.
John Seon Keun Yi; Joshua R. Minot; Dokyun Lee arXiv: 2609.04442
Standard RAG retrieves isolated passages without tracking cross-document evidence or quantifying uncertainty. GRACE deconstructs claims into a graph-grounded reflective copilot loop so expert-in-the-loop knowledge expansion stays tied to evidence relationships.
Key insight: Ground high-stakes copilots in cross-document evidence graphs, not isolated RAG hits.
Ismail Erbas; Xavier Intes; Vikas Pandey arXiv: 2609.04490
In recurrent nets the quantized state is stored and returned next step, so the write-back rule can alter subsequent computation. Isolating recurrent-state write-back in a compact GRU encoder–decoder shows how low-precision temporal inference can corrupt memory-like state.
Key insight: Quantized recurrent write-back can silently corrupt temporal state—choose the store rule deliberately.
Milos Gravara; Andrija Stanisic; Stefan Nastic arXiv: 2609.04513
Compound AI workflows expose many model variants and placements per stage. Atlas searches execution plans that select models and place them on heterogeneous clusters so deployment cost and performance trade-offs are explicit.
Key insight: Treat compound-workflow deployment as joint model-selection and placement search.
Gnaneswar Villuri; Hashmath Shaik; Alex Doboli arXiv: 2609.04570
When one routine's meaning depends on another's runtime behavior, static context binding fails. The paper studies dynamic LLM context adaptation for code generation under coupled semantics that textual descriptions alone cannot resolve.
Key insight: Adapt context dynamically when routine correctness depends on joint runtime behavior.
Xing Chen; Hengshuai Yao arXiv: 2609.04575
Fine-grained MoE renormalization calibrates expert gain to the training top-k, so cutting k at inference changes both which experts fire and branch strength. Separating those effects enables training-free halving of activated experts with controlled quality impact.
Key insight: Halve activated MoE experts at inference only after accounting for renormalization gain.
Chenyu Zhou et al. arXiv: 2609.04629
A runtime gate in a ReAct loop is a search operator over proposals, not merely a filter. SiLR studies post-violation recovery admission with structure-preserving criteria and process reward so progress can continue while the system is still in violation.
Key insight: Design tool-agent gates as structure-preserving admission, not reject-and-retry filters.
Cheng Li et al. arXiv: 2609.04678
Post-training often mismatches production tokens and controls when simplified envs or offline log reconstruction distort prompts. A fidelity-aware coupling keeps trainer-side sampling aligned with the tokens and controls the deployed coding agent actually sees.
Key insight: Post-train coding agents on the exact tokens and controls you deploy.
Xinyu Li et al. arXiv: 2609.04715
Per-user fine-tuning personalizes well but does not scale. PLUME uses low-rank user modulation in shared subspaces so personalization quality stays high without per-user full adapters.
Key insight: Personalize with shared-subspace low-rank user modulation instead of full per-user finetunes.
Zehao Wang et al. arXiv: 2609.04749
Multi-agent failures require tracing natural-language interactions to the decisive earliest error. DCFA uses dual-view causal-inspired attribution to reason about which agent action caused system-level failure.
Key insight: Attribute multi-agent failures to the earliest decisive error with dual-view causal tracing.
Hyun Bin Park et al. arXiv: 2609.04773
On-policy distillation matches student next-token distributions to a teacher, but as rollouts enter states the teacher would not visit the gap accumulates. Persistent teacher anchoring stabilizes tool-agent distillation when teacher quality varies across the trajectory.
Key insight: Anchor tool-agent distillation to the teacher persistently as rollouts leave teacher states.
Xinyu Mao et al. arXiv: 2609.04801
Visual personalization can retrieve a true record yet apply it to the wrong subject. Record authorization requires subject presence, record-edge validity, and answer support; violations are visual memory misbinding.
Key insight: Authorize personalized records with presence×edge×support—retrieval alone is not enough.
Aziz Ben Amor et al. arXiv: 2609.04898
Repository-scale refactoring needs agents to propagate one change across interdependent files without altering behavior. RefactorPlatform holds the environment fixed and varies design axes—model backbone, prompts, tools—so success factors are isolable.
Key insight: Evaluate repo-scale refactoring agents with a harness that varies one design axis at a time.
Zukang Xu et al. arXiv: 2609.05228
Fixed top-k MoE routing wastes computation on low-contribution experts. ACE skips experts adaptively without calibration data or extra training by estimating actual routed-expert contribution at inference.
Key insight: Skip MoE experts adaptively without calibration when contribution estimates are reliable.
10a Labs et al. arXiv: 2609.05241
An ecosystem removes safety guardrails from open-weight models and redistributes them at scale. Profiling producers, reproductions, and applications from 2024–2026 shows redistribution itself acting as the persistence layer for uncensoring—not a single model drop.
Key insight: Treat redistribution networks as the persistence layer for uncensoring open-weight models.
Jiazheng Sun et al. arXiv: 2609.05261
Flat trajectory retrieval and shallow skill summarization ignore temporal dependencies and outcome-conditioned topology. Trace2Tower distills raw trajectories into multi-level skills with transition-aware EigenTrace induction.
Key insight: Induce multi-level skills from transition-aware traces, not flat trajectory summaries.
Jianxin Gao et al. arXiv: 2609.05279
Production multi-agent systems often assume role-matched agents are interchangeable. Forming teams independently, then trading role-matched agents while each keeps a private notebook, tests what actually breaks under swap.
Key insight: Measure role-swap cost empirically before treating agents as interchangeable.
Dain Kim et al. arXiv: 2609.05395
KOPA-Bench covers 145 real-world multi-step tasks over live Korean open public APIs under on-premise open-source constraints, plus a data-synthesis recipe for closing the gap where open models underperform.
Key insight: Benchmark and synthesize multi-step tool-calling on live public APIs for on-prem open models.
Happy Bhati arXiv: 2609.04681
As coding agents inspect repos, edit files, run tools, and open PRs for long stretches, the bottleneck shifts from typing code to reliability, verification, and cost. The survey frames where field productivity gains attenuate once agents own more of the SDLC.
Key insight: Agentic SDLC gains hinge on verification and cost controls, not autocomplete speed alone.
Chenqi Li et al. arXiv: 2609.04778
Diffusion language models refine tokens by iterative denoising rather than left-to-right decoding, enabling parallel updates and bidirectional context. This survey maps foundations, applications, and challenges for mobile-edge agentic AI.
Key insight: Watch diffusion LMs as a non-autoregressive option for mobile-edge agents as tooling matures.
Tianyidan Xie et al. arXiv: 2609.04802
Long-horizon embodied agents need object state transitions queryable in natural language across hours to days. Linguistic trajectory encoding compresses motion into language-addressable memory instead of dropping it in clip embeddings or keeping only raw coordinates.
Key insight: Encode long-horizon spatial change as linguistic trajectories for queryable persistent memory.
Linsen Zhu; Mengqing Cai arXiv: 2609.04894
Models become consequential when surrounding systems let outputs change external state—tools, interfaces, delegation, persistence, generated worlds, robots. This survey separates model competence, system integration, persistence, and safe authority across digital, social, virtual, and physical settings.
Key insight: Separate model skill from integration, persistence, and authority when mapping agentic progress.