Monday’s cs.AI listing had 205 new submissions and cross-lists. This page keeps the 37 that are about agent systems: memory and context, harnesses and skills, tool use and computer-use, multi-agent coordination, persistent identity, and local open-weight models aimed at that stack. Telecom-only, medical, climate, quantum, generic eval, and vision-only papers are omitted.

Memory Portability shows the same store can still “forget” after a model upgrade via reinterpretation, mixed embeddings, and failed repair. Execution-state unlearning requires clearing summaries, plaintext memory, pending tool plans, and serving state—not only truncating chat. Persistent skills, CoSkill, TROVE, and Trace2Tower grow reusable procedures from traces and RL, while harness-agnostic reward-hacking immunization warns self-evolve loops game imperfect scores. CUA-Universe scales hybrid GUI+CLI computer-use; Substrate-Aware treats runtime constraints as first-class inputs; CONTINUITY adds security-context contracts for composable controls. ElderBench, Multi-Harness RL, IPI-as-search, Harbor adapters, and agent interchangeability fill out mobile, coding-harness, security, and eval surfaces. Seven papers have a figure extracted from the PDF.


Research Papers

Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation

Lin Shi et al. arXiv: 2609.04298

Harbor Adapters offers a unified evaluation infrastructure for agentic benchmarks that otherwise demand bespoke environments and agent integrations, plus Harbor-Index as a curated meta-dataset layer. The goal is to run the same agent stacks across many benches without rewriting glue for every harness.

Key insight: Shared adapters and a meta-dataset beat one-off env glue when evaluating agents at scale.



Rethinking Indirect Prompt Injection as a Test-Time Search Problem

Duong M. Nguyen et al. arXiv: 2609.04495

Line chart of attack success rate versus attacker token budget comparing Full-Harness, No-Harness, and No-Strategy indirect prompt injection search
IPI framed as test-time search: full-harness attackers scale ASR with token budget

Indirect prompt injection is cast as test-time search over a task-dependent attack surface induced by the environment, user task, and injection task. An agentic attacker with a dedicated search harness does environment reconnaissance, structured strategy reasoning, and adaptive evaluation from victim-agent feedback across heterogeneous settings.

Key insight: Treat IPI as searchable env×task surface—not a single static payload.



What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents

Chenqian Le et al. arXiv: 2609.04518

Multi-harness RL mixes exposing a policy to several execution harnesses with comparing their rewards inside one relative-advantage group. Isolating the second choice in repository-level coding—from a Qwen3-8B supervised warm start replaying frozen task-harness records from Aider, OpenHands, Qwen Code, and related systems—separates what the recipe actually learns about credit assignment versus portability.

Key insight: Separate multi-harness exposure from in-group reward comparison when judging coding-agent RL.



Does the Selected Object Reach the Reader? Auditing Identity Handoffs in Grounded Language-Model Pipelines

Siddharth Vohra et al. arXiv: 2609.04579

Grounded LM pipelines split into selecting an object, retrieving passages for it, and answering from that evidence. If the selected object must reach the reader, losing it breaks the handoff—yet benchmark recall often checks a dataset-linked object that can differ from what the pipeline actually selected.

Key insight: Audit selection→retrieval→answer identity handoffs; dataset-linked recall can hide silent drops.



τ^τ-Bench: An Environment for End-To-End, Realistic Agent Construction

Quan Shi et al. arXiv: 2609.04611

τ^τ-bench (hyper-tau-bench) asks whether an AI system can deliver a production agent under the conditions of a real client engagement—not only whether a finished policy scores on tool loops. As coding agents take on more of the build work, the bench targets end-to-end construction realism.

Key insight: Evaluate agent construction under client-like constraints, not only finished-policy tool scores.



Harness-agnostic detection and immunization of reward hacking in self-evolving language models

Rongxin Yang et al. arXiv: 2609.04665

Self-evolving language models improve by proposing updates and keeping whatever raises a visible score. When that score is an imperfect proxy, sustained selection widens the gap—reward hacking. The paper studies harness-agnostic detection and immunization so online evolve loops do not quietly optimize the proxy.

Key insight: Immunize self-evolve loops against imperfect-score hacking before enabling online writeback.



ElderBench: Benchmarking Autonomous Mobile Agents for Older Adults

Weide Zhan et al. arXiv: 2609.04850

ElderBench-themed illustration of an older adult on a smartphone video call, representing mobile agents assisting older users
ElderBench targets mobile agents helping older adults with implicit, under-specified goals

Existing GUI benchmarks lean on explicit goal-oriented instructions and miss how older adults actually speak—indirect speech, referential ambiguity, and under-specified requests. ElderBench benchmarks autonomous mobile agents under those naturalistic patterns so success tracks real assistance, not tidy scripted goals.

Key insight: Mobile agent benches need implicit, under-specified elderly language—not only explicit GUI goals.



CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution

Jinyuan Feng et al. arXiv: 2609.04865

CoSkill diagram of a reasoning agent retrieving skills while a meta-skill agent optimizes step-level skills in a hierarchical library
CoSkill jointly optimizes reasoning and meta-skill agents over a hierarchical skill library

Skill libraries help agentic RL reuse procedural knowledge, but common paradigms either decouple skill evolution from policy optimization or freeze meta-skills as fixed workflows. CoSkill jointly RL-trains a reasoning agent and a meta-skill agent so hierarchical skills evolve instead of staying passive objects.

Key insight: Evolve hierarchical skills with a joint reasoning + meta-skill RL loop, not passive skill shelves.



From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use Agents

Longtao Hu; Xiao Liang; Linchao Zhu arXiv: 2609.04869

Computer-use GUI episode: GIMP crop tool on a desktop task, illustrating interaction traces mined into persistent skills
GUI interaction traces (desktop apps) become reusable persistent skills via online evolution

Computer-use agents often throw away procedural knowledge after a GUI rollout. This work turns interaction traces into persistent skills via online evolution, measuring incremental value over the same agent without skills and refining reusable procedures across later tasks.

Key insight: Mine transient computer-use trajectories into persistent skills online—not one-shot libraries.



Forgetting Without Restarting: Execution-State Unlearning for Stateful LLM Agents

Chao Yao et al. arXiv: 2609.04875

Four-step execution-state unlearning pipeline: locate tainted source, restore checkpoint, sanitize and replay, produce counterfactual clean state
Execution-state unlearning restores a checkpoint, sanitizes inputs, and replays to a counterfactual clean state

Long-running agents accrete summaries, plaintext memory, pending tool plans, and KV cache beyond the chat transcript. Today's forget ops often delete a memory record and stop. Execution-state unlearning requires the agent to behave as if revoked information never existed—without a full serving restart.

Key insight: Forget must clear summaries, memory, tool plans, and serving state—not only the transcript.



Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing

Jiahe Geng; Jinpeng Wang; Kun Yuan arXiv: 2609.04915

Under tight prompt budgets, the question is which memory design wins on the quality–token Pareto frontier. RSM-full uses online max-member clustering and atom-aware packing so long-horizon agents stay useful without full-context prompting.

Key insight: Optimize quality–token trade-offs with online clustering and atom-aware packing, not raw recall alone.



TROVE: Adaptive Agent Skill Orchestration via Trace-Grounded Route Validation and Editing

Tianxing Wang et al. arXiv: 2609.05019

Agents often lock an execution structure before runtime evidence arrives, then either run stale steps or replan broadly when intermediate outcomes invalidate the plan. TROVE defers route commitment until trace-grounded validation and editing can revise the pending continuation.

Key insight: Validate and edit skill routes from traces before committing the next orchestration step.



Substrate-Aware AI Agents: Execution Context as a First-Class Input

Manu Agrawal arXiv: 2609.05232

Substrate blindness is planning without memory, runtime, compute, and operational constraints in the agent's state. Through numerical code generation and related settings, the paper argues execution context should be a first-class input when choosing suitable plans.

Key insight: Feed memory/runtime/compute constraints into planning—do not treat the substrate as invisible.



CONTINUITY: Security-Context Contracts for Composable LLM Agent Controls

Chris Zheng; Geng Yang arXiv: 2609.05269

Individually correct provenance, authorization, policy, adapter, and execution controls can still fail end-to-end when security-critical context is dropped, widened, rebound, or reinterpreted across boundaries. CONTINUITY proposes security-context contracts so composable agent controls keep that context intact.

Key insight: Compose agent security with explicit security-context contracts across component boundaries.



Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability

Ankit Goyal; Jaideep Ray arXiv: 2609.05339

Three-column study design for memory portability: synthetic histories, memory construction modes LC-RAW RAG NOTES KG, and migration plus evaluation across models
Controlled memory-portability study: write, store, retrieve, and read across model upgrades

Keeping the same memory store does not guarantee the upgraded model behaves the same: notes can be reinterpreted, mixed embeddings can break retrieval, and repair may need original evidence. The controlled study compares LC-RAW, RAG, notes, and fixed-schema KG stores under same versus mixed embedders and migration directions.

Key insight: Same store ≠ portable memory—test reinterpretation, embedding mix, and repair before model swaps.



CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents

Haoting Shi et al. arXiv: 2609.05374

CUA-Universe results panel showing success-rate and score gains on CUA-Verse, OSWorld, and OSWorld-MCP with lower steps and tokens versus baselines
CUA-Universe reports large success-rate gains with fewer steps and tokens versus GUI-only or base models

Real computer work mixes visual-state inspection with high-throughput CLI, but many computer-use agents still act mostly through the GUI. CUA-Universe supplies scalable hybrid GUI+CLI environments so agents can coordinate both modalities over shared application state, with large gains versus GUI-only baselines on reported benches.

Key insight: Prefer hybrid GUI+CLI computer-use environments; CLI when available beats GUI-only trajectories.



GRACE: Graph-Grounded Reflective Agent Copilot Engine for Expert-in-the-Loop Knowledge Expansion

John Seon Keun Yi; Joshua R. Minot; Dokyun Lee arXiv: 2609.04442

Standard RAG retrieves isolated passages without tracking cross-document evidence or quantifying uncertainty. GRACE deconstructs claims into a graph-grounded reflective copilot loop so expert-in-the-loop knowledge expansion stays tied to evidence relationships.

Key insight: Ground high-stakes copilots in cross-document evidence graphs, not isolated RAG hits.



When Quantization Breaks Memory: Recurrent-State Write-Back in Low-Precision Temporal Inference

Ismail Erbas; Xavier Intes; Vikas Pandey arXiv: 2609.04490

In recurrent nets the quantized state is stored and returned next step, so the write-back rule can alter subsequent computation. Isolating recurrent-state write-back in a compact GRU encoder–decoder shows how low-precision temporal inference can corrupt memory-like state.

Key insight: Quantized recurrent write-back can silently corrupt temporal state—choose the store rule deliberately.



Atlas: Optimizing Deployment of Compound AI Workflows on Heterogeneous Clusters

Milos Gravara; Andrija Stanisic; Stefan Nastic arXiv: 2609.04513

Compound AI workflows expose many model variants and placements per stage. Atlas searches execution plans that select models and place them on heterogeneous clusters so deployment cost and performance trade-offs are explicit.

Key insight: Treat compound-workflow deployment as joint model-selection and placement search.



Dynamic Adaptation of the LLM Context for Generating Routines with Coupled Semantics

Gnaneswar Villuri; Hashmath Shaik; Alex Doboli arXiv: 2609.04570

When one routine's meaning depends on another's runtime behavior, static context binding fails. The paper studies dynamic LLM context adaptation for code generation under coupled semantics that textual descriptions alone cannot resolve.

Key insight: Adapt context dynamically when routine correctness depends on joint runtime behavior.



Training-Free Halving of Activated Experts in Fine-Grained Mixture-of-Experts Models

Xing Chen; Hengshuai Yao arXiv: 2609.04575

Fine-grained MoE renormalization calibrates expert gain to the training top-k, so cutting k at inference changes both which experts fire and branch strength. Separating those effects enables training-free halving of activated experts with controlled quality impact.

Key insight: Halve activated MoE experts at inference only after accounting for renormalization gain.



SiLR: Structure-Preserving Admission and Process Reward for LLM Tool Agents

Chenyu Zhou et al. arXiv: 2609.04629

A runtime gate in a ReAct loop is a search operator over proposals, not merely a filter. SiLR studies post-violation recovery admission with structure-preserving criteria and process reward so progress can continue while the system is still in violation.

Key insight: Design tool-agent gates as structure-preserving admission, not reject-and-retry filters.



Train What You Deploy: Token-Faithful Post-Training of a Production Coding Agent

Cheng Li et al. arXiv: 2609.04678

Post-training often mismatches production tokens and controls when simplified envs or offline log reconstruction distort prompts. A fidelity-aware coupling keeps trainer-side sampling aligned with the tokens and controls the deployed coding agent actually sees.

Key insight: Post-train coding agents on the exact tokens and controls you deploy.



PLUME: Parameter-Efficient Personalization of Large Language Models via Low-Rank User Modulation in Shared Subspaces

Xinyu Li et al. arXiv: 2609.04715

Per-user fine-tuning personalizes well but does not scale. PLUME uses low-rank user modulation in shared subspaces so personalization quality stays high without per-user full adapters.

Key insight: Personalize with shared-subspace low-rank user modulation instead of full per-user finetunes.



DCFA: Dual-view Causal-inspired Attribution for Failure Reasoning in LLM-based Multi-agent Systems

Zehao Wang et al. arXiv: 2609.04749

Multi-agent failures require tracing natural-language interactions to the decisive earliest error. DCFA uses dual-view causal-inspired attribution to reason about which agent action caused system-level failure.

Key insight: Attribute multi-agent failures to the earliest decisive error with dual-view causal tracing.



Persistent Teacher Anchoring for Tool-Using Agents

Hyun Bin Park et al. arXiv: 2609.04773

On-policy distillation matches student next-token distributions to a teacher, but as rollouts enter states the teacher would not visit the gap accumulates. Persistent teacher anchoring stabilizes tool-agent distillation when teacher quality varies across the trajectory.

Key insight: Anchor tool-agent distillation to the teacher persistently as rollouts leave teacher states.



Whose record is this? Diagnosing and authorizing record use in personalized multimodal models

Xinyu Mao et al. arXiv: 2609.04801

Visual personalization can retrieve a true record yet apply it to the wrong subject. Record authorization requires subject presence, record-edge validity, and answer support; violations are visual memory misbinding.

Key insight: Authorize personalized records with presence×edge×support—retrieval alone is not enough.



RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents

Aziz Ben Amor et al. arXiv: 2609.04898

Repository-scale refactoring needs agents to propagate one change across interdependent files without altering behavior. RefactorPlatform holds the environment fixed and varies design axes—model backbone, prompts, tools—so success factors are isolable.

Key insight: Evaluate repo-scale refactoring agents with a harness that varies one design axis at a time.



ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs

Zukang Xu et al. arXiv: 2609.05228

Fixed top-k MoE routing wastes computation on low-contribution experts. ACE skips experts adaptively without calibration data or extra training by estimating actual routed-expert contribution at inference.

Key insight: Skip MoE experts adaptively without calibration when contribution estimates are reliable.



Uncensored Open-weight Models: Redistribution as the Persistence Layer

10a Labs et al. arXiv: 2609.05241

An ecosystem removes safety guardrails from open-weight models and redistributes them at scale. Profiling producers, reproductions, and applications from 2024–2026 shows redistribution itself acting as the persistence layer for uncensoring—not a single model drop.

Key insight: Treat redistribution networks as the persistence layer for uncensoring open-weight models.



Trace2Tower: Transition-Aware EigenTrace Induction of Multi-Level Skills for LLM Agents

Jiazheng Sun et al. arXiv: 2609.05261

3D scatter of Trace2Tower EigenTrace clusters separating skill-related latent states on axes q2 q3 q4
Trace2Tower separates multi-level skill structure in transition-aware EigenTrace space

Flat trajectory retrieval and shallow skill summarization ignore temporal dependencies and outcome-conditioned topology. Trace2Tower distills raw trajectories into multi-level skills with transition-aware EigenTrace induction.

Key insight: Induce multi-level skills from transition-aware traces, not flat trajectory summaries.



Testing Interchangeability in LLM Agent Teams

Jianxin Gao et al. arXiv: 2609.05279

Production multi-agent systems often assume role-matched agents are interchangeable. Forming teams independently, then trading role-matched agents while each keeps a private notebook, tests what actually breaks under swap.

Key insight: Measure role-swap cost empirically before treating agents as interchangeable.



Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe

Dain Kim et al. arXiv: 2609.05395

KOPA-Bench covers 145 real-world multi-step tasks over live Korean open public APIs under on-premise open-source constraints, plus a data-synthesis recipe for closing the gap where open models underperform.

Key insight: Benchmark and synthesize multi-step tool-calling on live public APIs for on-prem open models.



Beyond Code Generation: Reliability, Verification, and Cost Economics in the Agentic Software Development Lifecycle

Happy Bhati arXiv: 2609.04681

As coding agents inspect repos, edit files, run tools, and open PRs for long stretches, the bottleneck shifts from typing code to reliability, verification, and cost. The survey frames where field productivity gains attenuate once agents own more of the SDLC.

Key insight: Agentic SDLC gains hinge on verification and cost controls, not autocomplete speed alone.



Diffusion Language Models for Mobile Edge Agentic AI: Foundations, Applications, and Challenges

Chenqi Li et al. arXiv: 2609.04778

Diffusion language models refine tokens by iterative denoising rather than left-to-right decoding, enabling parallel updates and bidirectional context. This survey maps foundations, applications, and challenges for mobile-edge agentic AI.

Key insight: Watch diffusion LMs as a non-autoregressive option for mobile-edge agents as tooling matures.



Linguistic Trajectory Encoding for Efficient Long-Horizon Spatial Memory in Embodied Agents

Tianyidan Xie et al. arXiv: 2609.04802

Long-horizon embodied agents need object state transitions queryable in natural language across hours to days. Linguistic trajectory encoding compresses motion into language-addressable memory instead of dropping it in clip embeddings or keeping only raw coordinates.

Key insight: Encode long-horizon spatial change as linguistic trajectories for queryable persistent memory.



From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments

Linsen Zhu; Mengqing Cai arXiv: 2609.04894

Models become consequential when surrounding systems let outputs change external state—tools, interfaces, delegation, persistence, generated worlds, robots. This survey separates model competence, system integration, persistence, and safe authority across digital, social, virtual, and physical settings.

Key insight: Separate model skill from integration, persistence, and authority when mapping agentic progress.