Friday's cs.AI announcement day (2026-10-02) lists 149 new and 232 cross-lists (replacements skipped; listing total 381). Stack filter for agent systems, memory/context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent, persistence/identity, and local/open models keeps 40 papers — local-first harnesses (Mingbird, K-Dense BYOK), GUI self-improvement (GUI-HARVEST, Component Routing), memory systems (Heavy-Tailed Memory, MemFit, Mem++, CAVE-Mem), and harness/skill evolution (Praxa, ActiveSaddler, VeriHarness, Chaining Skills to Hijack).


Research Papers

Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks

Hao Wang; Ting Huang arXiv: 2610.02001

Figure from Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks
Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks

Small open-weight models (2-9B) run on ordinary laptops, but under cloud-scale agent harnesses they rarely complete real tasks: tool prefill overflows the context, self-correction diverges, tool demonstrations loop, and tasks are silently abandoned.

Key insight: Mingbird targets local or open-weight models for agent harnesses under device constraints.

K-Dense BYOK: An Open-Source AI Research Assistant That Runs Locally and Keeps a Hash-Chained Lab Notebook

Aubrey M. Brueckner; Darshil Patel; Yuhuan He et al. arXiv: 2610.00074

Figure from K-Dense BYOK: An Open-Source AI Research Assistant That Runs Locally and Keeps a Hash-Chained Lab Notebook
K-Dense BYOK: An Open-Source AI Research Assistant That Runs Locally and Keeps a Hash-Chained Lab Notebook

K-Dense BYOK (bring your own keys) is a free, open-source AI research assistant for scientists in any field that runs on the researcher's own computer. The researcher supplies access to a model of their choice, hosted or running locally, and the application supplies everything else: a place for the work to run, a layer of scientific scaffolding, and a complete record.

Key insight: K-Dense BYOK targets local or open-weight models for agent harnesses under device constraints.

GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution

Geyi Yang; Zikun Qu; Xiang Li et al. arXiv: 2610.00948

Figure from GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution
GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution

The executable harness surrounding a GUI model determines how observations are assembled, actions are executed, and verification, recovery, and termination are controlled.

Key insight: GUI-HARVEST advances computer-use / GUI agent competence or self-improvement.

Not All Experience Belongs in the Weights: Component Routing for Self-Improving GUI Agents

Beining Wu; Zihao Ding; Jun Huang arXiv: 2610.01787

Figure from Not All Experience Belongs in the Weights: Component Routing for Self-Improving GUI Agents
Not All Experience Belongs in the Weights: Component Routing for Self-Improving GUI Agents

Self-improving GUI agents keep the trajectories they produce and return them to the agent, by fine-tuning or by retrieval into the prompt, and studies that compare the two destinations disagree.

Key insight: Not All Experience Belongs in the advances computer-use / GUI agent competence or self-improvement.

Heavy-Tailed Memory Traces in Long-Horizon Language Agents

Xinyuan Song; Zekun Cai arXiv: 2610.00010

Figure from Heavy-Tailed Memory Traces in Long-Horizon Language Agents
Heavy-Tailed Memory Traces in Long-Horizon Language Agents

Long-horizon language agents increasingly rely on external memory as a frozen world model, yet current memory systems are usually judged only by task success or token cost.

Key insight: Heavy-Tailed Memory Traces in Long-Horizon Language proposes a concrete memory or context mechanism for long-horizon agents.

What Should an Agent Remember? Disentangling Retention from Retrieval in Bounded-Memory Evaluation

Juli Huang arXiv: 2610.00366

Figure from What Should an Agent Remember? Disentangling Retention from Retrieval in Bounded-Memory Evaluation
What Should an Agent Remember? Disentangling Retention from Retrieval in Bounded-Memory Evaluation

A persistent agent must decide both what to retain as information arrives and what to surface once a query appears, yet memory evaluations can confound these decisions by comparing methods that differ in both retention and selection.

Key insight: What Should an Agent Remember? Disentangling proposes a concrete memory or context mechanism for long-horizon agents.

MemFit: Efficient Long-Term Agentic Memory

Mitchell Piehl; Muchao Ye arXiv: 2610.00872

Figure from MemFit: Efficient Long-Term Agentic Memory
MemFit: Efficient Long-Term Agentic Memory

Long-term memory systems for large language models (LLMs) have gained popularity for extending reasoning capabilities across applications. Current memory systems rely on LLM agents to organize and consolidate memory, resulting in costly, inefficient write operations.

Key insight: MemFit proposes a concrete memory or context mechanism for long-horizon agents.

Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States

Yu Luo; Jiamin Jiang; Yimin Zuo et al. arXiv: 2610.01415

Large language model (LLM) agents can now undertake increasingly complex tasks, but the way they organize interaction history into memory does not ensure a coherent understanding of the current world.

Key insight: Beyond Memory proposes a concrete memory or context mechanism for long-horizon agents.

Mem++: Non-Destructive Memory for Long-Term Organizational LLM Agents

Ahmad Yehia; Aly O. Abdelkareem; Islam Ahmed et al. arXiv: 2610.02002

Large Language Model (LLM) agents now take part in organizational work, where many authors record decisions across documents over months. Because a revised decision arrives as a new document rather than an edit, answering a question requires knowing which version held at a given time.

Key insight: Mem++ proposes a concrete memory or context mechanism for long-horizon agents.

CAVE-Mem: Boundary-Aware Experience Validation for Memory Search

Xinyu Li arXiv: 2610.00238

Long-term memory agents increasingly rely on it- erative search and reusable experience to answer questions over large personal, factual, or narrative histories.

Key insight: CAVE-Mem proposes a concrete memory or context mechanism for long-horizon agents.

Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval

Arman Behnam; Binghui Wang arXiv: 2610.02070

Memory-augmented large language models must decide which memories to retain, and recent systems do so by estimating each memory's effect on task performance. However, these estimates rely entirely on retrieved memories.

Key insight: Causal Memory Policy proposes a concrete memory or context mechanism for long-horizon agents.

Madeleine: Learning Involuntary Recall for Conversational Memory from Simulated Lives

Zhiyun Shi arXiv: 2610.01118

A long-term conversational assistant must recall the right memory at the right moment, yet the memory that matters most is often not similar to what the user says now.

Key insight: Madeleine proposes a concrete memory or context mechanism for long-horizon agents.

Nous: Learning and Certifying Memory Decisions Before Source Calibration

Pranav Singh arXiv: 2610.00094

Belief-based agent memory needs reliable decisions about current state, yet its evidence may be noisy, copied, or stale. Must a memory calibrate its sources before it can improve its decisions?

Key insight: Nous proposes a concrete memory or context mechanism for long-horizon agents.

From Proposal to Verified Effect: Praxa, an Evidence-Bound Harness for Governed AI Agent Execution

Stefan G. Creadore arXiv: 2610.00015

Figure from From Proposal to Verified Effect: Praxa, an Evidence-Bound Harness for Governed AI Agent Execution
From Proposal to Verified Effect: Praxa, an Evidence-Bound Harness for Governed AI Agent Execution

Large-language-model agents can propose and execute actions, but proposal, authority, dispatch, verified external effect, and serving promotion are different claims.

Key insight: From Proposal to Verified Effect: Praxa, treats skills or harnesses as the evolvable control surface around the model.

Measuring the Microtask Eligibility Gap: When Is an Off-the-Shelf SLM Enough for an Agent Harness?

Jundong Hu; Shekar Ramachandran arXiv: 2610.00025

Agent harnesses increasingly want to run small language models (SLMs) on the microtasks around a frontier large language model (LLM) planner: auto-approving shell commands, writing memory, selecting tools, ranking past turns.

Key insight: Measuring the Microtask Eligibility Gap treats skills or harnesses as the evolvable control surface around the model.

ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

Sungho Park; Wonjoong Kim; Jue Zhang et al. arXiv: 2610.00906

Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback.

Key insight: ActiveSaddler treats skills or harnesses as the evolvable control surface around the model.

Finding the Right Fit: Model-Harness Interactions across Agent Tasks

Yixuan Li; Yiyun Zhou; Yao Long Teng et al. arXiv: 2610.00917

Choosing an agent system means choosing both a language model and the harness through which it acts. We ask whether a strong model, harness, or pairing stays strong when the setting changes.

Key insight: Finding the Right Fit treats skills or harnesses as the evolvable control surface around the model.

VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

Caiqi Zhang; Rujun Han; Zifeng Wang et al. arXiv: 2610.00972

As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time.

Key insight: VeriHarness treats skills or harnesses as the evolvable control surface around the model.

When Harnesses Lose the Signal: Causal Evaluation of Recovery in LLM Agents

Shuyao Xiao; Shengling Wang; Xuan Chen et al. arXiv: 2610.00372

Large language model agents rely on external harnesses to pass information between the model and its environment and to recover from execution errors. Yet recovery is usually judged only by average task success.

Key insight: When Harnesses Lose the Signal treats skills or harnesses as the evolvable control surface around the model.

VISTA: A Visual Harness for Reasoning in an Interactive World

Qiushi Han; Keya Hu; Linlu Qiu et al. arXiv: 2610.02200

We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments.

Key insight: VISTA treats skills or harnesses as the evolvable control surface around the model.

MCRI: A Four-Dimensional Framework for Analyzing and Evaluating Agent Skills

Zongrui Yang; Li Xintong; Runchen Xu et al. arXiv: 2610.01506

As agents evolve from single-tool systems into modular, composite architectures, skills are becoming an important mechanism for capability development and distribution.

Key insight: MCRI treats skills or harnesses as the evolvable control surface around the model.

Chaining Skills to Hijack LLM Agents

Tian Dong; Zixuan Ma; Haodong Zhao et al. arXiv: 2610.01564

LLM agents use skills to improve performance on specialized tasks. To complete a user request, an agent may invoke several skills in sequence, allowing information produced under one skill to guide the next.

Key insight: Chaining Skills to Hijack LLM Agents treats skills or harnesses as the evolvable control surface around the model.

YouRA: A Persistent-State Architecture for Evidence-Traceable Autonomous Research Agents

Yoonkyu Woo; Woojin Lee; Jin-Xia Huang arXiv: 2610.01097

End-to-end research agents can now produce complete scientific papers, yet manuscript claims often diverge from executed experiments. This gap is structural: research state, failure histories, and claim-evidence alignment are not maintained as persistent, verifiable state across long-horizon pipelines.

Key insight: YouRA addresses persistence, identity, or authority across agent runtimes.

Sapien: A Stateful Policy Engine for Autonomous AI Agents

Corinn Tiffany; Wen Zhang; Eugene Bagdasarian et al. arXiv: 2610.00797

Contextual security defenses prevent AI agents from taking rogue actions by synthesizing a task-specific policy and enforcing it on the agent's tool calls. In multi-step tasks, however, which actions are valid often depends on what the agent has already done and learned.

Key insight: Sapien addresses persistence, identity, or authority across agent runtimes.

Actions with Receipts: Jointly Binding Claims, Evidence, and Execution for Replayable Tool-Agent Auditing

Miaobo Hu; Shuhao Hu; Xiaobo Guo et al. arXiv: 2610.00327

Tool-using agents can expose citations and execution logs while leaving a critical association unaudited: whether the claim shown to a user is the claim emitted by the committed execution and supported by the cited source.

Key insight: Actions with Receipts tightens tool-call provenance, receipts, or capability enforcement for agents.

PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents

Fengpeng Li; Qizhou Wang; Yuke Hu et al. arXiv: 2610.01349

Tool-using large language model (LLM) agents turn generated text into real side effects, so poisoned tool metadata, retrieved pages, memory, and reusable skills can steer the next call.

Key insight: PACE tightens tool-call provenance, receipts, or capability enforcement for agents.

Authorization for Self-Modifying AI Agent Populations: Conserving Authority across Replacement, Forking, and Rollback

Genliang Zhu; Chu Wang arXiv: 2610.00347

Self-modifying AI agents can replace, fork, and roll back identity-bearing software while descendants remain executable. Per-successor authorization does not constrain the resulting population: siblings may duplicate quotas, combine permissions, survive ancestor cuts, or overlap predecessors during promotion.

Key insight: Authorization for Self-Modifying AI Agent Populations: addresses persistence, identity, or authority across agent runtimes.

Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems

Yunbei Zhang; Saiyue Lyu; Janet Wang et al. arXiv: 2610.00371

Multi-agent systems derive their capabilities from sharing evidence, delegating tasks, and combining information across agents. The same process creates a safety problem: contributions that are admissible in isolation can jointly enable a prohibited use.

Key insight: Deny Without Disabling studies multi-agent collaboration, failure modes, or workflow optimization.

Empty Commitments: When Agents Promise What Their Runtime Cannot Deliver

Jiaqi Tang; Lan Wei; Bingyu Shen et al. arXiv: 2610.01045

A chatbot that says "I will remind you tomorrow" will not run again until the user writes. We call such a promise an empty commitment: a promise of an action after the current turn that nothing in the agent's tools or runtime can carry out.

Key insight: Empty Commitments addresses persistence, identity, or authority across agent runtimes.

Auditing Action Settlement in LLM Agent Environments: Order, Progress, and Replay

Haotian Chen; Bowen Ye; Yuning Zhang et al. arXiv: 2610.01138

Concurrent actions in large language model (LLM) agent environments require arbitration even when each proposal is individually valid. We implement a typed snapshot-settlement contract and audit three distinct properties: order sensitivity, useful progress, and replay consistency.

Key insight: Auditing Action Settlement in LLM Agent tightens tool-call provenance, receipts, or capability enforcement for agents.

Fault-Tolerant Budget Conservation in Distributed Multi-Agent Delegation

Genliang Zhu; Chu Wang arXiv: 2610.00349

Resource limits are becoming an authorization boundary for AI agents that delegate work across concurrent and failure-prone workers. Parent-child allocation constraints, affine objects, and distributed escrow do not by themselves prevent overspend when replies are lost, effects complete after timeout, messages repeat, branches partition, or DAG joins alias one lineage.

Key insight: Fault-Tolerant Budget Conservation in Distributed Multi-Agent studies multi-agent collaboration, failure modes, or workflow optimization.

Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams

Sahan Paliskara; Nattaput Namchittai; Andrew Lampinen arXiv: 2610.00583

People are increasingly delegating tasks to AI agents, and those agents are increasingly encountering other people's agents over shared resources such as a codebase, a calendar, or a budget.

Key insight: Worse Together studies multi-agent collaboration, failure modes, or workflow optimization.

Pay for the Fault, Not the Flow: Label-Free In-Flow Multi-Agent Workflow Optimization

Xuehang Guo; Haoyu Wang; Shengyu Chen et al. arXiv: 2610.01017

Large language models (LLMs) increasingly construct multi-agent workflows that decompose a complex task and assign specialist agents from a pool. However, building such a workflow well remains challenging: how finely to divide the task, which agent to trust with each subtask, and when to create a new specialist are all critical decisions a workflow constructor needs to settle…

Key insight: Pay for the Fault, Not the Flow studies multi-agent collaboration, failure modes, or workflow optimization.

Beyond Final Accuracy: Auditing Communication in LLM Multi-Agent Systems

Shixuan Li; Wei Yang; Peiyu Zhang et al. arXiv: 2610.01042

Multi-agent communication aims to help agents benefit from one another's information. Yet improvements in system performance leave a fundamental ambiguity: do they reflect effective communication, a favorable agent architecture, or simply additional reasoning?

Key insight: Beyond Final Accuracy studies multi-agent collaboration, failure modes, or workflow optimization.

Global Coherence: When Every Agent Is Right and the Team Is Still Wrong - A Local-to-Global Semantic Foundation for Multi-Agent Collaboration

Xin Heng arXiv: 2610.02036

AI agents can each make locally valid decisions yet jointly produce an invalid result. We call this the global coherence problem: a failure of shared state, not merely of model intelligence.

Key insight: Global Coherence studies multi-agent collaboration, failure modes, or workflow optimization.

Memetic Trojans: Social Contagions as Carriers of Adversarial Payloads in Agent Networks

Birk Torpmann-Hagen; Finn Schwall; Leon Moonen arXiv: 2610.00430

Autonomous large language model (LLM) agents increasingly interact in network environments where adversarial content can propagate between agents. Known attacks include agent worms, which spread through self-replicating prompt injections or configuration compromises.

Key insight: Memetic Trojans studies multi-agent collaboration, failure modes, or workflow optimization.

ReLiveGym: Evaluating Long-Lived Agents over Weeks of Replayed Reality

Xisen Jin; Jingheng Li; Zhenglun Chen et al. arXiv: 2610.00710

As large language model (LLM) agents become widely adopted, they are increasingly deployed for tasks that require persistent monitoring or recurring actions (e.g., market analysis).

Key insight: ReLiveGym advances long-horizon evaluation, reliability, or professional-work benchmarks.

DAYJOB: A Benchmark for Long-Horizon Professional Work

Stephanie Finley; Liudas Panavas; Thomas Mikkelson et al. arXiv: 2610.01306

Professional work often starts with a brief request that leaves the professional to work out what is needed, which documents matter, and whether the request's premise holds.

Key insight: DAYJOB advances long-horizon evaluation, reliability, or professional-work benchmarks.

Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard

Michael Hardy; Ruhana Azam; Anka Reuel et al. arXiv: 2610.00651

Agent evaluations are increasingly used to compare LLMs and inform deployment decisions, yet ranks can reflect not only the model but also the effects of the evaluation conditions such as the scaffolds or tasks.

Key insight: Agent Evaluation Reliability advances long-horizon evaluation, reliability, or professional-work benchmarks.

Incident-Arena: Getting agents to the last nine of reliability

Andre Fu; Malik Drabla; Leon Liu et al. arXiv: 2610.00648

AI coding agents are ubiquitous in engineering workflows amongst industry and academia. Yet, despite their use in app coding, relatively less attention has been paid to their ability to execute on production incident response.

Key insight: Incident-Arena advances long-horizon evaluation, reliability, or professional-work benchmarks.