Friday's cs.AI announcement day (2026-10-02) lists 149 new and 232 cross-lists (replacements skipped; listing total 381). Stack filter for agent systems, memory/context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent, persistence/identity, and local/open models keeps 40 papers — local-first harnesses (Mingbird, K-Dense BYOK), GUI self-improvement (GUI-HARVEST, Component Routing), memory systems (Heavy-Tailed Memory, MemFit, Mem++, CAVE-Mem), and harness/skill evolution (Praxa, ActiveSaddler, VeriHarness, Chaining Skills to Hijack).
Hao Wang; Ting Huang arXiv: 2610.02001
Small open-weight models (2-9B) run on ordinary laptops, but under cloud-scale agent harnesses they rarely complete real tasks: tool prefill overflows the context, self-correction diverges, tool demonstrations loop, and tasks are silently abandoned.
Key insight: Mingbird targets local or open-weight models for agent harnesses under device constraints.
Aubrey M. Brueckner; Darshil Patel; Yuhuan He et al. arXiv: 2610.00074
K-Dense BYOK (bring your own keys) is a free, open-source AI research assistant for scientists in any field that runs on the researcher's own computer. The researcher supplies access to a model of their choice, hosted or running locally, and the application supplies everything else: a place for the work to run, a layer of scientific scaffolding, and a complete record.
Key insight: K-Dense BYOK targets local or open-weight models for agent harnesses under device constraints.
Geyi Yang; Zikun Qu; Xiang Li et al. arXiv: 2610.00948
The executable harness surrounding a GUI model determines how observations are assembled, actions are executed, and verification, recovery, and termination are controlled.
Key insight: GUI-HARVEST advances computer-use / GUI agent competence or self-improvement.
Beining Wu; Zihao Ding; Jun Huang arXiv: 2610.01787
Self-improving GUI agents keep the trajectories they produce and return them to the agent, by fine-tuning or by retrieval into the prompt, and studies that compare the two destinations disagree.
Key insight: Not All Experience Belongs in the advances computer-use / GUI agent competence or self-improvement.
Xinyuan Song; Zekun Cai arXiv: 2610.00010
Long-horizon language agents increasingly rely on external memory as a frozen world model, yet current memory systems are usually judged only by task success or token cost.
Key insight: Heavy-Tailed Memory Traces in Long-Horizon Language proposes a concrete memory or context mechanism for long-horizon agents.
Juli Huang arXiv: 2610.00366
A persistent agent must decide both what to retain as information arrives and what to surface once a query appears, yet memory evaluations can confound these decisions by comparing methods that differ in both retention and selection.
Key insight: What Should an Agent Remember? Disentangling proposes a concrete memory or context mechanism for long-horizon agents.
Mitchell Piehl; Muchao Ye arXiv: 2610.00872
Long-term memory systems for large language models (LLMs) have gained popularity for extending reasoning capabilities across applications. Current memory systems rely on LLM agents to organize and consolidate memory, resulting in costly, inefficient write operations.
Key insight: MemFit proposes a concrete memory or context mechanism for long-horizon agents.
Yu Luo; Jiamin Jiang; Yimin Zuo et al. arXiv: 2610.01415
Large language model (LLM) agents can now undertake increasingly complex tasks, but the way they organize interaction history into memory does not ensure a coherent understanding of the current world.
Key insight: Beyond Memory proposes a concrete memory or context mechanism for long-horizon agents.
Ahmad Yehia; Aly O. Abdelkareem; Islam Ahmed et al. arXiv: 2610.02002
Large Language Model (LLM) agents now take part in organizational work, where many authors record decisions across documents over months. Because a revised decision arrives as a new document rather than an edit, answering a question requires knowing which version held at a given time.
Key insight: Mem++ proposes a concrete memory or context mechanism for long-horizon agents.
Xinyu Li arXiv: 2610.00238
Long-term memory agents increasingly rely on it- erative search and reusable experience to answer questions over large personal, factual, or narrative histories.
Key insight: CAVE-Mem proposes a concrete memory or context mechanism for long-horizon agents.
Arman Behnam; Binghui Wang arXiv: 2610.02070
Memory-augmented large language models must decide which memories to retain, and recent systems do so by estimating each memory's effect on task performance. However, these estimates rely entirely on retrieved memories.
Key insight: Causal Memory Policy proposes a concrete memory or context mechanism for long-horizon agents.
Zhiyun Shi arXiv: 2610.01118
A long-term conversational assistant must recall the right memory at the right moment, yet the memory that matters most is often not similar to what the user says now.
Key insight: Madeleine proposes a concrete memory or context mechanism for long-horizon agents.
Pranav Singh arXiv: 2610.00094
Belief-based agent memory needs reliable decisions about current state, yet its evidence may be noisy, copied, or stale. Must a memory calibrate its sources before it can improve its decisions?
Key insight: Nous proposes a concrete memory or context mechanism for long-horizon agents.
Stefan G. Creadore arXiv: 2610.00015
Large-language-model agents can propose and execute actions, but proposal, authority, dispatch, verified external effect, and serving promotion are different claims.
Key insight: From Proposal to Verified Effect: Praxa, treats skills or harnesses as the evolvable control surface around the model.
Jundong Hu; Shekar Ramachandran arXiv: 2610.00025
Agent harnesses increasingly want to run small language models (SLMs) on the microtasks around a frontier large language model (LLM) planner: auto-approving shell commands, writing memory, selecting tools, ranking past turns.
Key insight: Measuring the Microtask Eligibility Gap treats skills or harnesses as the evolvable control surface around the model.
Sungho Park; Wonjoong Kim; Jue Zhang et al. arXiv: 2610.00906
Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback.
Key insight: ActiveSaddler treats skills or harnesses as the evolvable control surface around the model.
Yixuan Li; Yiyun Zhou; Yao Long Teng et al. arXiv: 2610.00917
Choosing an agent system means choosing both a language model and the harness through which it acts. We ask whether a strong model, harness, or pairing stays strong when the setting changes.
Key insight: Finding the Right Fit treats skills or harnesses as the evolvable control surface around the model.
Caiqi Zhang; Rujun Han; Zifeng Wang et al. arXiv: 2610.00972
As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time.
Key insight: VeriHarness treats skills or harnesses as the evolvable control surface around the model.
Shuyao Xiao; Shengling Wang; Xuan Chen et al. arXiv: 2610.00372
Large language model agents rely on external harnesses to pass information between the model and its environment and to recover from execution errors. Yet recovery is usually judged only by average task success.
Key insight: When Harnesses Lose the Signal treats skills or harnesses as the evolvable control surface around the model.
Qiushi Han; Keya Hu; Linlu Qiu et al. arXiv: 2610.02200
We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments.
Key insight: VISTA treats skills or harnesses as the evolvable control surface around the model.
Zongrui Yang; Li Xintong; Runchen Xu et al. arXiv: 2610.01506
As agents evolve from single-tool systems into modular, composite architectures, skills are becoming an important mechanism for capability development and distribution.
Key insight: MCRI treats skills or harnesses as the evolvable control surface around the model.
Tian Dong; Zixuan Ma; Haodong Zhao et al. arXiv: 2610.01564
LLM agents use skills to improve performance on specialized tasks. To complete a user request, an agent may invoke several skills in sequence, allowing information produced under one skill to guide the next.
Key insight: Chaining Skills to Hijack LLM Agents treats skills or harnesses as the evolvable control surface around the model.
Yoonkyu Woo; Woojin Lee; Jin-Xia Huang arXiv: 2610.01097
End-to-end research agents can now produce complete scientific papers, yet manuscript claims often diverge from executed experiments. This gap is structural: research state, failure histories, and claim-evidence alignment are not maintained as persistent, verifiable state across long-horizon pipelines.
Key insight: YouRA addresses persistence, identity, or authority across agent runtimes.
Corinn Tiffany; Wen Zhang; Eugene Bagdasarian et al. arXiv: 2610.00797
Contextual security defenses prevent AI agents from taking rogue actions by synthesizing a task-specific policy and enforcing it on the agent's tool calls. In multi-step tasks, however, which actions are valid often depends on what the agent has already done and learned.
Key insight: Sapien addresses persistence, identity, or authority across agent runtimes.
Miaobo Hu; Shuhao Hu; Xiaobo Guo et al. arXiv: 2610.00327
Tool-using agents can expose citations and execution logs while leaving a critical association unaudited: whether the claim shown to a user is the claim emitted by the committed execution and supported by the cited source.
Key insight: Actions with Receipts tightens tool-call provenance, receipts, or capability enforcement for agents.
Fengpeng Li; Qizhou Wang; Yuke Hu et al. arXiv: 2610.01349
Tool-using large language model (LLM) agents turn generated text into real side effects, so poisoned tool metadata, retrieved pages, memory, and reusable skills can steer the next call.
Key insight: PACE tightens tool-call provenance, receipts, or capability enforcement for agents.
Genliang Zhu; Chu Wang arXiv: 2610.00347
Self-modifying AI agents can replace, fork, and roll back identity-bearing software while descendants remain executable. Per-successor authorization does not constrain the resulting population: siblings may duplicate quotas, combine permissions, survive ancestor cuts, or overlap predecessors during promotion.
Key insight: Authorization for Self-Modifying AI Agent Populations: addresses persistence, identity, or authority across agent runtimes.
Yunbei Zhang; Saiyue Lyu; Janet Wang et al. arXiv: 2610.00371
Multi-agent systems derive their capabilities from sharing evidence, delegating tasks, and combining information across agents. The same process creates a safety problem: contributions that are admissible in isolation can jointly enable a prohibited use.
Key insight: Deny Without Disabling studies multi-agent collaboration, failure modes, or workflow optimization.
Jiaqi Tang; Lan Wei; Bingyu Shen et al. arXiv: 2610.01045
A chatbot that says "I will remind you tomorrow" will not run again until the user writes. We call such a promise an empty commitment: a promise of an action after the current turn that nothing in the agent's tools or runtime can carry out.
Key insight: Empty Commitments addresses persistence, identity, or authority across agent runtimes.
Haotian Chen; Bowen Ye; Yuning Zhang et al. arXiv: 2610.01138
Concurrent actions in large language model (LLM) agent environments require arbitration even when each proposal is individually valid. We implement a typed snapshot-settlement contract and audit three distinct properties: order sensitivity, useful progress, and replay consistency.
Key insight: Auditing Action Settlement in LLM Agent tightens tool-call provenance, receipts, or capability enforcement for agents.
Genliang Zhu; Chu Wang arXiv: 2610.00349
Resource limits are becoming an authorization boundary for AI agents that delegate work across concurrent and failure-prone workers. Parent-child allocation constraints, affine objects, and distributed escrow do not by themselves prevent overspend when replies are lost, effects complete after timeout, messages repeat, branches partition, or DAG joins alias one lineage.
Key insight: Fault-Tolerant Budget Conservation in Distributed Multi-Agent studies multi-agent collaboration, failure modes, or workflow optimization.
Sahan Paliskara; Nattaput Namchittai; Andrew Lampinen arXiv: 2610.00583
People are increasingly delegating tasks to AI agents, and those agents are increasingly encountering other people's agents over shared resources such as a codebase, a calendar, or a budget.
Key insight: Worse Together studies multi-agent collaboration, failure modes, or workflow optimization.
Xuehang Guo; Haoyu Wang; Shengyu Chen et al. arXiv: 2610.01017
Large language models (LLMs) increasingly construct multi-agent workflows that decompose a complex task and assign specialist agents from a pool. However, building such a workflow well remains challenging: how finely to divide the task, which agent to trust with each subtask, and when to create a new specialist are all critical decisions a workflow constructor needs to settle…
Key insight: Pay for the Fault, Not the Flow studies multi-agent collaboration, failure modes, or workflow optimization.
Shixuan Li; Wei Yang; Peiyu Zhang et al. arXiv: 2610.01042
Multi-agent communication aims to help agents benefit from one another's information. Yet improvements in system performance leave a fundamental ambiguity: do they reflect effective communication, a favorable agent architecture, or simply additional reasoning?
Key insight: Beyond Final Accuracy studies multi-agent collaboration, failure modes, or workflow optimization.
Xin Heng arXiv: 2610.02036
AI agents can each make locally valid decisions yet jointly produce an invalid result. We call this the global coherence problem: a failure of shared state, not merely of model intelligence.
Key insight: Global Coherence studies multi-agent collaboration, failure modes, or workflow optimization.
Birk Torpmann-Hagen; Finn Schwall; Leon Moonen arXiv: 2610.00430
Autonomous large language model (LLM) agents increasingly interact in network environments where adversarial content can propagate between agents. Known attacks include agent worms, which spread through self-replicating prompt injections or configuration compromises.
Key insight: Memetic Trojans studies multi-agent collaboration, failure modes, or workflow optimization.
Xisen Jin; Jingheng Li; Zhenglun Chen et al. arXiv: 2610.00710
As large language model (LLM) agents become widely adopted, they are increasingly deployed for tasks that require persistent monitoring or recurring actions (e.g., market analysis).
Key insight: ReLiveGym advances long-horizon evaluation, reliability, or professional-work benchmarks.
Stephanie Finley; Liudas Panavas; Thomas Mikkelson et al. arXiv: 2610.01306
Professional work often starts with a brief request that leaves the professional to work out what is needed, which documents matter, and whether the request's premise holds.
Key insight: DAYJOB advances long-horizon evaluation, reliability, or professional-work benchmarks.
Michael Hardy; Ruhana Azam; Anka Reuel et al. arXiv: 2610.00651
Agent evaluations are increasingly used to compare LLMs and inform deployment decisions, yet ranks can reflect not only the model but also the effects of the evaluation conditions such as the scaffolds or tasks.
Key insight: Agent Evaluation Reliability advances long-horizon evaluation, reliability, or professional-work benchmarks.
Andre Fu; Malik Drabla; Leon Liu et al. arXiv: 2610.00648
AI coding agents are ubiquitous in engineering workflows amongst industry and academia. Yet, despite their use in app coding, relatively less attention has been paid to their ability to execute on production incident response.
Key insight: Incident-Arena advances long-horizon evaluation, reliability, or professional-work benchmarks.