Wednesday’s cs.AI announcement day lists 72 new and 123 cross-lists (replacements skipped; listing total 195). Stack filter for agent systems, memory/context, computer-use / GUI / tools / skills / harnesses, multi-agent, persistence/identity, and local/open models keeps 36 papers — GUI replay memory, lifelong latent memory, 24 GiB local serving, agent-skill governance, and social harnesses for multi-principal agents.


Research Papers

Where Should the KV Cache Live? Placement Policies Across GPU, CPU, and SSD for Long-Lived Sessions

Srikanta Datta Tumkur; Jay Iyer; Mehar Simhadri; Sai Pavan Kumar; … arXiv: 2609.16215

Figure from Where Should the KV Cache Live? Placement Policies Across GPU, CPU, and SSD for Long-Lived Sessions
Where Should the KV Cache Live? Placement Policies Across GPU, CPU, and SSD for Long-Lived Sessions

GPU high bandwidth memory is scarce and expensive, and KV caches consume much of it as chats, agent loops, and document question answering accumulate state.

Key insight: KV cache placement across GPU/CPU/SSD is a first-class policy for long-lived agent sessions.


CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design

Zihan Dong; Yuanzhe Liu; Zhiyuan Ma; Qishi Zhan; … arXiv: 2609.16251

Figure from CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design
CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design

Computer-use agents are increasingly evaluated in realistic desktop environments, but existing benchmarks provide limited coverage of professional engineering workflows whose outputs are persistent, structured artifacts.

Key insight: Computer-use benchmarks must cover long-horizon CAD workflows with persistent structured artifacts.


BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents

Sadia Asif; Mohammad Mohammadi Amiri; Momin Abbas; Tejaswini Pedapati; … arXiv: 2609.16305

Figure from BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents
BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents

Large language model (LLM) agents increasingly operate over long-horizon interactions involving tool use, persistent state, evolving authorization, and external environment feedback.

Key insight: Long-horizon tool-using agents need safety/refusal calibration across multi-turn tool and auth state.


EchoPath: Execution-Level Replayable Memory for GUI Agents

Yao Zhao; Aditya Shanmugham; Swastik Roy; Yanxun Xu arXiv: 2609.16635

Figure from EchoPath: Execution-Level Replayable Memory for GUI Agents
EchoPath: Execution-Level Replayable Memory for GUI Agents

Computer-use agents increasingly operate browsers, software, and desktop applications via CLI or API portals, but graphical user interface (GUI) still plays an important role in common industrial production scenarios.

Key insight: GUI agents can reuse execution-level replayable memory instead of fresh observe-plan-act loops each time.


LSREP: A Longitudinal State-Replay Protocol for Evaluating Conversational Memory, with ICE v2 as an Audited Local-First Architecture

Deepesh Sonar arXiv: 2609.16730

Conversational memory changes during use, so endpoint question answering alone cannot establish how a persistent state accumulates, ages, or incorporates revisions.

Key insight: Conversational memory needs longitudinal state-replay evaluation, not only endpoint QA.


ThinkFlow: Self-Evolving Probabilistic Latent Memory for Lifelong Conversational Agents

Cai Ke; Xin Liu; Han Zhang; Jiangyue Yan; … arXiv: 2609.17010

Figure from ThinkFlow: Self-Evolving Probabilistic Latent Memory for Lifelong Conversational Agents
ThinkFlow: Self-Evolving Probabilistic Latent Memory for Lifelong Conversational Agents

Lifelong conversational agents rely on memory systems to maintain deep, context-aware interactions with users.

Key insight: Lifelong conversational agents need self-evolving latent memory beyond static textual memory pipelines.


Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs

Dushyant Rajput arXiv: 2609.17109

A common small-model deployment runs one shared backbone with several LoRA specialists that answer over the same context.

Key insight: Shared-prefix KV reuse across LoRA specialists avoids re-prefilling the same context.


JustFit: 200K-Token LLM Serving on a 24 GiB Laptop with Just-in-Time State Management

Yuhua Chen arXiv: 2609.17475

Capable open-weight models make local coding and reasoning attractive, but their context and execution state strain laptop memory.

Key insight: Just-in-time KV/state management enables 200K-token open-weight serving on a 24 GiB laptop.


Retrieval-Driven Memory Reconsolidation for Long-Term LLM Agents

Yuanyi Song; Yukai Wang; Xinbei Ma; Zhihui Fu; … arXiv: 2609.16053

Figure from Retrieval-Driven Memory Reconsolidation for Long-Term LLM Agents
Retrieval-Driven Memory Reconsolidation for Long-Term LLM Agents

Long-term memory is essential for LLM-based agents operating over extended interactions.

Key insight: Long-term LLM agents can reconsolidate memory on retrieval rather than only append experiences.


Protocol-Preserving Context Trimming for Agentic Workflows: Benefits, Failure Regimes, and Budget Guardrails

Harish Gaggar arXiv: 2609.16461

Figure from Protocol-Preserving Context Trimming for Agentic Workflows: Benefits, Failure Regimes, and Budget Guardrails
Protocol-Preserving Context Trimming for Agentic Workflows: Benefits, Failure Regimes, and Budget Guardrails

Agentic large language model (LLM) systems rely on long interaction histories to preserve instructions, tool states, intermediate decisions, and unresolved dependencies, but unrestricted context growth increases computational cost and can reduce efficiency.

Key insight: Context trimming for agentic workflows must preserve protocol/tool state under budget pressure.


RepoAtlas: Guiding Coding Agents via Evolving Multimodal Repository Views

Yunxiang Zhang; Haiquan Wang; JiaWei Guo; Hanyang Xia; … arXiv: 2609.16936

Large language model (LLM)-powered coding agents have made rapid progress in automating software engineering tasks, yet repository-level issue resolution remains challenging.

Key insight: Coding agents benefit from evolving multimodal repository views for localization.


After the Party: Governing What a Viral Agent-Skill Ecosystem Left Behind

Yunpeng Xiong; Ting Zhang arXiv: 2609.17274

Figure from After the Party: Governing What a Viral Agent-Skill Ecosystem Left Behind
After the Party: Governing What a Viral Agent-Skill Ecosystem Left Behind

AI agents increasingly act through agent skills, i.e., natural-language instructions, that direct a host agent toward shell, network, credential, file, and process actions, and public registries distribute them at scale.

Key insight: Viral agent-skill ecosystems leave residual governance debt after the install boom.


Agentic Societies Need a Social Harness

Tapan Chugh; Vidushi Singh; Krish Jain; Arvind Krishnamurthy; … arXiv: 2609.17527

Figure from Agentic Societies Need a Social Harness
Agentic Societies Need a Social Harness

An agentic society is a collection of AI agents that coordinate autonomously across trust boundaries, on behalf of different principals whose objectives may only partially align.

Key insight: Multi-principal agentic societies need a social harness beyond single-agent tool harnesses.


Skill-based Agentic Evaluation for Real-time Data Science Tasks

Aniruddha Tamhane; Raghavendra Addanki; Ayushi Aggarwal; Aditya Bansal; … arXiv: 2609.16487

We present a framework for evaluating data-science agents on live, continuously updated data using executable ground truth and format-agnostic factoid scoring.

Key insight: We present a framework for evaluating data-science agents on live, continuously updated data using executable ground truth and format-agnostic factoid scoring.


Turn-level Multiscale Density Ratio Estimation for LLM Agents

Zishuo Zhao; Kai Chen; Ao Li; Yuan Liu arXiv: 2609.16760

With the rapid development of Large language model (LLM), agent systems enhanced by LLMs show huge potential in being able to deal with complex tasks, especially involving multi-step thinking or interaction with tools.

Key insight: With the rapid development of Large language model (LLM), agent systems enhanced by LLMs show huge potential in being able to deal with complex tasks,….


World Model Science: Self-Organized Criticality, Weak Chaos, and Metastable Belief Dynamics in Long-Horizon LLM Agents

Xinyuan Song; Zekun Cai arXiv: 2609.17419

Long-horizon LLM agents must maintain task state across extended sequences of observations, actions, tool calls, and intermediate beliefs.

Key insight: Long-horizon LLM agents must maintain task state across extended sequences of observations, actions, tool calls, and intermediate beliefs.


ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents

Shuhan Xue; Jianyuan Zhong; Ziyuan Nan; Wenbin Li; … arXiv: 2609.17523

We introduce and release ScienceBuddy, an interactive scientific research workspace that brings continually improving scientific agents into researchers' everyday workflows.

Key insight: We introduce and release ScienceBuddy, an interactive scientific research workspace that brings continually improving scientific agents into researchers'….


"Looking for Something Weird to Happen": How Humans Sustain AI Agent Novelty Amid Semantic Collapse

Shiyang Lai; Arna Woemmel; Hongkai Mao; Junsol Kim; … arXiv: 2609.16051

Semantic collapse, the progressive narrowing of what AI systems generate, has been studied mainly in closed settings, and remedies have targeted models and data.

Key insight: Semantic collapse, the progressive narrowing of what AI systems generate, has been studied mainly in closed settings, and remedies have targeted models and….


The Immutable Past: Formalizing State Mutability and Conflict Resolution in Mutable RAG

Hamed HaddadPajouh; Amir AmiriTabat arXiv: 2609.16073

Retrieval-Augmented Generation (RAG) serves as the primary memory architecture for long-horizon autonomous agents.

Key insight: Retrieval-Augmented Generation (RAG) serves as the primary memory architecture for long-horizon autonomous agents.


Coaching Qwen3 Coder 30B to Think Like a CodeClash Arena Agent

Ivy Ning Zhang arXiv: 2609.16096

Large language model coding agents have recently become useful for software tasks, but weaker or open-weight agents still struggle to reliably interpret user intent and execute complex multi-step workflows.

Key insight: Open Qwen3-Coder-30B can be coached toward arena-agent coding behavior.


Assurance Envelopes for Autonomous Coding Agents: Minimum-Cost Evidence for Software Change

Anjan Goswami arXiv: 2609.16302

When a coding agent returns to existing software, it inherits evidence from earlier engineering work: tests, type checks, proofs, static analyses, and traces.

Key insight: When a coding agent returns to existing software, it inherits evidence from earlier engineering work: tests, type checks, proofs, static analyses, and traces.


Grounding SWE-Agent Decisions in Architecture-0 Design: Navigating Unknown Unknowns through Physical Mapping

Zhongkai Wang; Yan Liu arXiv: 2609.17221

Autonomous Software Engineering Agents (SWE-Agents) excel in deterministic coding tasks but struggle with Architecture 0, the nascent system design phase plagued by implicit engineering constraints, or Unknown Unknowns (UUs) that are rarely stated explicitly.

Key insight: Autonomous Software Engineering Agents (SWE-Agents) excel in deterministic coding tasks but struggle with Architecture 0, the nascent system design phase….


Mo' Models, Mo' Problems: How to best select model pools when designing Multi-Agent Systems

Sara Vera Marjanović; Jiacheng Xu; Aleksandr Laptev; Grigor Nalbandyan; … arXiv: 2609.17306

Multi-agent Systems (MAS) combine multiple model outputs to solve complex reasoning tasks.

Key insight: Multi-agent system design hinges on how you select the model pool, not only the topology.


Where Should a Document Live: Context, Representations, or Parameters?

Nathanaël Carraz Rakotonirina; Momchil Hardalov; Gonzalo Iglesias; Adrià de Gispert arXiv: 2609.17346

To answer questions outside of their pre-training data, large language models (LLMs) need access to new information, which can be presented in the context window as documents, encoded into the model's parameters, or injected as latent representations.

Key insight: To answer questions outside of their pre-training data, large language models (LLMs) need access to new information, which can be presented in the context….


Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead

Fengshuo Liu; Ying Liu; Ruize Sun; Lie Luo; … arXiv: 2609.17394

Small differences on coding-agent leaderboards are often read as an ordering of systems.

Key insight: Top SWE-bench entries have converged — leaderboard deltas no longer order systems.


Metacognitive Steering: Learning the Structure of Scientific Judgment

Vincent Karpf; Joseph Reth; Eike Gerhardt; Audrey Wang; … arXiv: 2609.16245

Long-horizon scientific discovery requires agents to alternate between exploration, disciplined execution, and critical reassessment as evidence changes.

Key insight: Long-horizon scientific discovery requires agents to alternate between exploration, disciplined execution, and critical reassessment as evidence changes.


Beyond Episodic AI: Cognitive Field Networks for Biologically Inspired Persistent Cognition

Byung Gyu Chae arXiv: 2609.16752

Cognitive Field Theory (CFT) proposes that cognition arises from memory-dressed collective dynamics that generate a persistent macroscopic cognitive field.

Key insight: Cognitive Field Theory (CFT) proposes that cognition arises from memory-dressed collective dynamics that generate a persistent macroscopic cognitive field.


FlexEE: Self-Speculative and KV-Compatible Early Exiting for Offloading-Aware LLM Inference

Qihu Xie; Ziwei Li; Yi Kang arXiv: 2609.17008

Large language model (LLM) inference is often constrained by both computation and memory, especially in offloading-based deployments where model weights are transferred across memory hierarchies during autoregressive decoding.

Key insight: Large language model (LLM) inference is often constrained by both computation and memory, especially in offloading-based deployments where model weights are….


Interactive Memory Learning for Long-Term Conversations

Cai Ke; Jiangyue Yan; Han Zhang; Xin Liu; … arXiv: 2609.17088

Recent advancements in large language models have significantly enhanced the capabilities of agents in modeling long-term conversations.

Key insight: Recent advancements in large language models have significantly enhanced the capabilities of agents in modeling long-term conversations.


Symbolic Separation: Grounding Deep Agents in Knowledge Graphs for Trustworthy Operational Data Analytics

Baibek Davletiyarov; Junaid Ahmed Khan; Andrea Bartolini arXiv: 2609.17107

Generative AI promises natural language access to the massive numerical telemetry of data centers and Industry 4.0 installations, yet text-to-query and tool-using agents stay unreliable: even frontier models answer little more than half of real-world database questions, and far fewer of the multi-step, operational ones, because the LLM must compose how…

Key insight: Generative AI promises natural language access to the massive numerical telemetry of data centers and Industry 4.0 installations, yet text-to-query and….


Self-Emergence Agent Architecture:Behavior-Inertia HMM, Reflexive Metacognition,and Social-Contrastive Self-Modeling

Xiaoyang Liu arXiv: 2609.17331

Large language model (LLM) agents exhibit strong language-generation and problem-solving capabilities, yet suffer from three structural limitations: personality drift, non-evolutionary reflection, and the absence of a self-other boundary.

Key insight: Large language model (LLM) agents exhibit strong language-generation and problem-solving capabilities, yet suffer from three structural limitations:….


Never Stop Thinking: Continuous-Time Language Agents

Bojie Li; Noah Shi arXiv: 2609.17416

Voice agents built on LLMs follow a rigid listen-think-speak loop that inserts seconds of dead air before every reply.

Key insight: Voice agents built on LLMs follow a rigid listen-think-speak loop that inserts seconds of dead air before every reply.


Verifiable Social Reasoning for LLM Assistants

Amir Taubenfeld; Zorik Gekhman; Avigail Grinstein-Dabush; Itay Laish; … arXiv: 2609.17496

LLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since (i) it requires setups where the assistant learns about social situations from subjective user narratives, and (ii) social properties, such as others' intentions, typically lack verifiable ground truth.

Key insight: LLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since (i) it….


Universal Defenses for Tool-Integrated LLM Agents Against Adversarial Attacks

Xiaoyan Li; Yunli Wang arXiv: 2609.16098

Large Language Model (LLM) agents have demonstrated impressive capabilities across a variety of domains, particularly when integrated with external tools for multi-step task completion.

Key insight: Large Language Model (LLM) agents have demonstrated impressive capabilities across a variety of domains, particularly when integrated with external tools for….


Cognitive Admission Control: Risk-Conditioned Assurance for Consequential Actions in Agentic Distributed Systems

Jun He; Deying Yu arXiv: 2609.16313

In agentic distributed systems, an agent may be authorized to mutate external infrastructure while lacking evidence that the mutation is ready to execute.

Key insight: In agentic distributed systems, an agent may be authorized to mutate external infrastructure while lacking evidence that the mutation is ready to execute.


Interpreting and Steering LLM Agents for Social Simulations

Jiayue Gaveal Fan; Arul Murugan; Shreyas Krishnan; Abhishek Nagaraj arXiv: 2609.16436

Simulations based on large language models (LLMs) have proven to be powerful for understanding human behavior, making them valuable additions to the social scientific toolkit.

Key insight: Simulations based on large language models (LLMs) have proven to be powerful for understanding human behavior, making them valuable additions to the social….