Monday’s cs.AI announcement day (2026-09-21) lists 49 new and 76 cross-lists (replacements skipped; listing total 125). Stack filter for agent systems, memory/context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent, persistence/identity, and local/open serving keeps 32 papers — conversational long-term memory (AutoViewMem), procedural skill memory for tool-heavy agents (Designer-RSI), agentic coding RL (CodeMidas), long-running authorization quiescence, and multi-agent test-time communication.


Research Papers

AutoViewMem: Self-Configuring Orthogonal Views for Conversational Long-Term Memory

Zijie Cao; Xijun Qu; Zhicheng Gu; Xiaoshu Chen; … arXiv: 2609.21940

Figure from AutoViewMem: Self-Configuring Orthogonal Views for Conversational Long-Term Memory
AutoViewMem: Self-Configuring Orthogonal Views for Conversational Long-Term Memory

Long-term memory is essential for large language model (LLM) agents to maintain consistency and personalization over extended interactions.

Key insight: Self-configuring orthogonal memory views reduce semantic interference in long conversational agent memory.

Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design

Hongyang Du; Lan Yan; Christian Flores; Asim Kadav arXiv: 2609.22086

Figure from Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design
Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design

Professional graphic design is a long-horizon agentic task in which structured, editable artifacts emerge from many interdependent actions, yet outcomes admit no reliable programmatic oracle.

Key insight: Procedural skill memory can evolve from user traffic while a frozen frontier model drives 230+ design tools.

CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

Bowen Ye; Lei Li; Shicheng Li; Zihao Yue; … arXiv: 2609.22068

Figure from CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers.

Key insight: Agentic coding RL environments can be scaled from implemented code itself, not only issues and commits.

Authorization Revocation for Long-Running AI Agents: Root-Scoped Quiescence under Delegation and Asynchronous Execution

Genliang Zhu; Chu Wang arXiv: 2609.21284

Long-running AI agents outlive initiating processes through credentials, delegated tasks, queues, callbacks, reservations, and provider-side operations.

Key insight: Long-running agents need root-scoped authorization quiescence so revocation actually stops delegated work.

Scaling Discovery through Test-Time Communication

Jongho Park; Vasilis Kontonis; Shivam Garg; Akshay Krishnamurthy; … arXiv: 2609.21032

Figure from Scaling Discovery through Test-Time Communication
Scaling Discovery through Test-Time Communication

Science advances not in isolation but through collaboration, yet existing agentic systems capture little of this.

Key insight: Test-time multi-agent communication can beat independent parallel attempts when breakthroughs are shared.

DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement

Siyuan Liu (1 and 2); Fan Yu (1 and 2); Dongyu Ru (2); Yizhu Liu (2); … arXiv: 2609.21423

Figure from DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement
DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement

Online agent deployments produce abundant execution traces, while task-specific verification and expert annotation are costly to scale.

Key insight: Unlabeled agent traces can be distilled into evidence-grounded shortcut trees for self-refinement.

LEGIT: Credentialing Protocol for Trustworthy AI Agent Marketplaces

Steve Drew; Jiayu Zhou arXiv: 2609.21325

Agentic marketplaces are emerging where AI agents with varying capabilities autonomously complete specialized tasks for buyers.

Key insight: Agent marketplaces need credentialing that ties certification, reputation, and measurable task fit.

AI-GRACE: A Use-Case Operationalization Framework for Agentic AI: From Organizational Objectives and Obligations to Deployment Capabilities and Architecture

John Cuneo; David Chun; Gaurav Khanna arXiv: 2609.21192

Organizations deploying agentic artificial intelligence must determine more than whether a model is trustworthy; they must establish what to validate, control, and observe for a use case to deliver its intended outcome while meeting applicable obligations.

Key insight: Agentic AI deployments need a use-case operationalization bridge from objectives to controls and evidence.

Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw

Renkai Ma; Ruyuan Wan; Xuan Lu; Fan Yang; … arXiv: 2609.22067

Figure from Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw
Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw

Users increasingly delegate work to autonomous AI agents, yet evaluations typically measure task completion rather than the values users prioritize.

Key insight: Everyday OpenClaw users prioritize value groups like autonomy, dependability, and bounded agency—not only task success.

Efficient Benchmarking in Production: A Study of an Evolving LLM Agent

Yining She; Lei Lin arXiv: 2609.21267

Figure from Efficient Benchmarking in Production: A Study of an Evolving LLM Agent
Efficient Benchmarking in Production: A Study of an Evolving LLM Agent

Production LLM agents are evaluated repeatedly as they evolve, but full agent benchmarks are costly to rerun.

Key insight: Production LLM agents need cheap recurring evaluation strategies as the agent evolves.

RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models

Chuxu Song; Jiuqi Wei; Zhencan Peng arXiv: 2609.20971

Figure from RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models
RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models

Long-context large language model inference is increasingly limited by prefill, where dense self-attention processes the entire prompt before generation begins.

Key insight: Sparse long-context prefill fails when block centroids hide relevant tokens—radius-bounded selection counters mean dilution.

TinyCeNN-LM: Quality-Gated Conversion of Pretrained Attention with CeNN-Inspired Cellular-Recurrent Layers

Kabeh Mohsenzadegan; Vahid Tavakkoli; Kyandoghere Kyamakya arXiv: 2609.21139

Replacing attention in a pretrained language model is a compatibility problem: a plausible substitute may alter representations expected by later layers.

Key insight: Quality-gated conversion can replace pretrained attention with cellular-recurrent layers for local LMs.

Samsone: A Family of Open Small Audio Language Models for On-Device Inference

Piotr Masztalski; Michał K. Grzeszczyk; Olaf Sikorski arXiv: 2609.21666

The success of Large Audio Language Models has driven the development of massive multimodal networks exceeding billions of parameters.

Key insight: Open small audio language models can run full speech understanding on-device at ~134M parameters.

NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities

Jagadeesh Balam; Travis Bartley; Edresson Casanova; Sanjay Chauhan; … arXiv: 2609.21967

We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities.

Key insight: Full-duplex speech-to-speech models can expose tool calling for voice agents.

GameLogicBench: Evaluating Coding Agents on Runtime Game Logic with Tick-Level State Assertions

Xinyu Che; Yunfei Ge; Shihao Li; Yanchen Liu; … arXiv: 2609.21562

Coding agents can modify and test code across large software projects.

Key insight: Coding agents for games should be scored with tick-level rule assertions, not only final-state checks.

GameASG-Bench: Benchmarking Autonomous Software Generation for Game Development

Xiuhui Zhang; Yi Chen; Shusheng Xu; Fan Li; … arXiv: 2609.21293

Autonomous software generation (ASG) aims to turn human requirements into executable applications, but delivering these applications does not necessarily establish that their interacting components satisfy the specified behavioral requirements.

Key insight: Autonomous software generation needs an evaluation interface declared before generation for behavioral testability.

SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?

George Ma; Benjamin Mikek; Haoyu Li; Ferhat Erata; … arXiv: 2609.21190

Ensuring the correctness of LLM-generated code is a core challenge for modern software engineering.

Key insight: Formal, machine-checked proofs raise the bar beyond incomplete held-out tests for coding agents.

BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence

Chuxuan Hu; Yeye He; Penny Zhou; Wee Hyong Tok; … arXiv: 2609.20886

Business intelligence (BI) is a cornerstone of enterprise decision-making and is widely used by enterprise users in software such as Power BI and Tableau.

Key insight: End-to-end BI agents can cover table selection, transforms, joins, and answering in one workflow.

Verify, Don't Trust: Agentic Model Development for Video Discovery Retrieval at Scale

Hao Fu; Baiting Zhu; Minglei Chen; Yinjie Huang; … arXiv: 2609.21257

Large language model (LLM) agents can propose, implement, and evaluate model changes.

Key insight: Agentic model-development loops need verify-don't-trust checks against no-op diffs and evaluation leakage.

CIPL: A Channel-Aware Framework for Recoverable Privacy Leakage in LLM Agents

Tao Huang; Guosen Wu; Guolong Zheng; Jiayang Meng; … arXiv: 2609.21686

Privacy leakage in LLM agents is commonly evaluated within individual components such as memory, retrieval, or tool-use pipelines, which makes it difficult to distinguish internal exposure from information that an external observer can actually recover.

Key insight: Privacy leakage in LLM agents can be framed as recoverable channel-aware risk rather than all-or-nothing exposure.

TERMon: Detecting Persistent Behavioral Threats in Edge AI via Hardware-Native Ternary Runtime Monitor

Arish Sateesan; Edlira Dushku arXiv: 2609.21713

Edge AI accelerators are increasingly deployed in safety-critical environments, where model outputs may control physical actuators, make access-control decisions, or trigger alarms.

Key insight: Edge AI agents benefit from hardware-native runtime monitors for persistent behavioral threats.

Micro-Collaborative Poisoning: A Distributed Attack on RAG Systems

Pedro Pereira; Eva Maia; Isabel Praça arXiv: 2609.21573

Retrieval-Augmented Generation (RAG) improves large language models by grounding outputs in external knowledge sources, but this dependency also creates a surface for poisoning attacks.

Key insight: Distributed micro-collaborative poisoning is a realistic adversarial threat model for RAG-backed agents.

Bayesian Belief Layer for Controllable Opinion Dynamics in LLM Agents

Hafsa Akbar; Daniel Platnick; Marjan Alirezaie; Hossein Rahnama arXiv: 2609.21997

LLM agents in social simulation revise their opinions implicitly, in context: how open an agent is to persuasion can neither be specified nor verified, and collective outcomes inherit the model's training prior.

Key insight: A Bayesian belief layer can control opinion dynamics in LLM agents separately from token generation.

From Memory to Behavior: A Behavior-Aware Role-Playing Framework for Social Media Influencers

Ji-Lun Peng; Yi-Zhen Zhang; Chun-Nan Chou; Yun-Nung Chen arXiv: 2609.21349

Large language models have shown strong potential as role-playing agents for real individuals, yet faithful impersonating remains challenging.

Key insight: Role-play agents need situation-conditioned internal state linking memory to behavior, not static personas.

Do Personality-Tuned LLMs Make Better Social Agents?

Tim Krabbe; Xiaodan Shi arXiv: 2609.21857

LLMs are increasingly used in social simulations for socially interactive agents and robots, offering more flexibility than rule-based systems.

Key insight: Personality-tuned fine-tuning can improve consistency of social agents over instruction prompting alone.

An Agentic Just-in-Time Adaptive Intervention System for Personalized Sleep Support: Proof-of-Concept Study with N of 1 Data

Nick Rezaee; Chelsea Boccagno arXiv: 2609.21805

Background: Just-in-time adaptive interventions (JITAIs) can use behavioral data to adapt support to changing contexts, but many rely on predefined rules and manual configuration.

Key insight: An agentic JITAI on Home Assistant can adapt personal sleep interventions with human-reviewable decisions.

One Prompt Does Not Fit All: Self-Meta-Evolve for Personalized Information Extraction

Hongliang Li; Lu Wang; Yong Xu; Hanyang Chen; … arXiv: 2609.21626

Large language models (LLMs) are increasingly deployed for enterprise information extraction (IE), where the same document must be reorganized differently for each user.

Key insight: Enterprise IE needs per-user prompt evolution rather than one globally optimized prompt.

What Should We Ask Next? Retrieval-Aware Question Learning under Partial Evidence

Lyucheng Qian; John Yuehan Zhang; Pingyu Wang arXiv: 2609.21924

Interactive retrieval under partial evidence is a sequential information-acquisition problem: an agent must decide which question will create the most useful evidence for the next retrieval update.

Key insight: Interactive retrieval agents should learn which next question creates the most useful evidence under partial context.

AutoRecLab: Describe the Experiment, Get the Code!

Moritz Baumgart; Philipp Meister; Justus Krell; Michael Schmidt; … arXiv: 2609.21863

Empirical evaluation is central to recommender-systems (RecSys) research, but turning experimental designs into executable code remains a manual and error-prone task.

Key insight: Natural-language experiment ideas can drive an autonomous lab that builds and validates executable code.

Can Agents Design Better Chips with a Higher Level Abstraction?

Zijian Ding; Yang Zou; Yizhou Sun; Jason Cong arXiv: 2609.21157

Large Language Model (LLM) agents are increasingly being explored for chip design, but most existing approaches operate directly at RTL.

Key insight: Chip-design agents perform better when they operate at higher abstractions (HLS) before RTL refinement.

AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents

Seoyeon An; Hyeonseo Jang; Minsu Kim; Chanho Lee; … arXiv: 2609.21386

Comprehensive video understanding is crucial for advancing artificial intelligence toward the intricate dynamics of the physical world.

Key insight: MLLM agents need multi-hop video benchmarks that require multi-step inference, not only scene summaries.

GUARD: Natural Forgetting in Large Reasoning Models via Guided Answer-Reasoning Distillation

Zeyu Yan; Guanghao Zhou; Minghui Qiu; Ming Gao; … arXiv: 2609.21677

Recent advances in large reasoning models (LRMs) have made machine unlearning more challenging, as protected facts or unsafe rationales may surface in intermediate chain-of-thought (CoT) traces before the final answer is produced.

Key insight: Unlearning for reasoning models must shape post-forgetting CoT trajectories, not only final answers.