Monday’s cs.AI announcement day (2026-09-21) lists 49 new and 76 cross-lists (replacements skipped; listing total 125). Stack filter for agent systems, memory/context, computer-use / GUI / tools / MCP / skills / harnesses, multi-agent, persistence/identity, and local/open serving keeps 32 papers — conversational long-term memory (AutoViewMem), procedural skill memory for tool-heavy agents (Designer-RSI), agentic coding RL (CodeMidas), long-running authorization quiescence, and multi-agent test-time communication.
Zijie Cao; Xijun Qu; Zhicheng Gu; Xiaoshu Chen; … arXiv: 2609.21940
Long-term memory is essential for large language model (LLM) agents to maintain consistency and personalization over extended interactions.
Key insight: Self-configuring orthogonal memory views reduce semantic interference in long conversational agent memory.
Hongyang Du; Lan Yan; Christian Flores; Asim Kadav arXiv: 2609.22086
Professional graphic design is a long-horizon agentic task in which structured, editable artifacts emerge from many interdependent actions, yet outcomes admit no reliable programmatic oracle.
Key insight: Procedural skill memory can evolve from user traffic while a frozen frontier model drives 230+ design tools.
Bowen Ye; Lei Li; Shicheng Li; Zihao Yue; … arXiv: 2609.22068
Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers.
Key insight: Agentic coding RL environments can be scaled from implemented code itself, not only issues and commits.
Genliang Zhu; Chu Wang arXiv: 2609.21284
Long-running AI agents outlive initiating processes through credentials, delegated tasks, queues, callbacks, reservations, and provider-side operations.
Key insight: Long-running agents need root-scoped authorization quiescence so revocation actually stops delegated work.
Jongho Park; Vasilis Kontonis; Shivam Garg; Akshay Krishnamurthy; … arXiv: 2609.21032
Science advances not in isolation but through collaboration, yet existing agentic systems capture little of this.
Key insight: Test-time multi-agent communication can beat independent parallel attempts when breakthroughs are shared.
Siyuan Liu (1 and 2); Fan Yu (1 and 2); Dongyu Ru (2); Yizhu Liu (2); … arXiv: 2609.21423
Online agent deployments produce abundant execution traces, while task-specific verification and expert annotation are costly to scale.
Key insight: Unlabeled agent traces can be distilled into evidence-grounded shortcut trees for self-refinement.
Steve Drew; Jiayu Zhou arXiv: 2609.21325
Agentic marketplaces are emerging where AI agents with varying capabilities autonomously complete specialized tasks for buyers.
Key insight: Agent marketplaces need credentialing that ties certification, reputation, and measurable task fit.
John Cuneo; David Chun; Gaurav Khanna arXiv: 2609.21192
Organizations deploying agentic artificial intelligence must determine more than whether a model is trustworthy; they must establish what to validate, control, and observe for a use case to deliver its intended outcome while meeting applicable obligations.
Key insight: Agentic AI deployments need a use-case operationalization bridge from objectives to controls and evidence.
Renkai Ma; Ruyuan Wan; Xuan Lu; Fan Yang; … arXiv: 2609.22067
Users increasingly delegate work to autonomous AI agents, yet evaluations typically measure task completion rather than the values users prioritize.
Key insight: Everyday OpenClaw users prioritize value groups like autonomy, dependability, and bounded agency—not only task success.
Yining She; Lei Lin arXiv: 2609.21267
Production LLM agents are evaluated repeatedly as they evolve, but full agent benchmarks are costly to rerun.
Key insight: Production LLM agents need cheap recurring evaluation strategies as the agent evolves.
Chuxu Song; Jiuqi Wei; Zhencan Peng arXiv: 2609.20971
Long-context large language model inference is increasingly limited by prefill, where dense self-attention processes the entire prompt before generation begins.
Key insight: Sparse long-context prefill fails when block centroids hide relevant tokens—radius-bounded selection counters mean dilution.
Kabeh Mohsenzadegan; Vahid Tavakkoli; Kyandoghere Kyamakya arXiv: 2609.21139
Replacing attention in a pretrained language model is a compatibility problem: a plausible substitute may alter representations expected by later layers.
Key insight: Quality-gated conversion can replace pretrained attention with cellular-recurrent layers for local LMs.
Piotr Masztalski; Michał K. Grzeszczyk; Olaf Sikorski arXiv: 2609.21666
The success of Large Audio Language Models has driven the development of massive multimodal networks exceeding billions of parameters.
Key insight: Open small audio language models can run full speech understanding on-device at ~134M parameters.
Jagadeesh Balam; Travis Bartley; Edresson Casanova; Sanjay Chauhan; … arXiv: 2609.21967
We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities.
Key insight: Full-duplex speech-to-speech models can expose tool calling for voice agents.
Xinyu Che; Yunfei Ge; Shihao Li; Yanchen Liu; … arXiv: 2609.21562
Coding agents can modify and test code across large software projects.
Key insight: Coding agents for games should be scored with tick-level rule assertions, not only final-state checks.
Xiuhui Zhang; Yi Chen; Shusheng Xu; Fan Li; … arXiv: 2609.21293
Autonomous software generation (ASG) aims to turn human requirements into executable applications, but delivering these applications does not necessarily establish that their interacting components satisfy the specified behavioral requirements.
Key insight: Autonomous software generation needs an evaluation interface declared before generation for behavioral testability.
George Ma; Benjamin Mikek; Haoyu Li; Ferhat Erata; … arXiv: 2609.21190
Ensuring the correctness of LLM-generated code is a core challenge for modern software engineering.
Key insight: Formal, machine-checked proofs raise the bar beyond incomplete held-out tests for coding agents.
Chuxuan Hu; Yeye He; Penny Zhou; Wee Hyong Tok; … arXiv: 2609.20886
Business intelligence (BI) is a cornerstone of enterprise decision-making and is widely used by enterprise users in software such as Power BI and Tableau.
Key insight: End-to-end BI agents can cover table selection, transforms, joins, and answering in one workflow.
Hao Fu; Baiting Zhu; Minglei Chen; Yinjie Huang; … arXiv: 2609.21257
Large language model (LLM) agents can propose, implement, and evaluate model changes.
Key insight: Agentic model-development loops need verify-don't-trust checks against no-op diffs and evaluation leakage.
Tao Huang; Guosen Wu; Guolong Zheng; Jiayang Meng; … arXiv: 2609.21686
Privacy leakage in LLM agents is commonly evaluated within individual components such as memory, retrieval, or tool-use pipelines, which makes it difficult to distinguish internal exposure from information that an external observer can actually recover.
Key insight: Privacy leakage in LLM agents can be framed as recoverable channel-aware risk rather than all-or-nothing exposure.
Arish Sateesan; Edlira Dushku arXiv: 2609.21713
Edge AI accelerators are increasingly deployed in safety-critical environments, where model outputs may control physical actuators, make access-control decisions, or trigger alarms.
Key insight: Edge AI agents benefit from hardware-native runtime monitors for persistent behavioral threats.
Pedro Pereira; Eva Maia; Isabel Praça arXiv: 2609.21573
Retrieval-Augmented Generation (RAG) improves large language models by grounding outputs in external knowledge sources, but this dependency also creates a surface for poisoning attacks.
Key insight: Distributed micro-collaborative poisoning is a realistic adversarial threat model for RAG-backed agents.
Hafsa Akbar; Daniel Platnick; Marjan Alirezaie; Hossein Rahnama arXiv: 2609.21997
LLM agents in social simulation revise their opinions implicitly, in context: how open an agent is to persuasion can neither be specified nor verified, and collective outcomes inherit the model's training prior.
Key insight: A Bayesian belief layer can control opinion dynamics in LLM agents separately from token generation.
Ji-Lun Peng; Yi-Zhen Zhang; Chun-Nan Chou; Yun-Nung Chen arXiv: 2609.21349
Large language models have shown strong potential as role-playing agents for real individuals, yet faithful impersonating remains challenging.
Key insight: Role-play agents need situation-conditioned internal state linking memory to behavior, not static personas.
Tim Krabbe; Xiaodan Shi arXiv: 2609.21857
LLMs are increasingly used in social simulations for socially interactive agents and robots, offering more flexibility than rule-based systems.
Key insight: Personality-tuned fine-tuning can improve consistency of social agents over instruction prompting alone.
Nick Rezaee; Chelsea Boccagno arXiv: 2609.21805
Background: Just-in-time adaptive interventions (JITAIs) can use behavioral data to adapt support to changing contexts, but many rely on predefined rules and manual configuration.
Key insight: An agentic JITAI on Home Assistant can adapt personal sleep interventions with human-reviewable decisions.
Hongliang Li; Lu Wang; Yong Xu; Hanyang Chen; … arXiv: 2609.21626
Large language models (LLMs) are increasingly deployed for enterprise information extraction (IE), where the same document must be reorganized differently for each user.
Key insight: Enterprise IE needs per-user prompt evolution rather than one globally optimized prompt.
Lyucheng Qian; John Yuehan Zhang; Pingyu Wang arXiv: 2609.21924
Interactive retrieval under partial evidence is a sequential information-acquisition problem: an agent must decide which question will create the most useful evidence for the next retrieval update.
Key insight: Interactive retrieval agents should learn which next question creates the most useful evidence under partial context.
Moritz Baumgart; Philipp Meister; Justus Krell; Michael Schmidt; … arXiv: 2609.21863
Empirical evaluation is central to recommender-systems (RecSys) research, but turning experimental designs into executable code remains a manual and error-prone task.
Key insight: Natural-language experiment ideas can drive an autonomous lab that builds and validates executable code.
Zijian Ding; Yang Zou; Yizhou Sun; Jason Cong arXiv: 2609.21157
Large Language Model (LLM) agents are increasingly being explored for chip design, but most existing approaches operate directly at RTL.
Key insight: Chip-design agents perform better when they operate at higher abstractions (HLS) before RTL refinement.
Seoyeon An; Hyeonseo Jang; Minsu Kim; Chanho Lee; … arXiv: 2609.21386
Comprehensive video understanding is crucial for advancing artificial intelligence toward the intricate dynamics of the physical world.
Key insight: MLLM agents need multi-hop video benchmarks that require multi-step inference, not only scene summaries.
Zeyu Yan; Guanghao Zhou; Minghui Qiu; Ming Gao; … arXiv: 2609.21677
Recent advances in large reasoning models (LRMs) have made machine unlearning more challenging, as protected facts or unsafe rationales may surface in intermediate chain-of-thought (CoT) traces before the final answer is produced.
Key insight: Unlearning for reasoning models must shape post-forgetting CoT trajectories, not only final answers.