Monday’s cs.AI announcement day lists 69 new and 98 cross-lists (replacements skipped; listing total 167). Stack filter for agent systems, memory/context, computer-use / GUI / tools / skills / harnesses, multi-agent, persistence/identity, and local/open models keeps 27 papers — a full readable day after a quiet weekend.


Research Papers

Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work

Chen; Wenhui; Cheng; Shiwen; Dong; Hao; … arXiv: 2609.11977

Figure from Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work
Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work

Co-work agents execute complex workflows that combine information gathering, tool use, coding, and file manipulation across many model invocations. Because cost and latency accumulate over the full episode, their practical value depends not only on peak capability but also on how efficiently that capability is delivered.

Key insight: Open 35B MoE co-work model: optimize harness-trained execution efficiency, not only peak reasoning.


Harness or Model? Isolating the Harness Effect in Agentic Coding with a Contamination-Controlled Private Suite

Arjmandi; Mohsen arXiv: 2609.11987

An agentic coding system couples a language model to a harness: the tools, prompts and control flow that turn a chat model into an autonomous software engineer. Vendors ship harnesses tuned to their own models, and practitioners assume the vendor-native pairing solves more tasks.

Key insight: Same model, different harness: measure harness contribution with contamination-controlled paired contrasts.


Can LLMs in Draft-Verify-Revise Pipelines Resolve Deictic Ambiguity?

Ekekezie; Obinna I. arXiv: 2609.12162

Draft-verify-revise is a common LLM orchestration pattern for scaling inference-time compute. One LLM drafts, a second critiques the draft and provides feedback, and a third uses that feedback to revise the draft into the final output.

Key insight: Draft-verify-revise multi-LLM pipelines can silently shift deixis across stages — track referent continuity.


AIM: A Privacy-Aware Interoperable Memory Framework for Multi-Agent Multi-User LLM Systems

Johnson; Zachary; Kumankumah; Nigel Boachie; Chatterjee; Somya; … arXiv: 2609.12320

Figure from AIM: A Privacy-Aware Interoperable Memory Framework for Multi-Agent Multi-User LLM Systems
AIM: A Privacy-Aware Interoperable Memory Framework for Multi-Agent Multi-User LLM Systems

Traditional large language models (LLMs) are scoped to individual user sessions, limiting their knowledge to a single conversation and preventing them from learning user preferences that evolve over time. Existing agentic memory systems address this limitation but generally operate at the individual-user level, restricting the public knowledge that could be shared across users to improve downstream responses.

Key insight: Multi-user multi-agent memory needs interoperable + privacy-aware sharing, not only per-session stores.


Affective Agent: On-Device Personalized Intervention Reasoning for Wearable Systems

Mun; Reina; Wan; Zishen; Reddi; Vijay Janapa arXiv: 2609.12322

Figure from Affective Agent: On-Device Personalized Intervention Reasoning for Wearable Systems
Affective Agent: On-Device Personalized Intervention Reasoning for Wearable Systems

Affective computing has advanced wearable state inference, but on-device reasoning about whether, when, and how to intervene remains challenging. We present Affective Agent, a three-layer reference architecture for personalized intervention reasoning under uncertainty on wearable-class hardware.

Key insight: On-device wearable agent: intervene whether/when/how under uncertainty with a sub-billion LM.


Toward Robust Personalized Alignment for LLMs: Mitigating Persona Drift in Multi-Turn Dialogue

Zhang; Youyuan; Li; Siyuan; Liu; Fangming; … arXiv: 2609.12373

Persona drift remains a central challenge for personalized language models, as user profiles evolve over long interactions rather than remain permanently fixed. Models must therefore revise persistent persona states when preferences genuinely change, while avoiding updates driven by transient, ambiguous, or unresolved observations.

Key insight: Separate turn-local evidence from persistent persona revision to fight multi-turn persona drift.


BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI Agents

Ye; Tong; Han; Kunyang; Wang; Guozhi; … arXiv: 2609.12394

Figure from BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI Agents
BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI Agents

Mobile GUI agents are shifting from multi-module frameworks to native models trained end-to-end, yet industrial deployment faces three persistent gaps. Sandbox training produces a distribution mismatch with production environments; expensive real-device failures remain underutilized; and fixed benchmarks saturate, losing the power to guide iteration.

Key insight: Mobile GUI agents need a real-device flywheel; sandbox-only training mismatches production.


VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgets

Bai; Yu; Miao; Yukai; Wang; Dawei; … arXiv: 2609.12404

Learning from trial and error is a promising way to improve language agents on complex tasks such as computer control. Reflexion introduced verbal reinforcement learning, which turns failed trials into text that guides later attempts without updating model parameters.

Key insight: Evaluate verbal RL / trial-and-error agents under finite trial budgets, not unlimited Reflexion loops.


SoK: Rethinking Jailbreaking in the Era of Agentic AI: Attacks, Defenses, and Practical Consideration

Mia; Md Jueal; Wu; Yanzhao; Uluagac; Selcuk; … arXiv: 2609.12413

Figure from SoK: Rethinking Jailbreaking in the Era of Agentic AI: Attacks, Defenses, and Practical Consideration
SoK: Rethinking Jailbreaking in the Era of Agentic AI: Attacks, Defenses, and Practical Consideration

Large language models (LLMs) are rapidly evolving from conversational assistants into agentic AI systems that reason, plan, invoke tools, maintain persistent memory, communicate with other agents, and execute multi-step tasks. At the same time, modern models exhibit substantially stronger native safety alignment than earlier generations on which many jailbreak attacks and defenses were originally studied.

Key insight: Jailbreak literature must be re-scoped for tool-using, memory-keeping agentic systems.


LifeFuse-Mem: Lifecycle-Aware State Fusion Against Temporary Overwriting for Long-Term Memory

Zhao; Hanyu; Feng; Yuqian; Song; Zhenyu; … arXiv: 2609.12436

Figure from LifeFuse-Mem: Lifecycle-Aware State Fusion Against Temporary Overwriting for Long-Term Memory
LifeFuse-Mem: Lifecycle-Aware State Fusion Against Temporary Overwriting for Long-Term Memory

Long-running LLM agents require memory mechanisms that maintain coherent internal states across interactions. We study a lifecycle-labeled memory setting in which write episodes provide lifecycle metadata during training, and phase-aware readout is used during evaluation.

Key insight: Lifecycle-label memory writes so temporary context cannot overwrite long-term state.


Beyond Vector Similarity: Hierarchical Context-Aware Graph RAG vs Standard RAG in Enterprise Code Migration

Jaiswal; Nilesh; Agrawal; Aniket; Shukla; Arjit; … arXiv: 2609.12464

As enterprises modernize legacy monolithic systems to microservices, Large Language Models (LLMs) are heavily utilized for automated code translation. However, traditional vector-based Retrieval-Augmented Generation (Standard RAG) struggles to capture topological relationships.

Key insight: As enterprises modernize legacy monolithic systems to microservices, Large Language Models (LLMs) are heavily utilized for automated code translation.


When Does AI Augment Work? A Workflow-Level Framework for Human-Agent Collaboration

AI Collaboration; Wu; Jiaying; Ziems; Caleb; Chan; … arXiv: 2609.12482

We aim to characterise the value of artificial intelligence in the workplace. Current studies largely measure this value in terms of the current automation capabilities and public adoption of AI.

Key insight: We aim to characterise the value of artificial intelligence in the workplace.


From Collaboration to Capability: Internalizing Routed LLM Experts into Compact Reasoners

Nie; Frank; Wang; Shuyao; Liu; Ethan B. arXiv: 2609.12578

Figure from From Collaboration to Capability: Internalizing Routed LLM Experts into Compact Reasoners
From Collaboration to Capability: Internalizing Routed LLM Experts into Compact Reasoners

A compact controller can coordinate stronger experts by selecting whom to consult, formulating requests, and integrating their responses. We study whether learning from both the controller's decisions and the experts' reasoning and code improves its generation after expert removal.

Key insight: Internalize routed-expert collaboration into a compact controller so capability survives expert removal.


Reproducing and Evaluating the Generalizability of Subliminal Learning in Open-Weight Models

van der Weijden; Daan; Brack; Nathan; Santamaria; Selene Baez arXiv: 2609.12586

In this reproduction paper we investigate subliminal learning, a consequence of distillation where teacher models transmit behavioral preference traits through semantically unrelated data. The original paper explores two types of traits (animal preferences and misalignment), three data modalities (number sequences, code, and chain of thought), and several model families.

Key insight: In this reproduction paper we investigate subliminal learning, a consequence of distillation where teacher models transmit behavioral preference traits through semantically unrelated data.


Residual Vector-based Reconstruction as Long-Context Recall Regardless of Context Window Size

Ryu; MyungHoon; Piao; XinYu; Kim; Jong-Kook arXiv: 2609.12686

Figure from Residual Vector-based Reconstruction as Long-Context Recall Regardless of Context Window Size
Residual Vector-based Reconstruction as Long-Context Recall Regardless of Context Window Size

Large language models (LLMs) process long contexts, including long documents and lengthy conversations, but face token-level memory usage that increases proportionally to input length. Although model optimization and lossy prompt compression are widely used, these methods still fail to solve the long-context recall problem beyond pretrained and size-constrained context windows.

Key insight: Large language models (LLMs) process long contexts, including long documents and lengthy conversations, but face token-level memory usage that increases proportionally to input length.


Skill Issue: Lessons from Optimizing Repository SKILLs for Coding Agents

Kozyrev; Mykhailo; Kozyrev; Andrei; Podkopaev; Anton arXiv: 2609.12742

Coding agents increasingly read repository knowledge from SKILLs --- plain \texttt{.md} files versioned alongside the code. Recent work synthesizes these files automatically, by optimizing the document against a benchmark.

Key insight: Treat repository SKILL.md files as optimizable artifacts for coding agents, with care about overfitting.


What Drives Recovery in Agentic Text-to-Cypher? LAST-CQ: An LLM Agent Self-Refinement Framework

Prokopiou; Ioannis; Aidinis; Athanasios; Kyrmpatsos; Panagiotis-Christos; … arXiv: 2609.12746

Agentic pipelines for structured-query generation are rapidly expanding, but it is unclear which part of the loop produces the gain. We use LAST-CQ -- a five-agent, training-free, execution-grounded Text-to-Cypher framework -- as an instrumented testbed, running three counterfactuals over 2,471 live-database queries and six backbones spanning three vendor scale tiers.

Key insight: Agentic pipelines for structured-query generation are rapidly expanding, but it is unclear which part of the loop produces the gain.


K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments

Yu; Guangsheng; Jiang; Yanna; Wang; Qin; … arXiv: 2609.12808

Unlearning benchmarks such as TOFU and MUSE certify forgetting by reading the model's final answer, where a model that refuses to answer already counts as having forgotten. We show that this model-level certificate does not transfer once the model is deployed as an agent.

Key insight: Unlearning benchmarks such as TOFU and MUSE certify forgetting by reading the model's final answer, where a model that refuses to answer already counts as having forgotten.


SoulAuth: An Actor-native Identity Architecture and Rust Reference Implementation for Humans and Long-lived AI Actors

Yuan; Kun; Wang; Harold; Li; Echo; … arXiv: 2609.11258

Figure from SoulAuth: An Actor-native Identity Architecture and Rust Reference Implementation for Humans and Long-lived AI Actors
SoulAuth: An Actor-native Identity Architecture and Rust Reference Implementation for Humans and Long-lived AI Actors

As AI systems move from transient model invocations toward long-lived actors that persist across credentials, clients, sessions, and runtime instances, identity infrastructure must answer a basic question: where should the canonical continuity boundary be placed? This paper introduces Actor-native Identity and presents SoulAuth, an open-source Rust reference implementation for Humans and long-lived AIActors.

Key insight: Actor-native identity (SoulAuth) is a concrete persistence layer for humans and long-lived agents.


When Agent Metrics Measure Different Things: An Evidence-Grounded Audit of the Praxa AI Pipeline

Creadore; Stefan G.; Woakz; Peyton arXiv: 2609.12017

Agent evaluations can be numerically correct while measuring a different construct from the one implied by their labels. We present a retrospective measurement audit of selected Praxa AI implementation files, historical evaluation artifacts, and operational records.

Key insight: Agent evaluations can be numerically correct while measuring a different construct from the one implied by their labels.


Reality Is the Final Verifier: On Two Key Gaps in Agentic Software Engineering

Krentsel; Alexander; Agarwal; Shubham; Cemri; Mert; … arXiv: 2609.12039

Software development follows an implementation-verification loop in which developers or agents iteratively revise an implementation until an evaluator, such as a test suite, accepts it. The evaluator checks the implementation against a set of requirements under a model of the deployment environment.

Key insight: Agentic software engineering gaps: reality verification and measurement that actually match the pipeline.


Adaptive Agent Design

Velicheti; Raj Kiriti; Bose; Subhonmesh; Başar; Tamer arXiv: 2609.12486

We consider an agent acting against a general non-Markovian environment. The agent maintains its agent states, but is free to choose a transition kernel across those states and optimize its state-feedback control policies.

Key insight: We consider an agent acting against a general non-Markovian environment.


Evaluating Context Segmentation in Locally Deployable SLMs for Cybersecurity CTF Tasks

Nordio; Sebastiano; Lotto; Michele arXiv: 2609.12839

The proliferation of highly capable open-weight Small Language Models (SLMs) democratizes access to advanced cybersecurity capabilities, posing a escalating risk as these models can bypass proprietary API guardrails when deployed locally. However, SLMs deployed as autonomous agents often struggle with long-horizon, exploratory tasks like cybersecurity Capture The Flag (CTF) challenges due to context bloat and cognitive degradation from accumulated tool-call outputs.

Key insight: Local SLM agents need context segmentation against long-horizon CTF-style context bloat.


Behavior Quotient Learning for Low-Rank Adaptation of LLM Agents

Zhou; Pengyang; Tu; Xiaobin; Liu; Zhengxi; … arXiv: 2609.12896

LLM-based agents rely on heterogeneous interaction capabilities to accomplish complex tasks. Existing approaches often distribute these capabilities across multiple LoRA adapters, which increases adapter storage requirements and introduces routing overhead during inference.

Key insight: LLM-based agents rely on heterogeneous interaction capabilities to accomplish complex tasks.


SeqMoE: Toward Full-Load Performance via Predictive and Graph-Compatible MoE Offloading

Wang; Zihan; Wang; Yuqi; Gong; Lei; … arXiv: 2609.12978

Mixture-of-Experts (MoE) creates a structural advantage for offloading: only a small fraction of activated experts need to reside in device memory, and if they can be loaded in time for computation, offloading can in principle approach full-load performance, where all model weights reside in device memory. Yet translating MoE's structural advantage into practical offloading gains remains challenging.

Key insight: SeqMoE: MoE offloading that aims at full-load performance — relevant for 64GB local stacks.


Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models

Soni; Utkarsh; Murtaza; Syed Shariyar; Nie; Yifan; … arXiv: 2609.13005

Large language models (LLMs) have achieved strong performance on a wide range of natural language tasks, and recent benchmarks suggest that they are increasingly adept at multi-hop reasoning. However, these benchmarks are typically short-horizon, requiring only a small number of retrieval or inference steps, and provide limited evidence of reliability on real-world tasks that involve following manuals spanning hundreds of pages with complex, interdependent guidelines.

Key insight: Large language models (LLMs) have achieved strong performance on a wide range of natural language tasks, and recent benchmarks suggest that they are increasingly adept at multi-hop reasoning.


MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant

Shih; Yi-Jen; Kuan; Shih-Yun Shan; Lin; Guan-Ting; … arXiv: 2609.13076

Conversational voice agents have advanced significantly, offering increasingly natural human-machine interactions through both cascaded and end-to-end architectures. However, while recent benchmarks extensively evaluate dyadic interactions and passive audio comprehension, they largely overlook a prevalent real-world scenario: multi-party conversations.

Key insight: Conversational voice agents have advanced significantly, offering increasingly natural human-machine interactions through both cascaded and end-to-end architectures.