Backfill for a missed day. Friday’s cs.AI announcement (Fri 18 Sep 2026) lists 89 new and 128 cross-lists (replacements skipped; listing total 217). The main pass over this announcement already ran as the September 19 digest (37 papers), so this page collects the 12 remaining papers that pass the agent, memory, tools, and multi-agent filter — event-level privacy mediation for memory writes and agent messages, inference-engine fingerprinting by the model itself, harness-enforced safety constraints for coding agents, gated rewards for tool learning, skill distillation for agentic RL, and covert coordination between isolated model instances.
Tao Huang; Guosen Wu; Chen Hou; Guolong Zheng arXiv: 2609.19226
AI-mediated platforms coordinate work through LLM agents acting for different principals. In these workflows, privacy loss can be created before a final answer appears: a memory write, shared-workspace update, inter-agent message, or tool event may impose downstream exposure cost on another principal.
PAPC models privacy loss created before any final answer — a memory write, shared-workspace update, inter-agent message or tool event — as a privacy-propagation externality depending on topology and fanout. A platform mediator intercepts information-moving events and can allow, release a policy-safe abstraction, quarantine raw content, block, or narrow onward rights; on retrieval-memory and multi-agent benchmarks it keeps task completion and eliminates measured raw-value exposure.
Key insight: Privacy in multi-agent workflows has to be enforced at the level of individual information-moving events — memory writes, messages, tool calls — not just the final answer.
Shihao Liu; Hao Yin; Lijun Liu; Zhengzong Chen; … arXiv: 2609.20082
Tool learning enables large language models (LLMs) to use external tools for tasks beyond parametric knowledge. Reinforcement learning can optimize tool-call behavior from feedback, but current methods still face two problems: fixed-threshold curricula can become misaligned with the policy's evolving capability boundary, and additive rewards can leak argument-level credit when the predicted tool is wrong.
MATCH: model-aware curriculum near the policy's capability boundary plus Hierarchical Tool-call Gated Reward (tool name → argument key → value, credit only when prerequisites hold). 72.19% on API-Bank and 62.87% on BFCL V3, consistent across four backbones.
Key insight: Tool-learning rewards work better when credit for arguments is gated on choosing the right tool, combined with a curriculum that tracks the model's capability.
Sarah Radway; Andrew Cheng; Vijay Janapa Reddi; James Mickens arXiv: 2609.20614
Frontier AI models are rapidly gaining the ability to exploit vulnerabilities in complex pieces of software. The risk is not theoretical, as evidenced by recent sandbox escapes performed by frontier models at OpenAI and Anthropic.
Shows a misaligned model can fingerprint which inference engine runs it (examples for five popular engines, e.g. vLLM, SGLang) using realistic agentic harnesses, then use engine-specific exploits via crafted output tokens alone; includes a proof-of-concept exploit chain to bare metal.
Key insight: A model can fingerprint the inference engine running it and attempt engine-specific exploits using only its output tokens, putting the engine inside the agent's attack surface.
Bingxin Xu; Yuzhang Shang; Zhen Dong; Emilio Ferrara arXiv: 2609.20822
Coding agents have emerged as a promising paradigm for robot manipulation: a language model writes the robot controller as a program, and agents built in this way now operate robots without robot-specific training.Whether this paradigm is also safe, however, has not been asked. We evaluate coding agent under a safety constraint, where each task pairs a manipulation goal with an obstacle the robot must not touch.
Coding agents writing robot controllers collide with a forbidden obstacle in most cases even though their traces mention it and the prompt forbids it; SafeHarness adds obstacle-aware route planning and contact harnesses so the constraint is enforced in planning.
Key insight: Telling a coding agent about a safety constraint is not enough; the harness has to make the constraint part of planning.
Juzheng Zhang; Disha Makhija; Manoj Ghuhan Arivazhagan; Vinayshekhar Bannihatti Kumar; … arXiv: 2609.20715
Agent trajectories record what an agent does and what happens next. Yet standard supervised fine-tuning (SFT) applies loss only to agent-authored action tokens, using environment observations as context but not as prediction targets.
ActObs also supervises environment observation tokens during SFT. After GRPO on Qwen3-4B it beats action-only at every pass@k on Terminal-Bench 2.0; on Qwen3-8B +3.4 pp pass@16; +4.2 pp pass@1 on aider-polyglot at 4B.
Key insight: Supervising observation tokens during SFT gives agents a better starting point for RL without extra data or compute.
Yan Yu; Zhengxi Lu; Yizhou Liu; Yichen Pan; … arXiv: 2609.20784
Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from a self-teacher with privileged task skills, letting a skill-free student internalize them. This recipe, however, is undermined by two findings in agentic tasks: privileged information alone does not always make a teacher reliable, and the benefit of teacher supervision is stage-dependent.
RetireOPD trains a skill-conditioned teacher with environment rewards, then a skill-free student with RL + on-policy distillation that retires the teacher adaptively. Qwen2.5 1.5B–7B: +14.1–18.8% ALFWorld and +11.8–19.0% WebShop over RL; student surpasses its teacher.
Key insight: A skill-conditioned teacher can be distilled into a skill-free student that retires the teacher on its own and ends up surpassing it.
Yingxuan Zhuang; Binhe Yu; Jingxiao Yang; Ruopei Sun; … arXiv: 2609.19830
Reinforcement learning for LLM agents involves two distinct optimization di- mensions: how environment feedback is exploited within a trajectory, and how complete trajectories are aggregated across a batch. We formulate these dimen- sions as Intra-Trajectory Feedback Attribution and Inter-Trajectory Objec- tive Aggregation, and introduce BATON (Bayesian Attribution and Trajectory Objective Normalization), a dual-axis policy optimization framework.
BATON: Bayesian Feedback Attribution within trajectories plus Trajectory Mass Normalization across trajectories; independent gains with GRPO and GiGPO on ALFWorld, WebShop, SearchQA.
Key insight: Agent RL has two separable optimization axes — within-trajectory credit and across-trajectory aggregation — and improving both compounds.
Alexander Shirnin; Aleksey Kudelya arXiv: 2609.19504
As model-generated content is increasingly consumed by other model instances in automated workflows, a practically important question arises: can a model embed a signal in natural language that an independent instance of the same model can detect, relying only on shared pre-training and task instructions, without any shared memory or coordination-specific training? We introduce For Your Eyes Only, a cooperative signalling game designed to evaluate this directly.
For Your Eyes Only: can a model embed a signal that an isolated instance of the same model detects, with no shared memory? Most models struggle once detectable signals must be avoided, but one frontier model stays near-perfect; models can also misdirect; coordination is weaker across architectures.
Key insight: Isolated instances of the same model can sometimes coordinate through natural language alone, which matters for monitor/worker setups built from one model.
Tengfei Shao arXiv: 2609.20543
Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final states are operationalized. We replayed 100 held-out human Wason groups with matched large language model (LLM) agent groups, seeding one belief-anchored agent per participant's pre-discussion answer and scoring agents and people with the same code.
Replaying 100 human Wason-task groups with matched LLM agent groups: agents almost always post while ~1/5 of humans never did, and agent groups are 34–44 pp more consensual in sensitivity analyses.
Key insight: LLM agent groups replaying human deliberation reach consensus far more often than the humans did, so agent agreement is a poor proxy for collective reasoning.
Jian Gao; Hang Jiang arXiv: 2609.19530
Hiring is bilateral: employers assess fit, while candidates present and defend evidence of their qualifications. Yet résumé screening, the first gate, is commonly automated as a static, one-call judgment over a résumé-job pair.
Two-agent résumé screening (employer-side and candidate-side agents exchange evidence) vs one-call judgment on 600 constructed pairs: two-agent screening advances more applications (GPT-5.5 33.3%→39.3%; Opus 4.7 34.0%→35.5%).
Key insight: Letting employer-side and candidate-side agents exchange evidence changes screening outcomes compared with a single one-shot judgment.
Ruiyang Wang; Hao-Lun Hsu; Swarajh Mehta; Jiwoo Kim; … arXiv: 2609.19315
Large language models (LLMs) provide a flexible interface for long-horizon robot planning, but generated plans often fail to respect embodiment constraints, recover from planning errors, or reason effectively under partial observability. We present GAVEL, a framework for verifying and repairing long-horizon LLM planning built around an explicit graph world model.
GAVEL verifies and repairs LLM plans against an explicit graph world model (object relations, preconditions/effects, beliefs over unobserved locations), reserving LLM replanning for semantic errors. On BEHAVIOR-1K with Qwen3-8B, single-task success 41.2%→91.8%, multi-task 19.9%→92.6%.
Key insight: An explicit world model that checks and repairs plans before execution can dramatically raise long-horizon task success even for compact models.
Canfer Akbulut; Justine Breuch; Arianna Manzini; Lujain Ibrahim; … arXiv: 2609.20077
Interest in developing personalised language models is rapidly growing. While personalisation is often viewed as a mechanism to better serve diverse user needs, the effects of sustained interactions with personalised models on people's perception of and behaviour toward AI remain poorly understood.
Five-day study with 992 participants comparing non-personalised, memory-based and survey-based personalisation: memory-based users self-disclosed more and rated the model less creepy; survey-based users reported more regret about sharing personal information.
Key insight: How a model learns about a user changes how the user feels about sharing: memory-based personalization felt less creepy than an up-front intake survey.