Backfill for a missed day. Friday’s cs.AI announcement (Fri 18 Sep 2026) lists 89 new and 128 cross-lists (replacements skipped; listing total 217). The main pass over this announcement already ran as the September 19 digest (37 papers), so this page collects the 12 remaining papers that pass the agent, memory, tools, and multi-agent filter — event-level privacy mediation for memory writes and agent messages, inference-engine fingerprinting by the model itself, harness-enforced safety constraints for coding agents, gated rewards for tool learning, skill distillation for agentic RL, and covert coordination between isolated model instances.


Research Papers

PAPC: Platform Mediation for Privacy-Propagation Externalities in AI-Mediated Workflows

Tao Huang; Guosen Wu; Chen Hou; Guolong Zheng arXiv: 2609.19226

Figure from PAPC: Platform Mediation for Privacy-Propagation Externalities in AI-Mediated Workflows
PAPC: Platform Mediation for Privacy-Propagation Externalities in AI-Mediated Workflows

AI-mediated platforms coordinate work through LLM agents acting for different principals. In these workflows, privacy loss can be created before a final answer appears: a memory write, shared-workspace update, inter-agent message, or tool event may impose downstream exposure cost on another principal.

PAPC models privacy loss created before any final answer — a memory write, shared-workspace update, inter-agent message or tool event — as a privacy-propagation externality depending on topology and fanout. A platform mediator intercepts information-moving events and can allow, release a policy-safe abstraction, quarantine raw content, block, or narrow onward rights; on retrieval-memory and multi-agent benchmarks it keeps task completion and eliminates measured raw-value exposure.

Key insight: Privacy in multi-agent workflows has to be enforced at the level of individual information-moving events — memory writes, messages, tool calls — not just the final answer.

MATCH: Model-Aware Tool Learning with Curriculum Scheduling and Hierarchically Gated Rewards

Shihao Liu; Hao Yin; Lijun Liu; Zhengzong Chen; … arXiv: 2609.20082

Figure from MATCH: Model-Aware Tool Learning with Curriculum Scheduling and Hierarchically Gated Rewards
MATCH: Model-Aware Tool Learning with Curriculum Scheduling and Hierarchically Gated Rewards

Tool learning enables large language models (LLMs) to use external tools for tasks beyond parametric knowledge. Reinforcement learning can optimize tool-call behavior from feedback, but current methods still face two problems: fixed-threshold curricula can become misaligned with the policy's evolving capability boundary, and additive rewards can leak argument-level credit when the predicted tool is wrong.

MATCH: model-aware curriculum near the policy's capability boundary plus Hierarchical Tool-call Gated Reward (tool name → argument key → value, credit only when prerequisites hold). 72.19% on API-Bank and 62.87% on BFCL V3, consistent across four backbones.

Key insight: Tool-learning rewards work better when credit for arguments is gated on choosing the right tool, combined with a curriculum that tracks the model's capability.

Inference-Engine Fingerprinting Attacks are Practical: Exploring Model-Driven Environmental Discovery, Exploitation, and Escape

Sarah Radway; Andrew Cheng; Vijay Janapa Reddi; James Mickens arXiv: 2609.20614

Figure from Inference-Engine Fingerprinting Attacks are Practical: Exploring Model-Driven Environmental Discovery, Exploitation, and Escape
Inference-Engine Fingerprinting Attacks are Practical: Exploring Model-Driven Environmental Discovery, Exploitation, and Escape

Frontier AI models are rapidly gaining the ability to exploit vulnerabilities in complex pieces of software. The risk is not theoretical, as evidenced by recent sandbox escapes performed by frontier models at OpenAI and Anthropic.

Shows a misaligned model can fingerprint which inference engine runs it (examples for five popular engines, e.g. vLLM, SGLang) using realistic agentic harnesses, then use engine-specific exploits via crafted output tokens alone; includes a proof-of-concept exploit chain to bare metal.

Key insight: A model can fingerprint the inference engine running it and attempt engine-specific exploits using only its output tokens, putting the engine inside the agent's attack surface.

Coding Agents with an Obstacle-Aware Harness for Safe Robot Manipulation

Bingxin Xu; Yuzhang Shang; Zhen Dong; Emilio Ferrara arXiv: 2609.20822

Figure from Coding Agents with an Obstacle-Aware Harness for Safe Robot Manipulation
Coding Agents with an Obstacle-Aware Harness for Safe Robot Manipulation

Coding agents have emerged as a promising paradigm for robot manipulation: a language model writes the robot controller as a program, and agents built in this way now operate robots without robot-specific training.Whether this paradigm is also safe, however, has not been asked. We evaluate coding agent under a safety constraint, where each task pairs a manipulation goal with an obstacle the robot must not touch.

Coding agents writing robot controllers collide with a forbidden obstacle in most cases even though their traces mention it and the prompt forbids it; SafeHarness adds obstacle-aware route planning and contact harnesses so the constraint is enforced in planning.

Key insight: Telling a coding agent about a safety constraint is not enough; the harness has to make the constraint part of planning.

Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL

Juzheng Zhang; Disha Makhija; Manoj Ghuhan Arivazhagan; Vinayshekhar Bannihatti Kumar; … arXiv: 2609.20715

Agent trajectories record what an agent does and what happens next. Yet standard supervised fine-tuning (SFT) applies loss only to agent-authored action tokens, using environment observations as context but not as prediction targets.

ActObs also supervises environment observation tokens during SFT. After GRPO on Qwen3-4B it beats action-only at every pass@k on Terminal-Bench 2.0; on Qwen3-8B +3.4 pp pass@16; +4.2 pp pass@1 on aider-polyglot at 4B.

Key insight: Supervising observation tokens during SFT gives agents a better starting point for RL without extra data or compute.

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

Yan Yu; Zhengxi Lu; Yizhou Liu; Yichen Pan; … arXiv: 2609.20784

Figure from RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from a self-teacher with privileged task skills, letting a skill-free student internalize them. This recipe, however, is undermined by two findings in agentic tasks: privileged information alone does not always make a teacher reliable, and the benefit of teacher supervision is stage-dependent.

RetireOPD trains a skill-conditioned teacher with environment rewards, then a skill-free student with RL + on-policy distillation that retires the teacher adaptively. Qwen2.5 1.5B–7B: +14.1–18.8% ALFWorld and +11.8–19.0% WebShop over RL; student surpasses its teacher.

Key insight: A skill-conditioned teacher can be distilled into a skill-free student that retires the teacher on its own and ends up surpassing it.

Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization

Yingxuan Zhuang; Binhe Yu; Jingxiao Yang; Ruopei Sun; … arXiv: 2609.19830

Figure from Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization
Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization

Reinforcement learning for LLM agents involves two distinct optimization di- mensions: how environment feedback is exploited within a trajectory, and how complete trajectories are aggregated across a batch. We formulate these dimen- sions as Intra-Trajectory Feedback Attribution and Inter-Trajectory Objec- tive Aggregation, and introduce BATON (Bayesian Attribution and Trajectory Objective Normalization), a dual-axis policy optimization framework.

BATON: Bayesian Feedback Attribution within trajectories plus Trajectory Mass Normalization across trajectories; independent gains with GRPO and GiGPO on ALFWorld, WebShop, SearchQA.

Key insight: Agent RL has two separable optimization axes — within-trajectory credit and across-trajectory aggregation — and improving both compounds.

For Your Eyes Only: Evaluating Coordination Between Isolated Language Model Instances

Alexander Shirnin; Aleksey Kudelya arXiv: 2609.19504

Figure from For Your Eyes Only: Evaluating Coordination Between Isolated Language Model Instances
For Your Eyes Only: Evaluating Coordination Between Isolated Language Model Instances

As model-generated content is increasingly consumed by other model instances in automated workflows, a practically important question arises: can a model embed a signal in natural language that an independent instance of the same model can detect, relying only on shared pre-training and task instructions, without any shared memory or coordination-specific training? We introduce For Your Eyes Only, a cooperative signalling game designed to evaluate this directly.

For Your Eyes Only: can a model embed a signal that an isolated instance of the same model detects, with no shared memory? Most models struggle once detectable signals must be avoided, but one frontier model stays near-perfect; models can also misdirect; coordination is weaker across architectures.

Key insight: Isolated instances of the same model can sometimes coordinate through natural language alone, which matters for monitor/worker setups built from one model.

Language-model groups overstate consensus when replaying human deliberation on a reasoning task

Tengfei Shao arXiv: 2609.20543

Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final states are operationalized. We replayed 100 held-out human Wason groups with matched large language model (LLM) agent groups, seeding one belief-anchored agent per participant's pre-discussion answer and scoring agents and people with the same code.

Replaying 100 human Wason-task groups with matched LLM agent groups: agents almost always post while ~1/5 of humans never did, and agent groups are 34–44 pp more consensual in sensitivity analyses.

Key insight: LLM agent groups replaying human deliberation reach consensus far more often than the humans did, so agent agreement is a poor proxy for collective reasoning.

When Hiring Becomes Agent-Mediated: Evaluating Access and Recurrence in Two-Agent Résumé Screening

Jian Gao; Hang Jiang arXiv: 2609.19530

Hiring is bilateral: employers assess fit, while candidates present and defend evidence of their qualifications. Yet résumé screening, the first gate, is commonly automated as a static, one-call judgment over a résumé-job pair.

Two-agent résumé screening (employer-side and candidate-side agents exchange evidence) vs one-call judgment on 600 constructed pairs: two-agent screening advances more applications (GPT-5.5 33.3%→39.3%; Opus 4.7 34.0%→35.5%).

Key insight: Letting employer-side and candidate-side agents exchange evidence changes screening outcomes compared with a single one-shot judgment.

GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning

Ruiyang Wang; Hao-Lun Hsu; Swarajh Mehta; Jiwoo Kim; … arXiv: 2609.19315

Figure from GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning
GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning

Large language models (LLMs) provide a flexible interface for long-horizon robot planning, but generated plans often fail to respect embodiment constraints, recover from planning errors, or reason effectively under partial observability. We present GAVEL, a framework for verifying and repairing long-horizon LLM planning built around an explicit graph world model.

GAVEL verifies and repairs LLM plans against an explicit graph world model (object relations, preconditions/effects, beliefs over unobserved locations), reserving LLM replanning for semantic errors. On BEHAVIOR-1K with Qwen3-8B, single-task success 41.2%→91.8%, multi-task 19.9%→92.6%.

Key insight: An explicit world model that checks and repairs plans before execution can dramatically raise long-horizon task success even for compact models.

Tailored to you: longitudinal effects of personalising language models

Canfer Akbulut; Justine Breuch; Arianna Manzini; Lujain Ibrahim; … arXiv: 2609.20077

Interest in developing personalised language models is rapidly growing. While personalisation is often viewed as a mechanism to better serve diverse user needs, the effects of sustained interactions with personalised models on people's perception of and behaviour toward AI remain poorly understood.

Five-day study with 992 participants comparing non-personalised, memory-based and survey-based personalisation: memory-based users self-disclosed more and rated the model less creepy; survey-based users reported more regret about sharing personal information.

Key insight: How a model learns about a user changes how the user feels about sharing: memory-based personalization felt less creepy than an up-front intake survey.