Saturday’s digest covers leftovers from Friday’s cs.AI announcement day (171 new submissions and cross-lists). Most stack-relevant papers were already published on the Friday page; these five still pass the agent/memory/tools filter and were not yet covered. Two papers have a figure extracted from the HTML/PDF.
Kesheng Chen; Yamin Hu; Wenjian Luo arXiv: 2609.11636
LLM optimization agents usually treat each natural-language request in isolation. MAPLE keeps the optimization program, accepted plans, earlier updates, and candidate solutions so later NL revisions can reuse search state. It combines language-based problem construction with mathematical programming and evolutionary search, and introduces NLDO (15 trajectories, 180 updates). Main eval: completes all trajectories with online scalar quality 0.951 and Pareto hypervolume ratio 0.875.
Key insight: Keep the executable plan and candidate solutions across NL revisions, not only a chat summary.
Qinzhen Ma; Ruihai Wu arXiv: 2609.10873
Independent evaluation can reject harmful policy updates and also starve useful continual learning. The paper argues update admission needs both error control and retained learning opportunity at a stated interaction budget. Range-based confidence gates can admit zero useful updates; a paired-binomial construction recovers 31.6% of an update stream at 2,000 episodes per stage while unconditional replay still learns better closed-loop.
Key insight: Audit missed-opportunity rate alongside false-admit rate before shipping validate-then-write curators.
Jessica Pourleyli; Maitreyee Das Urmi; Glaucia Melo arXiv: 2609.10762
Defines Static-Pass Dynamic-Fail (SPDF): code that clears static gates but remains exploitable at runtime. A three-stage agentic pipeline (Bandit+Semgrep, LLM CWE reasoning, Docker exploit verification) finds that of 654 Bandit-Semgrep-clean Python samples, about 1 in 7 (95/654, 14.53%) still showed confirmed or partially confirmed runtime-exploitable issues.
Key insight: Treat static-clean as necessary but not sufficient before promoting generated Python into an agent harness.
Jeongyeon Kim; John Mitchell arXiv: 2609.11109
Empirical study of multi-agent LLM qualitative coding with independent code, debate, and reconcile stages. Accuracy depends on codebook length, data similarity, and agent disagreement. Intense and unresolved debates between agents were associated with higher coding accuracy, suggesting disagreement traces can be a quality signal rather than pure noise.
Key insight: Preserve disagreement traces in multi-agent review loops instead of forcing early consensus.
Spandan Ghose Chowdhury arXiv: 2609.11190
Seller-side multi-agent system for LLM-mediated ecommerce: defines Agentic Share-of-Search (ASoS), deploys query agents across AI shopping platforms, and uses a ReAct diagnostic agent for prioritized merchandising interventions. In a 100-trial ablation study, the system recovers the ablated signal in 39% of trials overall and 63.9% among high-correlation ablations.
Key insight: Query-swarm plus ReAct diagnosis is a concrete pattern for measuring visibility in LLM-mediated discovery.