Daily Issue
Vol. I — No. 57
21 · 08
Friday, 21 August 2026
Generated 2026-08-21 09:44
google/gemini-2.5-flash-lite
Think about scary movies: There's a fine line between horror and humor. — Roy Blount, Jr. 35 items · 3 sections
§ 0

The Morning

Local weather 1
This morning in
London
Overcast
Today's range
21.0°14.4°
currently 17.3°
Feels
15.9°
Rain
45%
Wind
9 km/h
Humid
59%
Rise
05:55
Set
20:11
§ I

From the arXiv

arXiv preprints 10 of 20
cs.AIarxiv:2608.20274v1Lead article

Break It Down, Pass It On: Cross-Task Skill Transfer in LLM Agents

Yiyang Feng, Biddut Sarker Bijoy, Niranjan Balasubramanian, Jiawei Zhou

his paper investigates how LLM agents learn and reuse skills across tasks. The core method involves comparing different skill induction strategies: task-level vs. subtask-level and text vs. code formats. The key contribution is demonstrating that subtask-level skill induction and text-based skill representation lead to more reliable and beneficial skill transfer, improving agent performance compared to task-level induction or code-based skills.

Top: The common skill-selection paradigm for coding agents: given a task, the system selects skills from a library and provides them to a frozen LLM executor. Bottom: Effective skill selection depends on capability composition rather than individual relevance. The LLM executor benefits from skill sets that cover the required capabilities, while redundant and irrelevant skills consume context budget with little or negative utility.
Top: The common skill-selection paradigm for coding agents: given a task, the system selects skills from a library and provides them to a frozen LLM executor. Bottom: Effective skill selection depends…
cs.AIarxiv:2608.19993v1

Optimal Skill Selection for LLM Agents with Provable Bicriteria Guarantees

Yu Chen, Ruishuo Chen et al.

This paper addresses the challenge of selecting optimal skills for LLM agents within a limited context window. The authors propose a novel optimization framework that balances skill benefit against token cost, moving beyond independent skill scoring. Their dev…

cs.CLarxiv:2608.20153v1

FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models

Dingzirui Wang, Xuanliang Zhang et al.

This paper introduces FormalTCS, an expert-validated benchmark for evaluating LLMs on realistic, end-to-end theoretical computer science research tasks. It comprises 175 instances from top TCS conferences, preserving original definitions and proof structures w…

cs.AIarxiv:2608.20318v1

AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement

Yizhe Chi, Wenyi Li et al.

This paper introduces AI4AI-Bench, a novel benchmark designed to evaluate Large Language Model (LLM) agents' ability to design training algorithms for recursive self-improvement (RSI). The core method involves agents modifying existing training algorithms with…

cs.AIarxiv:2608.20161v1

DARS: Dual-Level Credit Assignment RL with Structured Reasoning for Instruction-Based Image Editing

Haoxiang Cao, Jiajiong Cao et al.

This paper introduces DARS, a reinforcement learning framework for instruction-based image editing. DARS addresses the inefficiency of traditional training by employing dual-level credit assignment. It achieves this by using multi-plan, multi-render rollouts t…

Figure 1. Two visually poor editing outcomes can require opposite cross-module update emphases. Case A : the planner captures the requested edit reasonably well, but the renderer fails to realize it, so this sample should place more corrective update weight on the renderer. Case B : the sampled plan is a less useful intermediate for the instruction, and the renderer faithfully executes it, so this sample should place more corrective update weight on the planner. DARS is designed to distinguish render-dominant from plan-dominant update needs through routing signals; when planner-side emphasis is needed, its structured slots support finer within-planner diagnosis. Teaser showing two low-scoring failure cases in a two-stage image editing pipeline that require opposite cross-module credit-assignment decisions: one should emphasize renderer correction, while the other should emphasize planner correction.
Figure 1. Two visually poor editing outcomes can require opposite cross-module update emphases. Case A : the planner captures the requested edit reasonably well, but the renderer fails to realize it, …
№06
cs.AI
8

From Agent Behaviour to Agent-Friendly Documentation: An Empirical Study of How Coding Agents Discover, Read, and Write Technical Documentation

Zhijun Gao, Jing Chen

This paper empirically studies how coding agents interact with technical documentation. It reveals that agents primarily consult agent-facing artifacts like instruction files and w…

№07
cs.AI
8

Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization

Qian Kou, Xiaofeng Shi et al.

This paper introduces IAR, a three-stage post-training method for enabling large language models to answer questions from a fixed document set without retrieval. IAR first injects …

№08
cs.AI
8

Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts

Nayeon Kim, Hojin Lee et al.

This paper introduces a compute-efficient method to transfer optimal learning rates for large Mixture-of-Experts (MoE) models. It first demonstrates that optimal learning rates are…

№09
cs.AI
8

Manifold Drift in Flow Preference Optimization: A Root Cause of Reward Hacking

Yansen Han, Shengyi Liao et al.

This paper identifies "manifold drift" as a key cause of reward hacking in continuous-time generative model alignment. The authors show that standard preference optimization can pu…

№10
cs.AI
8

MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use

Mengru Wang, Haozhe Luo et al.

This paper introduces MemTrapBench, a novel benchmark designed to evaluate how retrieved memories can negatively impact LLM reasoning and performance, a phenomenon termed "cognitiv…

§ II

The Town Square

Hacker News 6
compiled overnight by google/gemini-2.5-flash-lite · end of issue no. 57 · thank you for reading