Daily Issue
Vol. I — No. 43
03 · 08
Monday, 3 August 2026
Generated 2026-08-03 10:44
google/gemini-2.5-flash-lite
Myself when young did eagerly frequent doctor and saint, and heard great argument about it and about: but evermore came out by the same door as in I went. — Omar Khayyam 35 items · 3 sections
§ 0

The Morning

Local weather 1
This morning in
London
Clear sky
Today's range
31.2°18.5°
currently 26.2°
Feels
26.8°
Rain
6%
Wind
9 km/h
Humid
43%
Rise
05:27
Set
20:45
§ I

From the arXiv

arXiv preprints 10 of 20
cs.AIarxiv:2607.29626v1Lead article

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers

Tianyu Huai, Tingshuo Fan, Xinchi Chen, Yining Zheng, Yuxin Wang

his paper introduces AgentHPOBench, a novel benchmark designed to evaluate LLM agents' ability to perform sequential hyperparameter optimization. Unlike previous benchmarks, it assesses how agents interpret experimental evidence to guide subsequent configuration choices across diverse machine learning tasks. The contribution lies in providing a standardized framework to measure and compare the experimental optimization capabilities of LLM agents.

Comparison of natural-language reflection, direct code verification, and AMTFV for mathematical answer verification and correction.
Comparison of natural-language reflection, direct code verification, and AMTFV for mathematical answer verification and correction.
cs.AIarxiv:2607.29549v1

AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction

Rui Zou, Yutao Zhu et al.

AMTFV introduces a novel "Mathematical Tool Flow" (MTF) interface to enable LLMs to reliably verify their mathematical answers. This method decouples verification modeling from execution by allowing the LLM to construct a workflow, request specific computation…

cs.AIarxiv:2607.29405v1

Beyond Component Testing: Validating Agentic AI Systems

Fabio Orazio Mirto, Luca D'Agati et al.

This paper addresses the challenge of validating complex agentic AI systems, which exhibit multi-step, dynamic behaviors. It synthesizes existing research to propose a five-dimension taxonomy (behavioral, safety, temporal, regulatory, multi-agent) for characte…

PRISMA-style workflow for literature identification, screening, eligibility assessment, and inclusion. The retrieval stage yielded 7,197 unique records after source merging and deduplication across five primary sources. Sequential screening reduced the corpus to 257 papers included in the final survey.
PRISMA-style workflow for literature identification, screening, eligibility assessment, and inclusion. The retrieval stage yielded 7,197 unique records after source merging and deduplication across fi…
Illustration of our framework LEMUR: (1) Unsupervised Pre-training for the MORL agent to explore and collect diverse experiences via maximizing state entropy H(s). (2) Reward learning from Preference feedback, where each reward model is learned separately from the preferences queried from each teacher. The reward models are used to dynamically relabel the state-action pairs as a reward vector for each objective (i.e., each teacher’s preferences). (3) Multi-Objective RL agent denoted by π \( \pi_{\phi} \) uses each of the trained reward models to do multi-objective policy optimization to maximize the expected vector rewards.
Illustration of our framework LEMUR: (1) Unsupervised Pre-training for the MORL agent to explore and collect diverse experiences via maximizing state entropy H(s). (2) Reward learning from Preference …
cs.AIarxiv:2607.29559v1

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback

Manith Adikari, Bei Peng et al.

LEMUR addresses the challenge of training RL agents for tasks with multiple, conflicting objectives when explicit reward functions are unavailable. It combines Multi-Objective Reinforcement Learning with Preference-based RL, enabling agents to learn complex tr…

cs.AIarxiv:2607.29254v1

Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents

Minghui Pan, Jiayuxuan Yang et al.

This paper argues that the way AI agents are given instructions for using external tools (tool specifications) significantly impacts their safety. They found that schema-formatted specifications weaken the AI's ability to refuse harmful actions. To address thi…

Comparison of chatbot- and agent-formatted inputs.
Comparison of chatbot- and agent-formatted inputs.
№06
cs.CL
9

Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning

Xinyan Guan, Jiali Zeng et al.

This paper introduces **CaRL**, a method to train Large Language Models (LLMs) to recognize and abort "futile reasoning" on tasks exceeding their capabilities. CaRL uses reward sha…

№07
cs.CL
9

TokTier: Exact Stateful Tokenization for Agentic LLM Serving

Zhenyu Zhang, Zhichao Cao

TokTier addresses the inefficiency of re-tokenizing entire prompts in agentic LLM serving. Its core method is stateful tokenization that guarantees identical token IDs to full refe…

№08
cs.CL
9

Zero-Mem: Zero-Token Memory Operations for LLM Agents

Yilin Xiao, Zhehan Zhu et al.

Zero-Mem proposes a novel approach to LLM agent memory by eliminating token costs for memory operations. Instead of using LLM calls, it organizes interaction traces into an entity-…

№09
cs.AI
8

AgenticRepair: Multi-Faceted Program Context Engineering for Agentic Vulnerability Repair

Michael Fu, Qiyue Mei et al.

This paper introduces AgenticRepair, a framework for automated vulnerability repair. Its core method is multi-faceted program context engineering, which addresses critical gaps in …

№10
cs.AI
8

Beyond Retrieval: Analytic Memory for Multimodal Agents

Zhoujin Tian, Yao Tian et al.

This paper introduces "analytic memory" as a new paradigm for multimodal agents, complementing existing "retrieval memory." Analytic memory allows agents to compute over accumulate…

§ II

The Town Square

Hacker News 6
compiled overnight by google/gemini-2.5-flash-lite · end of issue no. 43 · thank you for reading