Daily Issue
Vol. I — No. 36
23 · 07
Thursday, 23 July 2026
Generated 2026-07-23 10:16
google/gemini-2.5-flash-lite
I don't dismiss the music that I was involved with, I don't think it was a joke, I don't think it was funny or a phase, I don't think it was just something I was doing back then, to me it was who I am. It connects all the way through. I don't distance myself from any of it. — Ian MacKaye 38 items · 3 sections
§ 0

The Morning

Local weather 1
This morning in
London
Overcast
Today's range
25.7°17.5°
currently 22.7°
Feels
22.2°
Rain
2%
Wind
5 km/h
Humid
46%
Rise
05:12
Set
21:02
§ I

From the arXiv

arXiv preprints 10 of 20
cs.AIarxiv:2607.20268v1Lead article

PoTRE: Test-Time Reasoning inspired by Cognitive Heterogeneity

Anmol Kankariya, Sercan Ö. Arık

oTRE is a novel framework that enhances LLM reasoning by employing a heterogeneous ensemble of four specialized agents: adversarial refinement, hierarchical planning, spectrum search, and direct chaining. These agents' diverse perspectives are dynamically integrated by a task-adaptive aggregation layer to produce robust solutions for complex reasoning tasks. This approach significantly improves performance on challenging benchmarks like Humanity's Last Exam, achieving state-of-the-art results.

PRO-LONG matches or exceeds state-of-the-art ARC-AGI-3 results at 4.2 4.2 – 5.8 × 5.8\( \times \) lower token cost. Left: ARC-AGI-3 results on the public game set with a 500 action limit, grouped by model and harness. Filled bars show pass@1, with bootstrap confidence intervals where multiple runs are available (five for Codex, two for Claude Code). Outlined bars show best@ k k . We rescore the released runs of WorldModeler, Arcgentica, and Schema for a consistent comparison. 3.1 Schema reports best@2; the others report pass@1. Right: Performance versus billed tokens per game for PRO-LONG and the strongest prior harness on Codex and Claude Code, across budgets from 100 to 500 actions. PRO-LONG stays within 2 2 – 4 4 points of the strongest prior harness at 4.2 4.2 – 5.8 × 5.8\( \times \) lower cost. ∗ PRO-LONG (Fable 5) at a 2 , 000 2{,}000 -action limit; this is a lower bound on best@2, as certain games we only ran once.
PRO-LONG matches or exceeds state-of-the-art ARC-AGI-3 results at 4.2 4.2 – 5.8 × 5.8\( \times \) lower token cost. Left: ARC-AGI-3 results on the public game set with a 500 action limit, grouped by m…
cs.AIarxiv:2607.20064v1

PRO-LONG: Programmatic Memory Enables Long-Horizon Reasoning

Alexis Fox, Junlin Wang et al.

PRO-LONG introduces a programmatic memory framework for LLM agents to tackle long-horizon reasoning tasks. It addresses the challenge of context management by maintaining a complete, structured interaction log and leveraging recent advancements to efficiently …

cs.AIarxiv:2607.20090v1

Reinforcement Learning for Large Language Model Selective Evidence Adoption from Contaminated Retrieval Results

Yanyu Chen, Yue Li et al.

This paper addresses the challenge of Large Language Models (LLMs) selectively adopting evidence from potentially contaminated retrieval results. Their core method involves post-training LLMs using Reinforcement Learning with Direct Preference Optimization (DA…

A practical instance of data generation for our problem setting, which leverages a classifier ℋ \( \mathcal{H} \) that detects an harmful output by the LLM ℳ \( \mathcal{M} \) under a fixed prompt 𝐱 \( \mathbf{x} \) .
A practical instance of data generation for our problem setting, which leverages a classifier ℋ \( \mathcal{H} \) that detects an harmful output by the LLM ℳ \( \mathcal{M} \) under a fixed prompt 𝐱 …
cs.AIarxiv:2607.20286v1

Sound Probabilistic Safety Bounds for Large Language Models

Mahdi Nazeri, Anne-Kathrin Schmuck et al.

This paper introduces a framework for calculating rigorous probabilistic safety bounds for Large Language Models (LLMs), ensuring they don't generate harmful content. Their core method applies Clopper-Pearson confidence intervals and a novel algorithm that use…

cs.CLarxiv:2607.20410v1

LKValues: Aligning Large Language Models with Sri Lankan Societal Values

Nethmi Muthugala, Supryadi et al.

This paper introduces LKValues, a novel resource suite to address the Western bias in Large Language Model (LLM) value alignment. It contributes a survey-grounded set of 40 Sri Lankan societal values, an instruction corpus (LKvaluesIT) in Sinhala and English, …

The flowchart shows the process for deriving Sri Lankan societal values, starting with selecting questions from established surveys, followed by manual and LLM-assisted value elicitation. This results in 51 candidate values, with 40 values retained after calculating endorsement percentages from 205 participants, using finite population correction.
The flowchart shows the process for deriving Sri Lankan societal values, starting with selecting questions from established surveys, followed by manual and LLM-assisted value elicitation. This results…
№06
cs.CL
9

Notes to Self: Can LLMs Benefit from Experiential Abstractions?

Chang Liu, Xinyu Li et al.

This paper investigates if Large Language Models (LLMs) can improve their problem-solving abilities by learning from their own past experiences, similar to how humans create reusab…

№07
cs.CL
9

Solar Open 2 Technical Report

Sungrae Park, Sanghoon Kim et al.

Solar Open 2 is a 250B-parameter Mixture-of-Experts model designed for long-horizon agentic tasks. Its core innovation is a novel 1M-token attention mechanism that interleaves soft…

№08
cs.AI
8

Co-Evolving LLM Evaluators and Policies via DynamicRubric

Beining Wang, Weihang Su et al.

This paper addresses the challenge of improving large language models (LLMs) when evaluator feedback on similar quality responses becomes less informative. The core method, Dynamic…

№09
cs.LG
8

Statistical Inference for Rank Allocation in Low-Rank Adaptation

Yihang Gao, Vincent Y. F. Tan

This paper introduces StatLoRA, a novel method for allocating rank in Low-Rank Adaptation (LoRA) for large language models. Instead of relying on heuristic importance scores, StatL…

№10
cs.CL
8

Gotta Catch them all: the modes of Sycophancy

Shreyans Jain, Alexandra Yost et al.

This paper challenges the view of sycophancy in LLMs as a single behavior. It identifies three distinct modes of sycophancy that, while producing similar outputs, have separable in…

§ II

The Town Square

Hacker News 9
compiled overnight by google/gemini-2.5-flash-lite · end of issue no. 36 · thank you for reading