Daily Issue
Vol. I — No. 44
04 · 08
Tuesday, 4 August 2026
Generated 2026-08-04 10:21
google/gemini-2.5-flash-lite
The family is one of nature's masterpieces. — George Santayana 37 items · 3 sections
§ 0

The Morning

Local weather 1
This morning in
London
Clear sky
Today's range
29.4°20.0°
currently 26.7°
Feels
26.5°
Rain
39%
Wind
8 km/h
Humid
35%
Rise
05:29
Set
20:44
§ I

From the arXiv

arXiv preprints 10 of 20
cs.AIarxiv:2608.02407v1Lead article

Antares: Foundation Models for Agentic Vulnerability Localization

Supriti Vijay, Aman Priyanshu, Didier Chapoteau, Arthur Goldblatt, Jianliang He

ntares is a family of compact foundation models designed for agentic vulnerability localization in software. Its core method involves a two-stage training pipeline combining supervised fine-tuning with reinforcement learning, enabling it to reason over codebases and identify vulnerabilities. Antares' key contribution is achieving state-of-the-art performance comparable to much larger models, while offering efficient, low-cost local inference.

F1 score versus model size on Vulnerability Localization Benchmark (VLoc Bench) [ manuscript-vlb ] , a repository-scale benchmark comprising 500 tasks across 290 unique real-world repositories, where models receive only a CWE category description and must identify vulnerable implementation files in real codebases. Antares models form the Pareto frontier among evaluated models, achieving the strongest localization quality at small parameter scales. Antares-3B reaches near-frontier closed-source performance while remaining orders of magnitude smaller than GPT-5.5 and Gemini-family baselines.
F1 score versus model size on Vulnerability Localization Benchmark (VLoc Bench) [ manuscript-vlb ] , a repository-scale benchmark comprising 500 tasks across 290 unique real-world repositories, where models receive only a CWE category description and must identify vulnerable impl…
Real agent traces ( ollama7b : qwen2.5:7b against real tools). One-class CUSUM score streams per failure class; dashed = threshold at the 5% validation FA budget, dotted = verified injection onset. The healthy run stays two orders of magnitude below the line, and every injected class — context corruption, goal drift , looping and tool cascade — alarms one step after onset. The sixth panel is grounding loss , and it is different in kind: fabrication cannot be injected into a live run, so this episode is a genuine one from the organic (non-injected) corpus, scored against that corpus’s own healthy null. The CUSUM stays flat – a fabricated figure perturbs no behavioural channel – and the class is caught instead by the deterministic grounding verifier. That is the blind spot the verifier exists to cover, shown rather than asserted. Note the y y axis is symlog: real streams span twelve orders of magnitude, because the CUSUM accumulates multiplicatively once a failure takes hold.
Real agent traces ( ollama7b : qwen2.5:7b against real tools). One-class CUSUM score streams per failure class; dashed = threshold at the 5% validation FA budget, dotted = verified injection onset. Th…
cs.AIarxiv:2608.02464v1

Real-Time Detection and Repair of LLM Agent Failures

Sunny Dubey

This paper proposes a cost-effective method for detecting LLM agent failures using observable step telemetry, avoiding expensive step-by-step validation. Their core contribution is a one-class echo-state-network ensemble with CUSUM alarms that can detect a sig…

cs.AIarxiv:2608.02442v1

Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks

Xuan Ren, Weiqi Zhai et al.

This paper introduces "Solution Hacking," a phenomenon where LLMs achieve correct answers on scientific reasoning benchmarks through invalid shortcuts rather than genuine reasoning. The authors demonstrate that this hacking significantly inflates accuracy scor…

Three solution paths for the same quadratic-equation problem. The upper path represents the expected reasoning process, in which the LLM exercises the targeted reasoning capability and derives the correct answer. The middle path represents a reasoning error, where an invalid derivation leads to an incorrect answer. The lower path represents Solution Hacking, where the LLM obtains the correct answer through enumeration, guessing, search, or answer-first verification without completing the targeted derivation. Traditional evaluation can distinguish the upper and middle paths, but may fail to distinguish the upper and lower paths because both yield the correct final answer.
Three solution paths for the same quadratic-equation problem. The upper path represents the expected reasoning process, in which the LLM exercises the targeted reasoning capability and derives the cor…
Motivation. We argue that independent skill retrieval may overlook skill dependencies, while SkillTrace connects atomic queries and skills through a query–skill graph, enabling dependency-aware traversal to retrieve a composable skill set.
Motivation. We argue that independent skill retrieval may overlook skill dependencies, while SkillTrace connects atomic queries and skills through a query–skill graph, enabling dependency-aware traver…
cs.AIarxiv:2608.02356v1

SkillTrace: Traversing a Query-Skill Graph for Composable LLM Agents

Yue Yao, Shengyuan Wang et al.

SkillTrace addresses the challenge of composing reusable skills for LLM agents by modeling skill relationships as a three-level graph. It organizes user queries semantically, matches them to skills, and propagates dependencies to find executable compositions. …

cs.LGarxiv:2608.02352v1

Qwen-CUA: Native Computer Use for (almost) Everything

Dunjie Lu, Shuai Bai et al.

Qwen-CUA is a native computer-use agent that operates software solely through screenshots and keyboard/mouse inputs, avoiding direct access to underlying code or APIs. Its core method involves a novel scaffold for managing long-term visual history and a large-…

№06
cs.CL
9

AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling

Jiajun Liang, Yucheng Liao et al.

AURORA-LM introduces a novel approach to continuous-latent diffusion language modeling by decoupling representation learning from distribution modeling. It constructs a high-capaci…

№07
cs.AI
8

A Taxonomy of Cognitive Capability Gaps in Generative and Agentic AI

Taye Akinrele, Sindhuja Penchala et al.

This paper identifies and categorizes key cognitive capability gaps hindering the development of advanced Cognitive AI, moving beyond simple generation and task execution. It propo…

№08
cs.AI
8

Can AI Agents Simulate A/B Test Outcomes? A Validation Framework for Agentic Experimentation

Stefan Hut, Lorenzo Masoero

This paper proposes a framework called Simulated Randomized Controlled Trial (S-RCT) to assess if AI agents can accurately predict A/B test outcomes. The core method involves decom…

№09
cs.AI
8

Can Foundation Models Hear What Made That Sound? A Tiered Benchmark of Audio-Language Models and Traditional Classifiers for Closed-Set Sound Source Identification

Sajjad Abdoli, Ghassan Al-Sumaidaee et al.

This paper benchmarks various audio classification models, including foundation models and traditional classifiers, on a sound source identification task. It introduces a tiered ev…

№10
cs.AI
8

Cooperative Coevolution for Resource-Constrained Agentic LLM Post-Training

Zhiyuan Wang, Shengcai Liu et al.

This paper introduces Cooperative Parameter-subspace Evolution Strategy (CoPES) to address the memory and computational demands of post-training tool-using LLM agents. CoPES decomp…

§ II

The Town Square

Hacker News 8
713
research.jfrog.com3 Aug
178
Ask HN: Who is hiring? (August 2026)
3 Aug
118
Ask HN: Who wants to be hired? (August 2026)
3 Aug
compiled overnight by google/gemini-2.5-flash-lite · end of issue no. 44 · thank you for reading