The Morning
From the arXiv
MemHarness: Memory Is Reconstructed, Not Replayed
emHarness proposes a novel approach to memory augmentation for LLM agents, moving beyond simple verbatim replay. Its core method involves a unified policy model that actively reconstructs retrieved past experiences based on the agent's current state. This allows agents to adapt and ground memories in the present context, mitigating negative transfer and improving decision-making.

One Human, $N$ Agents: Audit-Budget Allocation for LLM Agent Fleets under Miscalibrated, Correlated Confidence
This paper addresses the challenge of a human auditor allocating a limited budget to audit a fleet of $N$ LLM agents, whose self-reported confidence is unreliable due to miscalibration and correlated errors. The core method models this as budgeted noisy inspec…
ORCA-bench: How Ready Are Language Model Agents for Oncall?
This paper introduces ORCA-bench, a novel benchmark designed to evaluate the readiness of language model agents for on-call incident response. The benchmark simulates a production-fidelity microservice environment with real telemetry data and source code, pres…


When Specifications Conflict: A Symmetry-Based Framework for Measuring LLM Preferences
This paper introduces a novel framework for analyzing how Large Language Models (LLMs) handle conflicting instructions. By creating controlled experimental setups with explicit specification conflicts and employing a symmetry-based design, the framework allows…
(Towards) Scalable Reliable Automated Evaluation with Large Language Models
This paper introduces a scalable and reliable automated evaluation framework for LLM outputs. It uses pairwise comparisons between LLM-generated texts, aggregated via an Elo rating system, to approximate expert assessments without relying on explicit reference…

HARGO: Heterogeneity-Aware Reward-Guided Optimization for RL Post-Training of LLMs on HPC Tasks
This paper introduces HARGO, a novel RL post-training method for LLMs on HPC tasks. HARGO addresses the challenge of extreme task heterogeneity by dynamically weighting rewards bas…
LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger
This paper introduces LedgerMind, a novel framework for multimodal agents that treats their reasoning process as a provenance-constrained state machine. Its core method involves or…
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
This paper introduces Frontis-MA1, an AI model designed for recursive self-improvement in machine learning engineering (MLE). Its core method involves training a meta-evolution age…
A foundation model of numerical intelligence with cross-disciplinary generalization
This paper introduces UNICON, a foundation model designed to exhibit "numerical intelligence" by learning predictive relationships from numerical data presented as graph-based exam…
A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports
This paper introduces EndoCLIP, a vision-language foundation model specifically trained for colonoscopy. Its core method involves recovering lesion-level image-text pairs from rout…
The Town Square
Gemini Robotics 2, powered by Google DeepMind, enables robots to understand and interact with their environment using whole-body intelligence, allowing for more complex and adaptable actions.
Workshops
This repository offers a comprehensive, 12-week curriculum with 24 lessons designed to introduce Artificial Intelligence concepts to a broad audience, making AI accessible for everyone.
This repository curates a comprehensive list of resources, including libraries, strategies, and educational materials, for individuals interested in systematic trading.