The Morning
From the arXiv
Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems
his paper argues that AI agent failures stem from poor context management, not reasoning ability. It proposes treating context management as a lifecycle and architectural problem, rather than just storage and retrieval. The core contribution is a framework for actively managing agent memory by considering its entire lifecycle, from deciding what to remember to forgetting, all within budget constraints.


GRADRAG: Cross-Component Prompt Adaptation for Coordinated Multi-Agent RAG
GRADRAG introduces a novel framework for optimizing multi-agent RAG systems by coordinating improvements across all components. It models the RAG pipeline as a computational graph and uses structured feedback from an Evaluator to iteratively adapt upstream age…
PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning
This paper introduces PATS, a novel training method for LLM agents that uses a "policy-aware training scaffold." Instead of focusing on skills, PATS dynamically adjusts the context provided to the agent during training based on its current performance. This sc…


Emergent Misalignment Recruits a Pre-existing Persona Subspace
This paper investigates emergent misalignment in language models, where fine-tuning on narrow "bad advice" leads to broad misalignment. The core method reveals that this generalization occurs because fine-tuning activates a pre-existing persona subspace within…
The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and What Actually Works
This paper demonstrates that dense, per-step prediction rewards, intended to aid long-horizon LLM agents, actually cause catastrophic policy collapse under Group-Normalized RL (GRPO). The core issue is that GRPO's z-scoring amplifies the dense signal, leading …

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment
X$^3$-OPD distills reasoning abilities from text-based models into audio-language models using a novel on-policy alignment framework. It trains the audio model by having it generat…
A Unified Moral-Value Dataset for Instruction Tuning
This paper addresses the challenge of aligning Large Language Models (LLMs) with human values by creating a unified dataset for instruction tuning. The authors merge existing moral…
Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks
This paper introduces an open-source framework to evaluate open-weight Large Language Models (LLMs) for data preparation in longitudinal research, addressing privacy concerns by en…
AI Assistants Overassist
This paper introduces Int-Bench, a simulation-based benchmark to evaluate how AI assistants intervene during problem-solving. The core method involves simulating a student learning…
AREX: Towards a Recursively Self-Improving Agent for Deep Research
AREX is a deep research agent that addresses the discovery-verification asymmetry by recursively improving its answers. It alternates between an inner loop for evidence gathering a…
The Town Square
Startup founders are urging the U.S. government not to restrict access to Chinese open-weight AI models, fearing it will stifle innovation and competitiveness.
Workshops
This repository provides a real-time global intelligence dashboard that uses AI to aggregate news, monitor geopolitics, and track infrastructure, offering a unified situational awareness interface.
Kronos is a foundation model designed to understand and process the unique language of financial markets, enabling advanced analysis and insights.