The Morning
From the arXiv
DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents
his paper introduces DBA-Bench, a novel benchmark designed to accurately evaluate LLM-based database agents in production-like environments. It addresses key gaps by simulating multi-turn read-write interactions with live databases, handling complex observations, and allowing for diverse remediation strategies. DBA-Bench's core contribution is its production fidelity, enabling more realistic and reliable assessment of these agents' capabilities.


From Isolated Tasks to Structured Capabilities: A Multilayer Taxonomy for Large Language Models
This paper introduces a novel multilayer taxonomy of LLM capabilities, organized by human cognitive science principles rather than LLM architecture. This framework, comprising 14 capability domains and 91 subskills across Primitive, Constructed, and Integrativ…
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding
HiKV addresses the KV cache memory bottleneck in LLM decoding by compressing it hierarchically. It first evicts unimportant tokens and then further compresses retained tokens by keeping only significant elements. This algorithm-hardware co-design, featuring a …


IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation
This paper introduces IDEAgent, a multi-agent framework for research idea generation that treats ideation as a Quality-Diversity (QD) search. Unlike previous methods that optimize for quality or diversity separately, IDEAgent jointly drives both objectives. It…
Learning on the Job: Continual Learning from Deployment Feedback for Frozen-Weights Agents
This paper proposes a continual learning method for deployed AI agents with frozen weights. It leverages deployment feedback, such as outcome verdicts and corrections, to train an external memory that stores natural-language rules. This approach significantly …
Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Mode
Nanbeige4.2-3B is a compact 3B parameter agentic model that achieves strong performance in code, office, and tool-use tasks, along with competitive reasoning. Its core method invol…
The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents
This paper introduces the "regression tax" to analyze the impact of adding procedural skills to LLM agents. Instead of just measuring average improvement, it quantifies how skills …
Towards Trustworthy and Cost-Efficient Data Integration: From Naïve RAG to Agentic RAG
This paper proposes Agentic RAG as a solution to improve trustworthiness and cost-efficiency in LLM-based data integration. It builds upon existing RAG methods by introducing auton…
Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills
This paper introduces Skill Self-Play (Skill-SP), a novel framework for LLM training that addresses the trade-off between task diversity and verification reliability. Skill-SP uses…
A Roadmap to Impactful Pluralistic Alignment Research
This paper argues that pluralistic AI alignment research, aiming to represent diverse human values, is currently failing to impact real-world AI systems. The authors find no eviden…
The Town Square
A US citizen was charged after his GrapheneOS phone automatically wiped itself during an airport search, raising concerns about data privacy and border security.
Workshops
Ego-lite is a zero-cost, zero-config browser designed for AI agents to perform browser automation by sharing logged-in states without user interruption.
Block/buzz is a decentralized communication platform enabling a hive mind by facilitating peer-to-peer message exchange and collective intelligence.