The Morning
From the arXiv
CHARM: A Multimodal Graph Foundation Model with Hierarchical Context Modeling for Zero-Shot Transfer
HARM is a multimodal graph foundation model that addresses zero-shot transfer by modeling hierarchical context across different modalities. Its core method involves learning transferable cross-modal relations and disentangling domain-specific information from generalizable node representations. This allows CHARM to generalize to new graph domains and tasks without requiring any downstream fine-tuning.


HiSkill: Empowering LLM Agents with Hierarchical Skill Graphs
HiSkill addresses the limitations of flat skill representations in LLM agents by introducing a hierarchical skill graph. This framework organizes skills and actions into a directed graph, capturing complex relationships like decomposition and temporal transiti…
MemLens: A Value-Aware Memory Management System with Interactive Analytics for LLM-based Agents
MemLens introduces a value-aware memory management system for LLM agents, treating memory records as first-class objects. Its core method involves Shapley-style evaluation to identify and prioritize valuable memory content, enabling efficient storage and retri…


Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL
This paper introduces a "self-speculating agent" that unifies task execution and next tool call prediction within a single model. By training this agent using a joint reinforcement learning method, it learns to predict its future tool calls by leveraging its o…
Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction
This paper introduces HYSET, a novel method for LLM agents to retrieve tool sets. Instead of evaluating tools individually or sequentially, HYSET treats the entire tool set as a unit, predicting hyperedges on a tool co-invocation graph to capture joint utility…

RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement
This paper introduces RSIBench-Data, a benchmark designed to isolate and evaluate the data-centric research capabilities of LLM agents for recursive self-improvement. The core meth…
DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space
DecoEvo addresses the limitations of fixed evaluation in text-space LLM optimization by introducing a decoupled co-evolutionary approach. It simultaneously trains a solver to impro…
How Do LLMs Read Bug Reports? An Empirical Study of Attention in LLMs for Automated Program Repair
This paper empirically studies how Large Language Models (LLMs) attend to information within bug reports when performing automated program repair. By analyzing attention patterns o…
Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
Messier is a large, standardized corpus of 957,253 records from 30 benchmarks and 714 agents, designed to unify and enable cross-benchmark evaluation of AI agents. Its core contrib…
Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models
This paper introduces an "input-only" method to suppress specific internal activations in Large Language Models without modifying the model itself. By optimizing prompts, they aim …
The Town Square
This article explores DeltaNet, a family of linear attention variants designed to improve efficiency and performance in transformer models.
Workshops
Jenkins is an open-source automation server that enables developers to reliably build, test, and deploy their software through a vast ecosystem of plugins.
Airi is a self-hosted, user-owned Grok companion that brings virtual waifus and cyber beings to life, capable of real-time voice chat and playing games like Minecraft and Factorio across Web, macOS, and Windows.