The Morning
From the arXiv
Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents
his paper introduces "Digital Pantheon," a novel multi-agent framework for simulating political coalition formation using LLMs. It combines SFT, DPO, and RAG to create partisan agents that are both ideologically aligned and factually grounded. The framework's contribution lies in enabling realistic, interpretable simulations of complex political negotiations, demonstrated on a real-world election scenario.


Mask-Aware Policy Gradients for Diffusion Language Models
This paper introduces a novel reinforcement learning method for Masked Diffusion Language Models (MDLMs) by treating generation as a two-stage action Markov Decision Process. This approach decomposes the policy gradient into token prediction and masking decisi…
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios
OmniaBench introduces a comprehensive benchmark for evaluating general AI agents by creating diverse, executable scenarios derived from real-world applications. Its core method involves constructing a hierarchical taxonomy of domains and synthesizing tasks acr…
Scaling Behavior Foundation Model for Humanoid Robots
This paper investigates how to effectively scale Behavior Foundation Models (BFMs) for humanoid robots. Their core method involves coordinating three key components: a motion tracking learning paradigm, specific behavioral data, and model architecture. The mai…
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
This paper introduces SearchOS-V1, a multi-agent framework for robust open-domain information seeking. Its core method is to represent search progress as explicit, shared state, moving beyond the limitations of implicit tracking in current systems. This explic…

When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space
This paper investigates the distinction between text-based safety and physically grounded danger in Large Language Models (LLMs). It demonstrates that these two types of danger are…
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
LongStraw addresses the challenge of training Reinforcement Learning (RL) agents with extremely long contexts (over 2 million tokens) within a limited GPU budget. Its core method i…
ANet Patu-1: The Value of Connection in the Agent Network
This paper introduces ANet Patu-1, a self-organizing consensus protocol for AI agents. It models the value of agent networks based on coordination group size, deriving properties f…
Can LLMs Build a MaxSAT Solver from Papers? The CoreForge Experience
CoreForge demonstrates the feasibility of using LLMs like ChatGPT and Codex to construct an unweighted MaxSAT solver by interpreting research papers. The project's core method invo…
RoboTTT: Context Scaling for Robot Policies
RoboTTT introduces a novel method for scaling robot policy context to 8,000 timesteps by integrating Test-Time Training (TTT) into foundation models. This allows the model to compr…
The Town Square
Open Frontier Intelligence has launched Kimi K3, a new large language model designed for advanced reasoning and complex tasks, aiming to push the boundaries of AI capabilities.
Workshops
Hallmark is a design skill that helps Claude Code, Cursor, and Codex avoid generating AI-generated "slop" by promoting better coding practices.
OpenCut is an open-source video editing application that provides a free and accessible alternative to CapCut, offering a suite of creative tools for video manipulation and production.