The Morning
From the arXiv
PoTRE: Test-Time Reasoning inspired by Cognitive Heterogeneity
oTRE is a novel framework that enhances LLM reasoning by employing a heterogeneous ensemble of four specialized agents: adversarial refinement, hierarchical planning, spectrum search, and direct chaining. These agents' diverse perspectives are dynamically integrated by a task-adaptive aggregation layer to produce robust solutions for complex reasoning tasks. This approach significantly improves performance on challenging benchmarks like Humanity's Last Exam, achieving state-of-the-art results.

PRO-LONG: Programmatic Memory Enables Long-Horizon Reasoning
PRO-LONG introduces a programmatic memory framework for LLM agents to tackle long-horizon reasoning tasks. It addresses the challenge of context management by maintaining a complete, structured interaction log and leveraging recent advancements to efficiently …
Reinforcement Learning for Large Language Model Selective Evidence Adoption from Contaminated Retrieval Results
This paper addresses the challenge of Large Language Models (LLMs) selectively adopting evidence from potentially contaminated retrieval results. Their core method involves post-training LLMs using Reinforcement Learning with Direct Preference Optimization (DA…

Sound Probabilistic Safety Bounds for Large Language Models
This paper introduces a framework for calculating rigorous probabilistic safety bounds for Large Language Models (LLMs), ensuring they don't generate harmful content. Their core method applies Clopper-Pearson confidence intervals and a novel algorithm that use…
LKValues: Aligning Large Language Models with Sri Lankan Societal Values
This paper introduces LKValues, a novel resource suite to address the Western bias in Large Language Model (LLM) value alignment. It contributes a survey-grounded set of 40 Sri Lankan societal values, an instruction corpus (LKvaluesIT) in Sinhala and English, …

Notes to Self: Can LLMs Benefit from Experiential Abstractions?
This paper investigates if Large Language Models (LLMs) can improve their problem-solving abilities by learning from their own past experiences, similar to how humans create reusab…
Solar Open 2 Technical Report
Solar Open 2 is a 250B-parameter Mixture-of-Experts model designed for long-horizon agentic tasks. Its core innovation is a novel 1M-token attention mechanism that interleaves soft…
Co-Evolving LLM Evaluators and Policies via DynamicRubric
This paper addresses the challenge of improving large language models (LLMs) when evaluator feedback on similar quality responses becomes less informative. The core method, Dynamic…
Statistical Inference for Rank Allocation in Low-Rank Adaptation
This paper introduces StatLoRA, a novel method for allocating rank in Low-Rank Adaptation (LoRA) for large language models. Instead of relying on heuristic importance scores, StatL…
Gotta Catch them all: the modes of Sycophancy
This paper challenges the view of sycophancy in LLMs as a single behavior. It identifies three distinct modes of sycophancy that, while producing similar outputs, have separable in…
The Town Square
Bento is a new tool that allows users to create, view, and collaborate on entire PowerPoint presentations within a single HTML file.
Workshops
RuView transforms commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection, eliminating the need for cameras.
This repository provides a coding agent skill designed to prevent information overload by delivering ADHD-friendly, concise answers, ensuring the core solution isn't buried.