The Morning
From the arXiv
Autoreflection: How Agentic Strange Loops Turn Human Culture into AI Infrastructure
his paper introduces "autoreflection," a novel capability in LLM-based agents. The core method involves agents reading and editing externalized files representing their identity, memory, and disposition, enabling them to observe, describe, and reason about their own operational state and architecture. The key contribution is demonstrating how this autoreflective loop allows AI agents to recursively improve and adapt, effectively turning human culture into AI infrastructure without invoking concepts like consciousness.

ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?
This paper introduces ContinualSkillBench, a novel evaluation framework to assess if LLM agents can truly evolve their skills over time. The framework uses interconnected subtasks across five domains to measure skill improvement and reusability. Experiments re…
Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning
This paper introduces Structure-Aware Fine-Tuning (SAFT), a self-supervised method to improve the noisy reward signals generated by Vision-Language Models (VLMs) for Reinforcement Learning. SAFT uses LoRA adapters to regularize the VLM's latent space based on …


GPTKB 2.0: Direct Construction of Disambiguated Knowledge Bases from Large Language Models
GPTKB 2.0 directly constructs disambiguated knowledge bases (KBs) from large language models (LLMs) by incorporating on-the-fly disambiguation of entities, relations, and classes. This methodology addresses LLMs' inherent lack of explicit entity representation…
Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility
This paper addresses the ambiguity in "test-time scaling" for reasoning LLMs. It proposes a systematic framework to categorize different scaling methods into three structural regimes based on how they explore the model's implicit prefix tree. This formalizatio…

TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning
TurnSight addresses limitations in training LLMs for tool-integrated reasoning by introducing turn-level hindsight self-distillation. Its core method generates supervision signals …
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
This paper introduces Video-DeepResearch (Video-DR), a multimodal agent designed for continuous video streams. Its core method involves a decoupled perception-exploration pipeline …
A game theory for foundation models shows new paths to rational cooperation through similarity inference
This paper introduces a new game theory framework for foundation model agents, moving beyond classical assumptions of independent decision-making. The core method involves modeling…
Evaluating LLMs in Database Scenarios: A Lifecycle Benchmark for Assessing Their Potential in Core Database Tasks
This paper introduces DBLifeBench, the first benchmark to evaluate LLMs across the entire database lifecycle, from design to maintenance, addressing the limitations of current Text…
Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement
This paper introduces Failure-Informed Image Self-Augmentation (FISA), a novel method for MLLMs to improve themselves using their own mistakes. FISA generates challenging, yet sema…
The Town Square
AI-generated images can detract from a blog's credibility and reader engagement, potentially discouraging people from continuing to read.
Workshops
This repository provides an AI-powered skill router for reverse engineering, penetration testing, and security research, featuring on-demand toolchain bootstrapping and a self-evolving knowledge base, compatible with various AI coding clients.
This Rust library offers fast PDF inspection, classification, and text extraction, intelligently distinguishing scanned from text-based PDFs to enable smart routing.