Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models Paper • 2609.26637 • Published 3 days ago • 12 • 2
X-Planner: Event-Structured Task Planning for Embodied Intelligence Paper • 2609.25187 • Published 4 days ago • 1 • 2
Calibration as a First-Class Criterion in LLM Evaluation Paper • 2609.26489 • Published 3 days ago • 2 • 2
StudentBench: AI and human tutoring yield equivalent GRE learning gains Paper • 2609.28470 • Published 2 days ago • 3 • 6
The Past Frames the Future: Memory for Autoregressive Video Generation Paper • 2609.28466 • Published 2 days ago • 36 • 2
All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation Paper • 2609.27901 • Published 2 days ago • 5 • 2
WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents Paper • 2609.27490 • Published 2 days ago • 8 • 2
PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing Paper • 2609.23784 • Published 5 days ago • 10 • 2
RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling Paper • 2609.22947 • Published 6 days ago • 21 • 2
GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression Paper • 2609.25963 • Published 3 days ago • 11 • 2
Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents Paper • 2609.27334 • Published 2 days ago • 34 • 2
Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery Paper • 2609.27980 • Published 2 days ago • 3 • 3
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms Paper • 2609.27321 • Published 2 days ago • 8 • 1
Self-Organizing Agent Teams Learn to Reason Together Paper • 2609.22682 • Published 6 days ago • 4 • 2
Knowledge Pull Requests for Continual Document Authoring Paper • 2609.26634 • Published 3 days ago • 2 • 2
InternW0: A Foundational Physical World Model for Efficient Real-World Interactions Paper • 2609.27656 • Published 2 days ago • 5 • 1
The Linear Representation Hypothesis Needs a Group Action Paper • 2609.27158 • Published 3 days ago • 2 • 2
SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue Paper • 2609.26780 • Published 3 days ago • 82 • 5