Base Models Can Reason By Taking a Cue From Training Data Paper • 2610.06851 • Published 4 days ago • 21
When to Switch: Reliable Action-Chunk Extension for Vision-Language-Action Models Paper • 2610.05719 • Published 4 days ago • 24
Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems Paper • 2609.39050 • Published 9 days ago • 17
Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding Paper • 2609.32019 • Published 14 days ago • 62
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 13 days ago • 325
Raven: The Harness of Harnesses for Composable Agentic Intelligence Paper • 2609.33439 • Published 12 days ago • 569
Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation Paper • 2609.35347 • Published 11 days ago • 180
Disaggregated Quantization: Specializing LLM Prefill and Decode Paper • 2609.26333 • Published 17 days ago • 93
Parts-of-Speech as Emergent Categories in SAE Latent Space Paper • 2609.29362 • Published 15 days ago • 19
Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs Paper • 2609.29845 • Published 15 days ago • 105
Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents Paper • 2609.29892 • Published 15 days ago • 34
Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures Paper • 2609.29429 • Published 15 days ago • 29
Knowledge Pull Requests for Continual Document Authoring Paper • 2609.26634 • Published 17 days ago • 12
LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay Paper • 2609.25053 • Published Sep 7 • 17
The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks Paper • 2609.25804 • Published 17 days ago • 164