AutoDataBench: A Data-centric Testbed for Accelerating Auto Research Paper • 2609.40097 • Published 3 days ago • 10
Jev thinks "I don't know'', but doesn't say it: Introducing Sys1Cal-v1 Dataset for Probability Calibration Paper • 2609.35342 • Published 5 days ago • 6
Raven: The Harness of Harnesses for Composable Agentic Intelligence Paper • 2609.33439 • Published 6 days ago • 510
SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale Paper • 2609.38822 • Published 3 days ago • 4
Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents Paper • 2609.27334 • Published 10 days ago • 52
PUBG Ally: A Conversational Embodied Agent as an AI Teammate Paper • 2609.29837 • Published 9 days ago • 25
Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures Paper • 2609.29429 • Published 9 days ago • 23
Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs Paper • 2609.29845 • Published 9 days ago • 100
Jev in the Wild: A Data-Driven Analysis of the Jev Model's Functionality, Applications and Ecosystem Paper • 2609.30216 • Published 9 days ago • 15
SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL Paper • 2609.29050 • Published 9 days ago • 13
Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents Paper • 2609.17653 • Published 18 days ago • 45
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 16 days ago • 110
Grounded Skill Synthesis from Code at Scale for Agentic Intelligence Paper • 2609.05571 • Published 29 days ago • 119