E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation Paper • 2608.30730 • Published 5 days ago • 15
Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses Paper • 2606.02373 • Published Jun 1 • 61
Beetle-HumanScale/beetle-bilingual-l2-50-sequential-33-67-b3-humanscale-nld-eng-seed42 Text Generation • 0.2B • Updated May 21 • 213 • 1
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence Paper • 2605.12882 • Published May 13 • 274