WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing Paper • 2609.20423 • Published 4 days ago • 35
Beyond Solver Verdicts: Generative Reward Models for Autoformalization Paper • 2609.11085 • Published 11 days ago • 34
Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space Paper • 2608.29188 • Published 23 days ago • 11
Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training Paper • 2608.26730 • Published 25 days ago • 154
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published Aug 18 • 159
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published Aug 14 • 283
TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement Paper • 2608.11951 • Published Aug 12 • 9
Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence Paper • 2608.12743 • Published Aug 13 • 44
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published Aug 1 • 264
The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads Paper • 2608.04570 • Published Aug 5 • 41
When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills Paper • 2608.03700 • Published Aug 4 • 9