Continual Learning Mechanisms Compose for Long-Horizon Memorization Paper • 2609.06986 • Published 11 days ago • 295
FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow Estimation Paper • 2609.11486 • Published 8 days ago • 34
TempCloze: Can Video-LLMs Identify the Missing Middle? Paper • 2609.01515 • Published 17 days ago • 32
Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation Paper • 2609.06373 • Published 12 days ago • 22
Encoded Early, Used Late: Where Transformers Begin to Act on an Inferred Partner's Expertise Paper • 2609.07139 • Published 11 days ago • 17
LatentPress: Context Compression Beyond Text and Vision Paper • 2609.01507 • Published 17 days ago • 121
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published Aug 18 • 159
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published Aug 14 • 283
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers Paper • 2608.06867 • Published Aug 7 • 113
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published Aug 1 • 264
The Next Screenshot Knows: Gated Hindsight Distillation for Mobile GUI Agents Paper • 2608.06065 • Published Aug 6 • 8
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Paper • 2607.28609 • Published Jul 30 • 74
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Paper • 2608.05747 • Published Aug 6 • 47
NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap Paper • 2608.04397 • Published Aug 5 • 24
MiniWorld: Democratizing the Training of Video World Models from Scratch Paper • 2608.01127 • Published Aug 2 • 20
ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction Paper • 2607.29677 • Published Jul 31 • 26