SpanCalib-VLM: Calibrated Hallucination Span Detection in Vision-Language Models Paper • 2608.29974 • Published 12 days ago • 5
Human-in-the-Loop Signature Bootstrapping for UAV Hyperspectral PFM-1 Mine Detection Paper • 2607.25310 • Published Jul 28 • 5
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Paper • 2607.25659 • Published Jul 28 • 84
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Paper • 2607.20911 • Published Jul 23 • 26
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model Paper • 2607.17977 • Published Jul 20 • 199
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Paper • 2607.14952 • Published Jul 16 • 213
Metacognition in LLMs: Foundations, Progress, and Opportunities Paper • 2607.11881 • Published Jul 13 • 31
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models Paper • 2607.12463 • Published Jul 14 • 109
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published Jul 13 • 87
OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers Paper • 2607.04033 • Published Jul 4 • 78