Toolcompass: Guiding Tool Trialing, Not Suppressing It Paper • 2609.25678 • Published 19 days ago • 2
VLMGuard: Bootstrapping Malicious Prompt Detectors from Unlabeled Vision-Language Prompts in the Wild Paper • 2410.00296 • Published Jul 4 • 5
OpenHalDet: A Unified Benchmark for Hallucination Detection across Diverse Generation Scenarios Paper • 2606.06959 • Published Jun 5 • 1
Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection in Large Reasoning Models Paper • 2607.22098 • Published Jul 24 • 9
Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection in Large Reasoning Models Paper • 2607.22098 • Published Jul 24 • 9
Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents Paper • 2606.26080 • Published Jun 24 • 13
MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent Systems Paper • 2505.18943 • Published May 25, 2025 • 24
Feed Two Birds with One Scone: Exploiting Wild Data for Both Out-of-Distribution Generalization and Detection Paper • 2306.09158 • Published Jun 15, 2023
OpenOOD v1.5: Enhanced Benchmark for Out-of-Distribution Detection Paper • 2306.09301 • Published Jun 15, 2023 • 1
How Does Unlabeled Data Provably Help Out-of-Distribution Detection? Paper • 2402.03502 • Published Feb 5, 2024
VLMGuard: Bootstrapping Malicious Prompt Detectors from Unlabeled Vision-Language Prompts in the Wild Paper • 2410.00296 • Published Jul 4 • 5
HaloScope: Harnessing Unlabeled LLM Generations for Hallucination Detection Paper • 2409.17504 • Published Sep 26, 2024