Beyond Solver Verdicts: Generative Reward Models for Autoformalization Paper • 2609.11085 • Published 7 days ago • 34
A Neurosymbolic Approach to Natural Language Formalization and Verification Paper • 2511.09008 • Published Nov 12, 2025
TurboFuzzLLM: Turbocharging Mutation-based Fuzzing for Effectively Jailbreaking Large Language Models in Practice Paper • 2502.18504 • Published Feb 21, 2025 • 2
Zero-knowledge LLM hallucination detection and mitigation through fine-grained cross-model consistency Paper • 2508.14314 • Published Aug 19, 2025 • 2
Zero-knowledge LLM hallucination detection and mitigation through fine-grained cross-model consistency Paper • 2508.14314 • Published Aug 19, 2025 • 2
TurboFuzzLLM: Turbocharging Mutation-based Fuzzing for Effectively Jailbreaking Large Language Models in Practice Paper • 2502.18504 • Published Feb 21, 2025 • 2
StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean Paper • 2609.09264 • Published 9 days ago • 9
VeriCoT: Neuro-symbolic Chain-of-Thought Validation via Logical Consistency Checks Paper • 2511.04662 • Published Nov 6, 2025 • 37