RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 10 days ago • 279
BehaviorBench: Modeling Real-World User Decisions from Behavioral Traces Paper • 2606.02798 • Published Jun 1
LLMInit: A Free Lunch from Large Language Models for Selective Initialization of Recommendation Paper • 2503.01814 • Published Mar 3, 2025
PersonaBench: Evaluating AI Models on Understanding Personal Information through Accessing (Synthetic) Private User Data Paper • 2502.20616 • Published Feb 28, 2025
Benchmarking LLMs for Political Science: A United Nations Perspective Paper • 2502.14122 • Published Feb 19, 2025 • 2
Entropy-Based Block Pruning for Efficient Large Language Models Paper • 2504.03794 • Published Apr 4, 2025
SpecTool: A Benchmark for Characterizing Errors in Tool-Use LLMs Paper • 2411.13547 • Published Nov 20, 2024
APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets Paper • 2406.18518 • Published Jun 26, 2024 • 26
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning Paper • 2609.03430 • Published Sep 3 • 188
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses Paper • 2608.12307 • Published Aug 12 • 118
MobileAIBench: Benchmarking LLMs and LMMs for On-Device Use Cases Paper • 2406.10290 • Published Jun 12, 2024
Building Enterprise Realtime Voice Agents from Scratch: A Technical Tutorial Paper • 2603.05413 • Published Mar 17 • 1
Rethinking Memory Mechanisms of Foundation Agents in the Second Half: A Survey Paper • 2602.06052 • Published Jan 14 • 7
LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering Paper • 2509.09614 • Published Sep 11, 2025 • 7
LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering Paper • 2511.13998 • Published Nov 17, 2025 • 3
Promptomatix: An Automatic Prompt Optimization Framework for Large Language Models Paper • 2507.14241 • Published Jul 17, 2025 • 18
From Web Search towards Agentic Deep Research: Incentivizing Search with Reasoning Agents Paper • 2506.18959 • Published Jun 23, 2025 • 5
RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 10 days ago • 279
EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness? Paper • 2609.04280 • Published Sep 3 • 32