Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning Paper • 2609.03430 • Published 6 days ago • 170
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 6 days ago • 83
open-llm-leaderboard-old/details_xformAI__facebook-opt-125m-qcqa-ub-6-best-for-KV-cache Updated Jan 23, 2024 • 26 • 1
open-llm-leaderboard-old/details_saarvajanik__facebook-opt-6.7b-gqa-ub-16-best-for-KV-cache Updated Jan 28, 2024 • 24 • 1
saarvajanik/facebook-opt-6.7b-gqa-ub-16-best-for-KV-cache Text Generation • Updated Jan 28, 2024 • 68 • 1