Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Paper • 2607.07508 • Published Jul 8 • 31
Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning Paper • 2012.13255 • Published Dec 22, 2020 • 6
Compacter: Efficient Low-Rank Hypercomplex Adapter Layers Paper • 2106.04647 • Published Jun 8, 2021 • 2
Training language models to follow instructions with human feedback Paper • 2203.02155 • Published Mar 4, 2022 • 26
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness Paper • 2205.14135 • Published May 27, 2022 • 16
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Paper • 2210.09261 • Published Oct 17, 2022 • 2
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers Paper • 2210.17323 • Published Oct 31, 2022 • 12
Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning Paper • 2303.10512 • Published Mar 18, 2023 • 5
Direct Preference Optimization: Your Language Model is Secretly a Reward Model Paper • 2305.18290 • Published May 29, 2023 • 71
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration Paper • 2306.00978 • Published Jun 1, 2023 • 15
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes Paper • 2306.13649 • Published Jun 23, 2023 • 38
Stack More Layers Differently: High-Rank Training Through Low-Rank Updates Paper • 2307.05695 • Published Jul 11, 2023 • 25
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning Paper • 2307.08691 • Published Jul 17, 2023 • 10
LoRA-FA: Memory-efficient Low-rank Adaptation for Large Language Models Fine-tuning Paper • 2308.03303 • Published Aug 7, 2023 • 4
LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models Paper • 2309.12307 • Published Sep 21, 2023 • 90
SWE-bench: Can Language Models Resolve Real-World GitHub Issues? Paper • 2310.06770 • Published Oct 10, 2023 • 13
A General Theoretical Paradigm to Understand Learning from Human Preferences Paper • 2310.12036 • Published Oct 18, 2023 • 20
Tied-Lora: Enhacing parameter efficiency of LoRA with weight tying Paper • 2311.09578 • Published Nov 16, 2023 • 17
A Rank Stabilization Scaling Factor for Fine-Tuning with LoRA Paper • 2312.03732 • Published Nov 28, 2023 • 13
LoRAMoE: Revolutionizing Mixture of Experts for Maintaining World Knowledge in Language Model Alignment Paper • 2312.09979 • Published Dec 15, 2023 • 3
KTO: Model Alignment as Prospect Theoretic Optimization Paper • 2402.01306 • Published Feb 2, 2024 • 24
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models Paper • 2402.03300 • Published Feb 5, 2024 • 149
DistiLLM: Towards Streamlined Distillation for Large Language Models Paper • 2402.03898 • Published Feb 6, 2024 • 5
GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection Paper • 2403.03507 • Published Mar 6, 2024 • 192
ORPO: Monolithic Preference Optimization without Reference Model Paper • 2403.07691 • Published Mar 12, 2024 • 74
PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language Models Paper • 2404.02948 • Published Apr 3, 2024 • 5
SimPO: Simple Preference Optimization with a Reference-Free Reward Paper • 2405.14734 • Published May 23, 2024 • 13
LoRA-XS: Low-Rank Adaptation with Extremely Small Number of Parameters Paper • 2405.17604 • Published May 27, 2024 • 4
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Paper • 2406.01574 • Published Jun 3, 2024 • 57
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision Paper • 2407.08608 • Published Jul 11, 2024 • 2
On the Impact of Fine-Tuning on Chain-of-Thought Reasoning Paper • 2411.15382 • Published Nov 22, 2024 • 1
Kimi k1.5: Scaling Reinforcement Learning with LLMs Paper • 2501.12599 • Published Jan 22, 2025 • 132
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning Paper • 2501.12948 • Published Jan 22, 2025 • 463
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training Paper • 2501.17161 • Published Jan 28, 2025 • 127
How Much Knowledge Can You Pack into a LoRA Adapter without Harming LLM? Paper • 2502.14502 • Published Feb 20, 2025 • 93
DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs Paper • 2503.07067 • Published Mar 10, 2025 • 34
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Paper • 2503.14476 • Published Mar 18, 2025 • 148
SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild Paper • 2503.18892 • Published Mar 24, 2025 • 32
Understanding R1-Zero-Like Training: A Critical Perspective Paper • 2503.20783 • Published Mar 26, 2025 • 61
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Paper • 2504.05118 • Published Apr 7, 2025 • 27
HD-PiSSA: High-Rank Distributed Orthogonal Adaptation Paper • 2505.18777 • Published May 24, 2025 • 2
Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training Paper • 2507.05386 • Published Jul 7, 2025 • 2
InfiR2: A Comprehensive FP8 Training Recipe for Reasoning-Enhanced Language Models Paper • 2509.22536 • Published Sep 26, 2025 • 3
PEFT-Bench: A Parameter-Efficient Fine-Tuning Methods Benchmark Paper • 2511.21285 • Published Nov 26, 2025 • 2
TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummies Paper • 2511.23225 • Published Nov 28, 2025 • 3
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Paper • 2601.11868 • Published Jan 17 • 38
Jet-RL: Enabling On-Policy FP8 Reinforcement Learning with Unified Training and Rollout Precision Flow Paper • 2601.14243 • Published Jan 20 • 25
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning Paper • 2601.18150 • Published Jan 26 • 11
FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling Paper • 2603.05451 • Published Mar 5 • 2
Scaling DoRA: High-Rank Adaptation via Factored Norms and Fused Kernels Paper • 2603.22276 • Published Mar 23 • 15
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence Paper • 2606.19348 • Published Apr 26 • 40
Fara-1.5: Scalable Learning Environments for Computer Use Agents Paper • 2606.20785 • Published Jun 18 • 9
MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training Paper • 2606.30406 • Published Jun 29 • 25
Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning Paper • 2607.08393 • Published Jul 9 • 19
Antares: Foundation Models for Agentic Vulnerability Localization Paper • 2608.02407 • Published 25 days ago • 3
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 4 days ago • 197
WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report Paper • 2608.24053 • Published 3 days ago • 62
Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence Paper • 2608.21156 • Published 7 days ago • 54
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces Paper • 2608.23041 • Published 4 days ago • 58