Abstract
Model merging efficiently combines specialized large language models (LLMs) without joint retraining, but can substantially alter expert routing in Mixture-of-Experts (MoE) models. Such routing drift is often interpreted as routing failure, raising a fundamental question that remains unclear: does routing drift after MoE merging actually indicate routing failure, and what evidence should justify repair? We investigate these questions across DeepSeekMoE, OLMoE, and Qwen3-MoE proposing a routing analysis toolkit for controlled counterfactual interventions and token-level analysis. By crossing source and merged router inputs and parameters, we attribute most expert reassignments to input shifts rather than parameter changes at the same layer. However, source-relative routing differences poorly predict next-token likelihood gains from source-route restoration, and different expert selections can produce directionally similar mixture outputs. We therefore operationalize routing failure as task loss recoverable under a specified routing intervention, with non-routing parameters fixed. These tests detect recoverable loss under deliberate router corruption, whereas source-route restoration does not establish reliable task benefits in the evaluated merged models. Motivated by these, we propose Selective Router Repair (SRR) as a case study, and find that source-specialist token-likelihood advantages do not reliably identify beneficial local corrections. Together, these findings show that routing drift alone is insufficient evidence of routing failure: source-informed corrections must be judged by their task-level intervention effects. The analysis toolkit and SRR code are released.
Community
This work conducts a comprehensive analysis across s DeepSeekMoE, OLMoE, and Qwen3-MoE, by
proposing a routing analysis toolkit for controlled counterfactual interventions and
token-level analysis. These findings show that routing drift alone is insufficient
evidence of routing failure: source-informed corrections must be judged by their
task-level intervention effects. The analysis toolkit and SRR code are released at https://github.com/wyy-code/SRR.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- RAZOR: Pruning Replaceable Experts in LLMs (2026)
- RARE: Decoupling Representation Steering from Expert Routing in Mixture-of-Experts Language Models (2026)
- ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration (2026)
- Beyond Routing: Decoupling Expert Dispatch and Aggregation in Sparse Mixture-of-Experts (2026)
- Router Prior Bias: Preserving Base Routing Structure in MoE Post-Training (2026)
- UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models (2026)
- Task-Aware Spectral Pruning: A Mixture-of-Masks Framework for Efficient LLM Inference (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.32821 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper