When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles Paper • 2607.23379 • Published 18 days ago • 13
Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders Paper • 2606.10029 • Published Jun 8 • 12
Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders Paper • 2606.10029 • Published Jun 8 • 12
A Geometric Account of Activation Steering through Angle-Norm Decomposition Paper • 2606.06735 • Published Jun 4 • 27
Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders Paper • 2606.07473 • Published Jun 5 • 15
A Geometric Account of Activation Steering through Angle-Norm Decomposition Paper • 2606.06735 • Published Jun 4 • 27
Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders Paper • 2606.07473 • Published Jun 5 • 15
A Geometric Account of Activation Steering through Angle-Norm Decomposition Paper • 2606.06735 • Published Jun 4 • 27
Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders Paper • 2606.07473 • Published Jun 5 • 15