Fish Audio S2 Pro rows in a multilingual voice-cloning benchmark
I included Fish Audio S2 Pro fp16 in a local voice-cloning benchmark across English, German, Modern Standard Arabic, Spanish, and Mandarin Chinese:
https://www.soniqo.audio/blog/voice-cloning-benchmarks
Fish Audio S2 Pro showed strong German and Arabic speaker similarity in this run, though it was slower than OmniVoice and Chatterbox on these rows.
The benchmark uses Google FLEURS references and includes reference audio, generated audio, speaker similarity, WER/CER, generated audio length, and RTF for each row.
This is an engineering benchmark rather than a MOS study, but I wanted to share the Fish Audio rows with the model community.
Thank you for your benchmark report! We will further enhance these languages' performance and try to lower the RTF.