Text-to-Speech
Safetensors
fish_qwen3_omni
instruction-following
multilingual

Portable C++/GGML implementation of Fish Audio S2 Pro. 3x faster than real time.

#27
by audio-cpp - opened

Measured on RTX 5090:

Warmed requests run about 3.1x-3.4x faster than real time. Longform (6000+ chars) runs about 3.3x faster than real time.

Try it at https://github.com/0xShug0/audio.cpp

Sign up or log in to comment