Instructions to use VoiceHub/DACFlow-EN-10k with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- AudioSeal
How to use VoiceHub/DACFlow-EN-10k with AudioSeal:
# Watermark Generator from audioseal import AudioSeal model = AudioSeal.load_generator("VoiceHub/DACFlow-EN-10k") # pass a tensor (tensor_wav) of shape (batch, channels, samples) and a sample rate wav, sr = tensor_wav, 16000 watermark = model.get_watermark(wav, sr) watermarked_audio = wav + watermark# Watermark Detector from audioseal import AudioSeal detector = AudioSeal.load_detector("VoiceHub/DACFlow-EN-10k") result, message = detector.detect_watermark(watermarked_audio, sr) - Notebooks
- Google Colab
- Kaggle
DACFlow-EN-10k: English text-to-speech that copies a voice
DACFlow-EN-10k reads English text aloud in a voice you choose. Give it a short recording of a voice (3-10 seconds, with the
exact words spoken in it) and any English text, and it speaks that text in that voice. This is zero-shot voice cloning:
the voice does not have to be in the training data. The output is 48 kHz audio.
The model has 206 M parameters and is trained from scratch only on the provided data: 9,543 hours of
synthetic English speech (3,400,911 clips, 3,587 voices) from SynDataLab-EN/echo-clones-4m-en (the curated part of its
~11,150 hours), nothing else. A hobby project; the weights are apache-2.0.
The name: DAC for the Semantic-DACVAE audio codec whose latents it generates, Flow for flow matching, EN for
English, 10k for its training data, about ten thousand hours of speech.
The code: kadirnar/dacvae-next, Python package mytts.
Status: training is finished. The model is 100k base + EN-45 quality post-training (300 updates): the EMA weights saved at step 100k of the main run, then a short quality post-training of 300 updates (Training). The main run was planned for 200k steps and stopped at step 194,271 on 2026-10-04; this page no longer changes by itself. Why this model: against step 100k alone, both with each prompt's own bandwidth as the condition (auto), it sounds more natural to the automatic raters (UTMOS +0.048 on echo-dev v2, +0.058 on seed-dev), copies real voices a little better (seed-dev SIM-o +0.010) and makes about as many word errors (WER +0.00 points on seed-dev, -0.02 on echo-dev v2). It passed a rule fixed before its evaluation: higher UTMOS and DNSMOS on echo-dev v2 (95 % intervals above 0), word error rate not worse on either set, seed-dev SIM-o lower by at most 0.01. Why step 100k as the base: later checkpoints sound a little better on held-out voices of the training data's kind, but copy real people's voices much worse. With the default bandwidth condition at both steps (auto was not measured at step 190k), speaker similarity on real voices (seed-dev SIM-o) falls from 0.50 at step 100k to 0.22 at step 190k, where 51 % of the outputs score below 0.2 (5.7 % at step 100k). Step 100k with each prompt's own bandwidth as the condition (auto) is the best balance; with that condition, it scores seed-dev SIM-o 0.526, with 2.4 % of its outputs below 0.2. Details: #177 (the base), #69 (the post-training).
On this page: Listen · Results · How to use · Decoder versions · Training · Data · Licence
Listen
The model, 100k base + EN-45 quality post-training (300 updates), reads five texts. Each text is spoken in a different voice, copied from the short voice prompt next to it. The model never heard these voices in training (held-out voices of the provided data). Each output is the best of 8 takes, picked by automatic scores (below).
| What the model reads | Voice prompt (the input) | Model output |
|---|---|---|
| Short sentence I left my umbrella at the office again, so I'm definitely getting soaked on the way home. | A higher voice (~202 Hz). It says: “I keep telling myself I just need to power through, like, push a little harder, and then it'll be fine. But it's not getting fine. It's getting— it's getting worse.” | Best of 8 takes (seed 1): every word right · band 15.6 kHz · UTMOS 4.43 · DNSMOS 3.48 |
| Question Have you ever noticed that the quietest person in the room usually has the most interesting story to tell? | A lower voice (~119 Hz). It says: “Look, I'm not saying it's easy, but you always pull it together. Just take it step by step, you know?” | Best of 8 takes (seed 4): every word right · band 15.2 kHz · UTMOS 4.34 · DNSMOS 3.38 |
| Numbers, dates, abbreviations Dr. Patel moved my appointment to Tuesday, March 3rd, at 4:15 p.m., and the co-pay went up from $20 to $35. | A higher voice (~201 Hz). It says: “Would you just— I know you mean well, but every time you mention it I feel like an idiot. It's probably nothing.” | Best of 8 takes (seed 0): band 15.2 kHz · UTMOS 4.42 · DNSMOS 3.44 · Whisper heard: “Dr. Patel moved my appointment to Tuesday, March 3rd at 4.15 PM and the copay went up from $20 to $35.” |
| Conversation (~10 s) So I finally tried that new ramen place downtown, and honestly, it was worth the wait. The broth was amazing, but next time I'm definitely skipping the extra spicy option. | A lower voice (~115 Hz). It says: “I'm sorry, but the subway doors closed on my bag. Again.” | Best of 8 takes (seed 4): every word right · band 16.0 kHz · UTMOS 4.47 · DNSMOS 3.42 |
| Long passage (~20 s) When the storm passed, the whole neighborhood came outside to look at the damage, and although a few fences had fallen and the old oak tree had lost its biggest branch, everyone was relieved that nobody was hurt. They spent the rest of the afternoon clearing the street and sharing whatever food they had left. | The deepest voice (~102 Hz). It says: “I mean, come on, it's not like I woke up and decided to have the worst day ever. Things just went wrong one after another, like dominoes!” | Best of 8 takes (seed 3): every word right · band 14.5 kHz · UTMOS 4.42 · DNSMOS 3.41 |
How the samples are made: each text was read 8 times (seeds 0 to 7) with the settings of How to use (32 Euler steps + sway, joint CFG w 4.0, initial noise 0.9, at most -16 LUFS (peak-limited), duration by the text rule (duration_mode text), each prompt's own bandwidth as the condition (auto, clamped to 9,798-16,839 Hz)), and the page plays the best take of each, picked by automatic scores, not by ear: first the takes in which Whisper-large-v3 hears every word right (WER 0), then those whose sound reaches 13 kHz (not muffled), then the fewest word errors, then the highest predicted naturalness (UTMOS + DNSMOS OVRL). So these are the model's better takes; a single take can sound worse. Every output carries an inaudible AudioSeal watermark, checked before upload; the prompts are the original recordings.
Results
How well the model, 100k base + EN-45 quality post-training (300 updates), does on voices and sentences it never saw in training (lower WER is better, higher SIM-o and UTMOS are better). The settings are those of How to use: each prompt's own bandwidth as the condition (auto, clamped to 9,798-16,839 Hz), except the length rule: these numbers use the default band rule, the samples and How to use the text rule (below).
| Test set | WER ↓ | SIM-o ↑ | UTMOS ↑ |
|---|---|---|---|
| seed-dev: real people's voices (545 sentences, two takes each) | 1.42 % | 0.536 | 3.86 (real speech: 3.52) |
| echo-dev v2: held-out voices of the training data's kind (198 voices, 1,000 sentences, one take each) | 0.50 % | 0.820 | 4.04 |
1.7 % of the seed-dev outputs score SIM-o below 0.2 (the voice is not copied). These numbers average every take of the test sets (nothing is picked, unlike the samples above): they show the model as it is.
- WER (word error rate): Whisper-large-v3 writes down what it hears and this is compared with the text; 2 % means about one word in fifty is wrong.
- SIM-o (speaker similarity, 0 to 1): how close the output's voice is to the prompt's (WavLM-large + ECAPA-TDNN).
- UTMOS (1 to 5): a predicted listener rating of naturalness (UTMOS22). In brackets: the score of real recordings of the same set, the practical ceiling.
- seed-dev: real people's voices (the dev half of Seed-TTS test-en, Common Voice recordings); harder, since the model learned only from synthetic voices. echo-dev v2: held-out voices of the training data's kind (synthetic EchoTTS voices, none of them trained on), the set the model was chosen on.
The length rule (duration_mode="text", PR #185): the samples and How to
use set the length of each output from its text (letters, word gaps and pause punctuation, counted by a small model fitted on the
training data) and the voice prompt's pace, instead of the prompt's pace alone (the band rule, which reads a little too fast). It
was measured on step 100k with the same settings, not on this model: against the band rule it changed seed-dev by WER +0.05
points, SIM-o +0.002, UTMOS -0.000 (95 % CI -0.014 to +0.013) and echo-dev v2 by WER +0.03 points, SIM-o +0.001, UTMOS -0.012 (95
% CI -0.025 to +0.000). On 50 items that real recordings of the same voices read with the same texts, the speech came closer to
those recordings: 2.88 words per second (band rule 3.30, real 3.05), 2.32 pauses per item (band rule 1.86, real 2.24), an F0
spread (semitones) of 3.78 (band rule 3.52, real 3.83).
How to use
1. Install (Python 3.10 or newer; a CUDA GPU is recommended):
git clone -b roadmap/en-echo https://github.com/kadirnar/dacvae-next
cd dacvae-next
git checkout 7af9fac # the code that made the samples above
pip install -e ".[codec]"
2. Copy a voice in Python. You need a short recording of the voice (3-10 s, clean, one speaker) and the exact words spoken in it:
import soundfile as sf
from huggingface_hub import snapshot_download
from mytts.flow.sampler import SamplerConfig
from mytts.infer import Synthesizer
ckpt = "checkpoints/grow_q_u300" # the model: 100k base + EN-45 quality post-training (300 updates)
local = snapshot_download("VoiceHub/DACFlow-EN-10k", allow_patterns=[f"{ckpt}/*"]) # downloads only this checkpoint
tts = Synthesizer.from_export(f"{local}/{ckpt}", device="cuda", bandwidth_range=(9797.6, 16839.0), duration_mode="text")
sc = SamplerConfig(steps=32, cfg_mode="joint", cfg_w=4.0, noise_scale=0.9) # the settings of the samples above
wav, sr = tts.synthesize(
"Any English text you like.",
lang="en",
sc=sc,
prompt_audio="my_voice.wav", # the voice to copy
prompt_text="The exact words spoken in my_voice.wav.",
bandwidth_hz="auto", # condition on the prompt's own bandwidth (clamped to the range above)
out_lufs=-16.0,
seed=0,
)
sf.write("output.wav", wav, sr) # 48 kHz, with the AudioSeal watermark
Or from the command line:
hf download VoiceHub/DACFlow-EN-10k --include "checkpoints/grow_q_u300/*" --local-dir DACFlow-EN-10k
python scripts/synthesize.py --model DACFlow-EN-10k/checkpoints/grow_q_u300 --bandwidth auto --bandwidth-range 9797.6 16839.0 --duration-mode text \
--text "Any English text you like." --prompt-audio my_voice.wav \
--prompt-text "The exact words spoken in my_voice.wav." --out output.wav
auto measures the voice prompt's frequency band and asks the model for an output of the same band (a dull prompt
gives a dull output).
The range is the 5th to 95th percentile of the training data's band; it is passed explicitly because a released
config.json records only the top of it. auto needs the code of PR #180 or later (the commit above has it).
duration_mode="text" (--duration-mode text) sets the length of the output from the text and the prompt's pace (see
Results). It needs the code of PR #185 or later (the commit above has it).
Tips: a clean prompt with an exact transcript works best. A long text is split at sentence ends and read piece by piece.
python scripts/watermark_check.py detect output.wav checks the watermark.
Outputs vary with the seed: like the samples above (the best of 8 seeds), try a few (seed=0, 1, 2, ...) and keep the best.
Decoder versions
The model makes audio latents; the decoder of the Semantic-DACVAE codec turns them into sound. Everything above uses the codec's original decoder (Aratako/Semantic-DACVAE-Japanese, MIT). This repo also holds 3 fine-tuned versions of that decoder, fitted to this model's outputs (100k base + EN-45 quality post-training (300 updates); PR #184). Only the decoder was trained; the encoder was frozen, so the latents, the model and the voice prompts are unchanged, and any version can be swapped in when the model is used.
Training data: 20,000 pairs (39.3 hours, 3,581 voices) made from clips of the training split of
SynDataLab-EN/echo-clones-4m-en (the catalog echo_en_v1; 413
held-out and test voices excluded), the provided data only. For each clip the model regenerated its latents from a partly noised
copy of them (sdedit starts 0.5 and 0.6), and the decoder learned to turn those model-made latents into the clip's real audio. No
other audio and no other model's outputs.
How they compare. The model's outputs for echo-dev v2 (1,000 utterances) and seed-dev (545) were generated once, and the same latents were decoded by every decoder, so only the decoder differs (one take each, the band length rule: not the evaluation of Results). The first row is the original decoder; the others are the change against it (paired; n.s. = the 95 % interval includes 0). HNR and shimmer: Praat voice-quality measures on 1,000 echo-dev v2 outputs (higher HNR and lower shimmer = smoother, less rough voicing).
| Decoder | Folder | Updates | UTMOS echo-dev v2 | UTMOS seed-dev | SIM-o echo-dev v2 | SIM-o seed-dev | WER echo-dev v2 | WER seed-dev | HNR | Shimmer |
|---|---|---|---|---|---|---|---|---|---|---|
| original | (in the base codec) | - | 4.040 | 3.855 | 0.820 | 0.540 | 0.47 % | 1.39 % | 10.54 dB | 0.1029 |
| ft_A, 5k updates | decoders/ft_A_u300_step_005000 |
5,000 | +0.055 | +0.061 | +0.003 | -0.006 | +0.03 (n.s.) | -0.07 | +0.57 dB | -0.0048 |
| ft_A, 10k updates | decoders/ft_A_u300_step_010000 |
10,000 | +0.065 | +0.065 | +0.004 | -0.010 | +0.08 (n.s.) | -0.05 (n.s.) | +0.65 dB | -0.0057 |
| ft_A, 20k updates (default) | decoders/ft_A_u300_step_020000 |
20,000 | +0.077 | +0.079 | +0.006 | -0.010 | +0.05 (n.s.) | -0.05 (n.s.) | +0.76 dB | -0.0068 |
The fine-tuned decoders sound more natural to the automatic rater (UTMOS +0.055 to +0.077 on echo-dev v2, +0.061 to +0.079 on seed-dev) with smoother voicing, with no significant increase in corpus word error rate (the significant change is seed-dev, ft_A, 5k updates: -0.07 points, fewer errors), although every version has slightly more word errors on echo-dev v2 (not significant). The cost: they copy real people's voices slightly less well (seed-dev SIM-o -0.010 to -0.006, significant), while the similarity on echo-dev v2 rises (+0.003 to +0.006). More updates move UTMOS, SIM-o and the voicing measures further in the same direction. The samples above use the original decoder.
To use a version, load it as the model's codec (the rest of How to use stays the same):
ckpt, dec = "checkpoints/grow_q_u300", "decoders/ft_A_u300_step_020000"
local = snapshot_download("VoiceHub/DACFlow-EN-10k", allow_patterns=[f"{ckpt}/*", f"{dec}/*"])
tts = Synthesizer.from_export(f"{local}/{ckpt}", device="cuda", codec=f"{local}/{dec}", bandwidth_range=(9797.6, 16839.0), duration_mode="text")
or --codec decoders/ft_A_u300_step_020000 (the downloaded folder) on the command line. Loading a decoder folder needs the code
of PR #187 or later (git checkout roadmap/en-echo if you checked out an
older commit above). Each folder holds decoder.safetensors (the decoder half only, pickle-free), config.json (the base codec,
the DACVAE settings, the training record and the sha256 of the encoder and decoder halves) and a README. The encoder comes from
the base codec's own weights, which the loader downloads and checks by hash. The weights are a fine-tune of the base codec's
MIT-licensed decoder (itself a fine-tune of facebook/dacvae-watermarked).
Training
- Model: the
basepreset, 206 M parameters: a diffusion transformer trained with flow matching. It generates the latents of the Semantic-DACVAE audio codec (48 kHz, 25 frames per second) from the text, continuing the voice prompt in context (no separate speaker encoder). - Data: the curated training catalog
echo_en_v1of the provided data (EN-10, #34 curation; see Data). - Recipe: planned 200k steps on one RTX 5090, about 27 minutes of speech per step (40,000 latent frames); learning rate 2.5e-4 after 5k warm-up steps, held, then lowered over the last 20 % of the steps (from step 160k). Training was stopped at step 194,271 on 2026-10-04. The model's base is the EMA export of step 100k, before the learning rate was lowered. The settings were chosen with small screening runs: the flow-matching noise schedule t_mean -0.8 / t_std 0.8 (EN-55, #100) and cross-utterance voice prompts with p_cross 0.6 (EN-29, #53).
- Quality post-training (#69): 300 updates starting from the step-100k export. GROW, an online reinforcement-learning method for flow-matching models: for each of 16 prompts per update the model reads the text 16 times, each reading is scored (intelligibility by a CTC speech recognizer, voice similarity by a speaker-verification model, predicted naturalness by UTMOS), and the model is moved toward its better readings, with a penalty that keeps it close to the base (learning rate 2e-6). The prompts are 25,000 texts read by 320 voices of the training data (its transcripts and voices plus templated hard sentences with numbers, names and abbreviations): no new audio, only the model's own readings and the scores of pretrained scoring models. The round was planned for 500 updates and stopped at update 304 when the machine restarted; of the saved updates (100, 200, 300), the one with the lowest word error rate on a held-out dev list, update 300, was kept, as the rule fixed before the round said.
- Code and history: the recipe configs/train/en_full.yaml (PR #154, schedule PR #156); the run EN-34, #58; the data processing EN-12, #36; the whole English plan EN-01, #78.
Data
- The training data, tokenized: VoiceHub/DACFlow-EN-10k-data. Exactly what the model trains on: every clip of
the training catalog
echo_en_v1as codec tokens (Semantic-DACVAE latents, 25 frames per second) with its transcript, and the catalog itself (the exact list of training clips). No audio: the codec turns the tokens back into sound. With it (and the speech-feature targets in the backup below) the training can be repeated without encoding the 1.5 TB of audio again. - The full working backup: VoiceHub/DACFlow-EN-10k-backup: everything else the training and its evaluation use (speech-feature (REPA) targets, annotations, evaluation sets, experiments), kept to resume them; not needed to use the model.
- The source audio: SynDataLab-EN/echo-clones-4m-en: about 4 million synthetic English clips (EchoTTS voice clones of 4,000 reference voices, ~11,150 hours; apache-2.0). 236 of its voices are held out for testing and never trained on.
- Nothing else is used for training: a checkpoint is published here only when its training catalog (and any checkpoint
it started from) records no other source (the allowlist
configs/data/en_train_sources.txt, checked byscripts/hub_showcase.pybefore every upload).
Licence
apache-2.0 (owner decision 2026-09-29, hobby project) for the weights and the samples. The voice prompts are clips of held-out voices of SynDataLab-EN/echo-clones-4m-en (apache-2.0 on its card).
Limitations and responsible use
- The model learned only from synthetic voices: copying real people's voices is its weakest point (compare the seed-dev and echo-dev SIM-o above). English only; rare words and long numbers can be mispronounced.
- Copy a voice only with its owner's consent, never to impersonate anyone or to deceive, and say that the speech is synthetic.
- Every output of this code carries an inaudible AudioSeal watermark (16-bit
message
0100110101010100;python scripts/watermark_check.py detect <files>finds it). The mark is added by the inference code, not by the weights: do not remove it.
- Downloads last month
- -