VoxCPM2: ONNX graphs for in-browser WebGPU inference
A converted and 4-bit quantized copy of openbmb/VoxCPM2
(2B text-to-speech model with voice design and voice cloning, 30 languages including Khmer),
built to run entirely in the browser with onnxruntime-web on WebGPU. No server needed.
Total download: about 1.9 GB. Start at manifest.json: it lists every file, its size, its SHA-256
and the config the decode loop needs.
Contents
| Path | What |
|---|---|
manifest.json |
file list, sizes, sha256, loop config |
q4f16/ |
feature encoder, base LM, residual LM step, LocDiT step (int4 weights, fp16 activations) |
vae/ |
AudioVAE encoder and decoder (fp32) |
embed_tokens.f16.bin |
token embedding table, raw fp16 (looked up in JS) |
tokenizer*.json, special_tokens_map.json |
upstream tokenizer, unchanged |
Graph input and output names and the decoding loop are specified by the project that produced these
files. The graphs use plain MatMul/Softmax attention (no GroupQueryAttention) and avoid ops that
the ORT WebGPU provider lacks. Noise is fed in as an input so results are seedable.
What was changed, and how well it matches the original
- Weights were converted from the upstream checkpoint into ONNX graphs and quantized to int4 with GPTQ calibration (plain round-to-nearest int4 was clearly worse and is not shipped).
- Against the fp32 PyTorch reference, teacher-forced patch cosine similarity was at least 0.974 (mean 0.985) across six test prompts, and the number of generated patches stayed within a few patches of the reference.
- Free-running output is chaotic, so the generated waveform is not identical to the original, even for fp16. Voice and prosody will differ from the PyTorch model for the same seed.
- Khmer intelligibility of this build has not been formally evaluated (no Khmer ASR was used). Please judge by listening.
- Tested on one NVIDIA GPU with Chrome (WebGPU), at roughly real-time speed. Other GPUs, browsers and phones are untested.
Responsible use
VoxCPM2 can clone voices. Upstream forbids use for impersonation, fraud or disinformation, and asks that AI-generated audio be clearly labeled. Only clone voices you own or have permission to use.
License and attribution
Apache License 2.0. See LICENSE and NOTICE. This is a modified derivative of VoxCPM2 by OpenBMB /
the VoxCPM Team. It is not endorsed by or affiliated with them. If you use it, please cite the
upstream work (see the VoxCPM2 model card).
Model tree for Uddom/voxcpm2-khmer-webgpu
Base model
openbmb/VoxCPM2