VoxCPM2: ONNX graphs for in-browser WebGPU inference

A converted and 4-bit quantized copy of openbmb/VoxCPM2 (2B text-to-speech model with voice design and voice cloning, 30 languages including Khmer), built to run entirely in the browser with onnxruntime-web on WebGPU. No server needed.

Total download: about 1.9 GB. Start at manifest.json: it lists every file, its size, its SHA-256 and the config the decode loop needs.

Contents

Path What
manifest.json file list, sizes, sha256, loop config
q4f16/ feature encoder, base LM, residual LM step, LocDiT step (int4 weights, fp16 activations)
vae/ AudioVAE encoder and decoder (fp32)
embed_tokens.f16.bin token embedding table, raw fp16 (looked up in JS)
tokenizer*.json, special_tokens_map.json upstream tokenizer, unchanged

Graph input and output names and the decoding loop are specified by the project that produced these files. The graphs use plain MatMul/Softmax attention (no GroupQueryAttention) and avoid ops that the ORT WebGPU provider lacks. Noise is fed in as an input so results are seedable.

What was changed, and how well it matches the original

  • Weights were converted from the upstream checkpoint into ONNX graphs and quantized to int4 with GPTQ calibration (plain round-to-nearest int4 was clearly worse and is not shipped).
  • Against the fp32 PyTorch reference, teacher-forced patch cosine similarity was at least 0.974 (mean 0.985) across six test prompts, and the number of generated patches stayed within a few patches of the reference.
  • Free-running output is chaotic, so the generated waveform is not identical to the original, even for fp16. Voice and prosody will differ from the PyTorch model for the same seed.
  • Khmer intelligibility of this build has not been formally evaluated (no Khmer ASR was used). Please judge by listening.
  • Tested on one NVIDIA GPU with Chrome (WebGPU), at roughly real-time speed. Other GPUs, browsers and phones are untested.

Responsible use

VoxCPM2 can clone voices. Upstream forbids use for impersonation, fraud or disinformation, and asks that AI-generated audio be clearly labeled. Only clone voices you own or have permission to use.

License and attribution

Apache License 2.0. See LICENSE and NOTICE. This is a modified derivative of VoxCPM2 by OpenBMB / the VoxCPM Team. It is not endorsed by or affiliated with them. If you use it, please cite the upstream work (see the VoxCPM2 model card).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Uddom/voxcpm2-khmer-webgpu

Base model

openbmb/VoxCPM2
Quantized
(16)
this model