Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up

litert-community
/
VibeVoice-ASR-BitNet

Automatic Speech Recognition
LiteRT-LM
LiteRT
VibeVoice
litertlm
on-device
edge
asr
speech-recognition
audio
bitnet
Model card Files Files and versions
xet
Community

Instructions to use litert-community/VibeVoice-ASR-BitNet with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

  • Libraries
  • LiteRT-LM

    How to use litert-community/VibeVoice-ASR-BitNet with LiteRT-LM:

    # LiteRT-LM runs on various platforms (Android, iOS, Windows, Linux, macOS, IoT, Web/WASM)
    # and supports many APIs (C++, Python, Kotlin, Swift, JavaScript, Flutter).
    # For platform-specific integration guides, please refer to the official developer website:
    # https://ai.google.dev/edge/litert-lm
    
    # To try LiteRT-LM, the easiest way is to use our CLI tool.
    # 1. Install the LiteRT-LM CLI tool:
    pip install -U litert-lm
    
    # 2. Download and run this model locally:
    # See: https://ai.google.dev/edge/litert-lm/cli
    litert-lm run \
      --from-huggingface-repo=litert-community/VibeVoice-ASR-BitNet \
      --prompt="Write me a poem"
  • LiteRT

    How to use litert-community/VibeVoice-ASR-BitNet with LiteRT:

    # No code snippets available yet for this library.
    
    # To use this model, check the repository files and the library's documentation.
    
    # Want to help? PRs adding snippets are welcome at:
    # https://github.com/huggingface/huggingface.js
  • VibeVoice

    How to use litert-community/VibeVoice-ASR-BitNet with VibeVoice:

    import torch, soundfile as sf, librosa, numpy as np
    from vibevoice.processor.vibevoice_processor import VibeVoiceProcessor
    from vibevoice.modular.modeling_vibevoice_inference import VibeVoiceForConditionalGenerationInference
    
    # Load voice sample (should be 24kHz mono)
    voice, sr = sf.read("path/to/voice_sample.wav")
    if voice.ndim > 1: voice = voice.mean(axis=1)
    if sr != 24000: voice = librosa.resample(voice, sr, 24000)
    
    processor = VibeVoiceProcessor.from_pretrained("litert-community/VibeVoice-ASR-BitNet")
    model = VibeVoiceForConditionalGenerationInference.from_pretrained(
        "litert-community/VibeVoice-ASR-BitNet", torch_dtype=torch.bfloat16
    ).to("cuda").eval()
    model.set_ddpm_inference_steps(5)
    
    inputs = processor(text=["Speaker 0: Hello!\nSpeaker 1: Hi there!"],
                       voice_samples=[[voice]], return_tensors="pt")
    audio = model.generate(**inputs, cfg_scale=1.3,
                           tokenizer=processor.tokenizer).speech_outputs[0]
    sf.write("output.wav", audio.cpu().numpy().squeeze(), 24000)
  • Notebooks
  • Google Colab
  • Kaggle
VibeVoice-ASR-BitNet
1.98 GB
Ctrl+K
Ctrl+K
  • 1 contributor
History: 6 commits
mlboydaisuke's picture
mlboydaisuke
Card: base_model_relation: quantized (list under the base model's Quantizations)
f119893 verified 1 day ago
  • .gitattributes
    1.59 kB
    VibeVoice-ASR-BitNet: int4-b128 ternary LM + int8 audio encoder (30 s window), generic audio path 1 day ago
  • LICENSE
    1.07 kB
    MIT LICENSE (inherited from microsoft/VibeVoice) 1 day ago
  • README.md
    8.22 kB
    Card: base_model_relation: quantized (list under the base model's Quantizations) 1 day ago
  • VibeVoice-ASR-BitNet.litertlm
    1.98 GB
    xet
    VibeVoice-ASR-BitNet: int4-b128 ternary LM + int8 audio encoder (30 s window), generic audio path 1 day ago
  • litertlm_manifest.json
    6.28 kB
    litertlm_manifest.json (public, derived from the bundle header + curated rows) 1 day ago