orukeet / transcribe-cpp /README.md
NathanRoll's picture
Expose transcribe.cpp Q8 at a Handy-compatible root filename
bc4e7f1 verified
|
Raw
History Blame Contribute Delete
3.28 kB

Orukeet for transcribe.cpp

This Q8_0 export uses transcribe.cpp's existing Parakeet TDT v3 implementation. The 12,288 fitted Gabor kernels are materialized as ordinary convolution weights. There are no additional runtime operators or dependencies.

  • File: orukeet-transcribe-cpp-Q8_0.gguf (739,508,608 bytes)
  • SHA-256: cad2f52ac91cad829279422301989687c2cf02e19157352ed25ea501b90dbb7e
  • Source: Orukeet r3, SHA-256 031c8ddab4845aeced904a7cde8e8aa57993b2e344716cf83a545b079c473b56
  • Weights: CC BY-SA 4.0. Orukeet is an adaptation of NVIDIA Parakeet TDT v3; retain attribution.

The root-level filename follows Handy's existing cache and delete layout. The earlier transcribe-cpp/orukeet-Q8_0.gguf path remains available and contains identical bytes.

Run

Build transcribe.cpp normally, then run:

./build/bin/transcribe-cli --model orukeet-transcribe-cpp-Q8_0.gguf recording.wav

Use 16 kHz mono audio. This is offline transcription in 25 languages, with language detection and token timestamps. Translation and streaming recognition are not supported. Use this layout with transcribe.cpp; the root-level orukeet-v0.1.0-q8.gguf is for NeMo-Speech.cpp.

Reproduce the export

At the converter commit in manifest.json, apply convert-orukeet.patch, then:

uv run --no-project --python 3.11 --with torch --with omegaconf --with sentencepiece --with gguf scripts/convert-parakeet.py orukeet-v0.1.0.nemo orukeet-F32.gguf --repo-id oruk/orukeet
./build/bin/transcribe-quantize orukeet-F32.gguf orukeet-Q8_0.gguf --quant Q8_0

Validation

The exact transcribe-cpp 0.2.0 Rust dependency pinned by Handy loads and transcribes this file on CPU and Apple Metal. Each device passed 120 multilingual clips, repeated decoding, 0.1/1/5-second silence, cancellation, and successful session reuse after cancellation. The supporting C++ build passed all 38 tests.

On the fixed six-language FLEURS validation sample (120 clips, 2,433 reference words), Q8 Metal and the NeMo FP32 source both score 5.34% pooled WER; Q8 CPU scores 5.30%. Metal matches the source text exactly on 108 clips and after word normalization on 112. The NeMo reference uses transcribe.cpp's pinned NeMo 2.8.0rc0 environment. All paired language scores and edit counts.

Q8 Metal WER Parakeet TDT v3 Orukeet
English 4.19% 3.40%
German 3.56% 4.07%
Spanish 3.41% 3.01%
French 4.07% 4.50%
Russian 7.77% 7.51%
Ukrainian 9.40% 11.60%
Pooled 5.14% 5.34%

Standard source/export tensor comparisons pass the existing upstream Parakeet tolerances and the JFK reference transcript matches exactly. Four additional sub-block probes exceed generic strict tolerances, with the largest differences at sequence boundaries. The runtime and NeMo use different valid-length padding paths; no runtime code or numerical tolerance was changed for this export.

These are integration checks for this export and runtime, separate from the report's NeMo benchmark. No Windows/Vulkan or comparative speed claim is made here.