Instructions to use cstr/index-echo-2b-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use cstr/index-echo-2b-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf cstr/index-echo-2b-GGUF:F16 # Run inference directly in the terminal: llama cli -hf cstr/index-echo-2b-GGUF:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf cstr/index-echo-2b-GGUF:F16 # Run inference directly in the terminal: llama cli -hf cstr/index-echo-2b-GGUF:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf cstr/index-echo-2b-GGUF:F16 # Run inference directly in the terminal: ./llama-cli -hf cstr/index-echo-2b-GGUF:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf cstr/index-echo-2b-GGUF:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf cstr/index-echo-2b-GGUF:F16
Use Docker
docker model run hf.co/cstr/index-echo-2b-GGUF:F16
- LM Studio
- Jan
- Ollama
How to use cstr/index-echo-2b-GGUF with Ollama:
ollama run hf.co/cstr/index-echo-2b-GGUF:F16
- Unsloth Desktop
- Pi
How to use cstr/index-echo-2b-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf cstr/index-echo-2b-GGUF:F16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "cstr/index-echo-2b-GGUF:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use cstr/index-echo-2b-GGUF with Docker Model Runner:
docker model run hf.co/cstr/index-echo-2b-GGUF:F16
- Lemonade
How to use cstr/index-echo-2b-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull cstr/index-echo-2b-GGUF:F16
Run and chat with the model
lemonade run user.index-echo-2b-GGUF-F16
List all available models
lemonade list
- Hermes Agent
How to use cstr/index-echo-2b-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf cstr/index-echo-2b-GGUF:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default cstr/index-echo-2b-GGUF:F16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use cstr/index-echo-2b-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf cstr/index-echo-2b-GGUF:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "cstr/index-echo-2b-GGUF:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Index-Echo S2TT 2B — GGUF for CrispASR
Conversion of IndexTeam/Index-Echo-S2TT-2B
for the index-echo backend in CrispASR.
Weights have passed the independent CPU/CUDA checks below. The backend is
available on CrispASR main, including C ABI loading, companion downloads and
a pinned nightly regression.
The released model transcribes Chinese and translates each subtitle into English, Japanese or Spanish. CrispASR preserves the three-line generation recipe and returns bilingual cues with timestamps and punctuation. English JFK transcription/English translation also matches the released Python model in the CPU fixtures; this is a port-parity result, not a language-wide accuracy claim. The 9B checkpoint is not included.
Files and companions
| Precision | Tower + connector | Qwen3.5 decoder | Combined |
|---|---|---|---|
| F16 | index-echo-2b-f16.gguf (1,313,848,896 bytes) |
index-echo-2b-decoder-f16.gguf (3,897,387,648 bytes) |
4.853 GiB |
| Q8_0 | index-echo-2b-q8_0.gguf (702,951,040 bytes) |
index-echo-2b-decoder-q8_0.gguf (2,076,674,688 bytes) |
2.589 GiB |
Both files of the chosen precision must be in the same directory. The tower's
index_echo.decoder_file metadata names its decoder; point CrispASR at the
tower, not the decoder. The decoder omits the absent MTP branch and contains
exactly the 24 text blocks used by the released speech inference path. Vision
weights are omitted.
For speech-boundary windows, also put ggml-silero-v6.2.0.bin from
ggml-org/whisper-vad beside them,
or pass --vad-model. Silero uses its separate MIT license. Without Silero,
CrispASR uses bounded 60-second windows. The source recipe keeps context from
up to five prior windows and resets history for each input file.
Use
Build CrispASR with the Index-Echo backend, then download a matching pair:
hf download cstr/index-echo-2b-GGUF \
index-echo-2b-q8_0.gguf index-echo-2b-decoder-q8_0.gguf \
--local-dir models/index-echo
hf download ggml-org/whisper-vad ggml-silero-v6.2.0.bin \
--local-dir models/index-echo
./build/bin/crispasr --backend index-echo \
-m models/index-echo/index-echo-2b-q8_0.gguf \
-f audio.wav --target-lang en -osrt
--target-lang en|ja|es selects the translation. --prompt accepts glossary
entries such as name:translation; --ask overrides the task instruction.
Greedy decoding is the default. Positive temperature enables sampling;
--seed and --max-new-tokens are forwarded. Custom instructions need the
three-line timestamp/transcript/translation format for parsed bilingual cues;
unformatted custom output is returned as a whole-window text segment.
Word timestamps, beam search and live streaming are not exposed by this recipe.
Validation and quantization
CPU validation run 36905507951 passes F16 and Q8_0 against independent CPU captures (F32 tower/connector, BF16 decoder) from the released Python inference class: English JFK, Chinese and a partial-hop JFK clip. Checks cover mel/convolutions, all 32 audio blocks, the connector, prompt IDs, all 24 decoder blocks, logits and 16 cached teacher-forced steps. The actual Python Session/C ABI loads a filename without a backend hint and reproduces all decoded text and centisecond timestamps exactly.
| Precision | Worst stage-row cosine | Largest stage relative L2 |
|---|---|---|
| F16 | 0.998464 | 1.415% |
| Q8_0 | 0.995173 | 3.561% |
A separately forced F32 decoder diagnostic now proves higher-precision math: run 36915483750 passes every F16 stage, magnitude, cache, direct-output and full-file gate, including English/Japanese/Spanish and two-window context. Minimum cosine is 0.999992; maximum stage relative L2 is 0.239%. Q8 passes stage/cache/direct checks against that reference too (0.997150 minimum, 3.094% maximum relative L2). The overall diagnostic run is red because Q8 has the documented full-file precision differences; this is not an all-cohort exact-parity claim.
Cosine alone is scale-blind, so magnitude is gated separately. F16 stages require cosine >= .999 (cached logits >= .998) and relative L2 <= 2%; quantized learned stages require cosine >= .99 and relative L2 <= 5%. Frontend stages retain strict gates. Prompt IDs, cached greedy IDs and full direct-window decoded cues are additional mandatory checks. Full-file quality acceptance separately requires the complete text of an independent source precision variant, with exact F16 timestamps and at most 20ms Q8 timestamp drift. No arbitrary text edits are accepted.
Plain Q4_K is rejected: worst cosine 0.781530, relative L2 up to 23.349%, and changed translations/punctuation. Selective Q4 is also rejected: it fails stage/output checks and saves only 2.8% versus Q8. Neither Q4 recipe is a recommended download.
Full-file checks exposed and fixed a Silero waveform-context discrepancy.
With that fix, F16 full-file English/Japanese/Spanish and two-window outputs
match the separate explicitly F32 decoder source run. The released BF16
decoder differs by 20ms on three Japanese boundaries and chooses a Spanish
synonym (creamos versus F32/native F16 pensemos). Q8 also shifts one English
boundary by 20ms. Exact mixed-precision parity is therefore claimed only for
the three direct-window fixtures, not all full-file multilingual outputs.
A same-host four-core ARM Neoverse-N2 profile shows Q8 warm speedups of 1.33–1.46× over the actual F32/BF16 Python blueprint, with peak RSS 4.03 versus 7.95GiB. A two-Tesla-T4 CUDA run passes stage/cache checks for both cohorts; GPU SRTs match native CPU controls byte for byte. Q8 inference is 4.5–5.3× realtime on those three short clips, excluding load. See CrispASR PERFORMANCE.md and its JSON receipts for methodology and limitations.
Reproducibility and license
Source revision: 5d98a34d9685869b11e9c94d01dee22a8b8e53b5.
The tower converter is models/convert-index-echo-to-gguf.py; the standard
Qwen3.5 decoder converter is llama.cpp commit
42d958167a748f2c04b1f888e84e7a58f609ddcb, with --outtype f16 --no-mtp.
Quantization uses crispasr-quantize; frontend constants and small affines
retain floating-point precision. Stage fixtures are in
cstr/crispasr-regression-fixtures
at 374efe4d7f5c4ff5dce32deffb14e23865fa8493: released mixed-precision
captures under index-echo-2b/, separately forced F32 decoder captures under
index-echo-2b-f32/. The nightly uses the F32 fixture; the original golden
is preserved. Both source full-pipeline captures and effective-dtype audit are
included.
The model weights are Apache-2.0; the source license is included as LICENSE.
CrispASR runtime code has its own repository license. This conversion is not
an official IndexTeam release.
- Downloads last month
- 134
8-bit
16-bit
Model tree for cstr/index-echo-2b-GGUF
Base model
IndexTeam/Index-Echo-S2TT-2B