Octo v1.0
Octo selects among supplied text candidates and returns probability distributions without generating answer tokens. Candidates are encoded independently using a custom attention mask, and a shared pointer head scores them against a decision representation.
Model
| Property | Value |
|---|---|
| Version | 1.0 (v1.0) |
| Maintainer | b1zantine |
| Base model | Qwen/Qwen3-1.7B-Base |
| Required base revision | ea980cb0a6c2ae4b936e82123acc929f1cec04c1 |
| Adapter | LoRA, rank 16, alpha 32 |
| Pointer dimension | 256 |
| Precision | float32 |
| Attention implementation | SDPA |
| Language | English; chess position and move descriptions |
This repository includes the adapters, pointer weights, tokenizer, and inference runtime. The base weights are downloaded separately. Octo requires its own loader because standard Transformers pipelines do not implement its attention layout and decision head.
Capabilities
- Choice: selects a supplied option and returns probabilities by option ID.
- Score: returns an expected ordinal level and a distribution over supplied criteria.
- Noul: returns the probability that a statement is true of the supplied state.
Measured quality is available for Choice only. Score and Noul are available in the runtime but their quality is unvalidated for v1.0.
Intended applications include support-intent routing among supplied categories and experimental tactical chess move selection among supplied legal moves. Candidate IDs stay outside model text. The model cannot select an answer that is absent from the supplied options.
Usage
Use Python 3.12 or newer in a virtual environment:
python -m pip install "huggingface-hub>=1,<2"
python - <<'PYTHON'
from huggingface_hub import snapshot_download
snapshot_download("b1zantine/octo", revision="v1.0", local_dir="octo-v1.0")
PYTHON
python -m pip install ./octo-v1.0
import torch
from octo import load_checkpoint
device = "cuda" if torch.cuda.is_available() else "cpu"
model, encoder, _ = load_checkpoint(
"octo-v1.0", device=device, local_files_only=False
)
request = {
"record_id": "example",
"source_group_id": "example",
"state": "I was charged twice for the same card purchase.",
"questions": [{
"id": "category",
"type": "choice",
"instruction": "Which support category matches this request?",
"options": [
{"id": "duplicate-charge", "description": "A card purchase was charged more than once"},
{"id": "card-delivery", "description": "A new bank card has not arrived"},
],
}],
}
with torch.inference_mode():
answers = model.predict(encoder.batch([request]).to(device))
print(answers)
The first load downloads the pinned base model if needed. Once cached, local_files_only=True enables offline loading. CPU, CUDA, and Apple MPS are supported; float32 base weights need several GB of memory plus inference overhead.
Choice responses contain selected_id, probabilities, and confidence. Score responses contain the expected level and distribution. Noul responses contain probability_true and a true/false distribution. Confidence describes distribution concentration; it is not a calibrated probability of correctness.
Input limits
| Limit | Value |
|---|---|
| Questions per record | 1 |
| Choice candidates | 2โ8 |
| Physical tokens | 512 |
| Logical positions | 512 |
| Dense attention-mask budget | 64 MiB |
Over-budget requests are rejected rather than truncated. Behavior with larger limits or packed questions is unvalidated.
Measured performance
| Choice benchmark | Examples | Accuracy | 95% interval | Uniform-choice baseline |
|---|---|---|---|---|
| BFSI core | 800 | 99.375% | 98.750โ99.875% | 27.589% |
| Chess core | 200 | 71.500% | 65.488โ77.500% | 27.014% |
BFSI covers banking, insurance, and wealth-management support categories. Both benchmarks use supplied candidate subsets containing the correct answer, with errors counted as incorrect. Intervals use source-family bootstrap sampling. These figures do not represent full-taxonomy intent classification, general chess strength, or results on the entire locked test corpus. Machine-readable summaries are in evaluation/.
Limitations
- Quality outside English support-intent and tactical chess tasks is unmeasured.
- Chess results measure selecting a tactical solution among supplied legal options, not Elo or full-game strength.
- Ambiguous evidence, misleading candidate descriptions, omitted answers, and distribution shifts can cause incorrect decisions.
- Probabilities are uncalibrated. Validate thresholds and error rates on the intended application.
- Support classification does not establish the accuracy of financial advice or eligibility decisions.
- Score, Noul, longer contexts, larger candidate sets, and multi-question requests have no measured task-quality guarantees.
Files
backbone/: LoRA adapter weights and configuration.pointer.pt: decision-head tensors, loaded withweights_only=True.tokenizer/: tokenizer files.octo.json: model configuration and input contract.octo/andpyproject.toml: installable inference runtime.evaluation/: model performance summaries.file_checksums.json: release-file SHA-256 checksums.
Use the v1.0 revision to pin this model version.
License
Octo and its bundled runtime are distributed under the MIT license, copyright b1zantine. The separately downloaded Qwen base weights use Apache 2.0.
Model tree for b1zantine/octo
Base model
Qwen/Qwen3-1.7B-Base