Octo v1.0

Octo selects among supplied text candidates and returns probability distributions without generating answer tokens. Candidates are encoded independently using a custom attention mask, and a shared pointer head scores them against a decision representation.

Model

Property Value
Version 1.0 (v1.0)
Maintainer b1zantine
Base model Qwen/Qwen3-1.7B-Base
Required base revision ea980cb0a6c2ae4b936e82123acc929f1cec04c1
Adapter LoRA, rank 16, alpha 32
Pointer dimension 256
Precision float32
Attention implementation SDPA
Language English; chess position and move descriptions

This repository includes the adapters, pointer weights, tokenizer, and inference runtime. The base weights are downloaded separately. Octo requires its own loader because standard Transformers pipelines do not implement its attention layout and decision head.

Capabilities

  • Choice: selects a supplied option and returns probabilities by option ID.
  • Score: returns an expected ordinal level and a distribution over supplied criteria.
  • Noul: returns the probability that a statement is true of the supplied state.

Measured quality is available for Choice only. Score and Noul are available in the runtime but their quality is unvalidated for v1.0.

Intended applications include support-intent routing among supplied categories and experimental tactical chess move selection among supplied legal moves. Candidate IDs stay outside model text. The model cannot select an answer that is absent from the supplied options.

Usage

Use Python 3.12 or newer in a virtual environment:

python -m pip install "huggingface-hub>=1,<2"
python - <<'PYTHON'
from huggingface_hub import snapshot_download
snapshot_download("b1zantine/octo", revision="v1.0", local_dir="octo-v1.0")
PYTHON
python -m pip install ./octo-v1.0
import torch
from octo import load_checkpoint

device = "cuda" if torch.cuda.is_available() else "cpu"
model, encoder, _ = load_checkpoint(
    "octo-v1.0", device=device, local_files_only=False
)
request = {
    "record_id": "example",
    "source_group_id": "example",
    "state": "I was charged twice for the same card purchase.",
    "questions": [{
        "id": "category",
        "type": "choice",
        "instruction": "Which support category matches this request?",
        "options": [
            {"id": "duplicate-charge", "description": "A card purchase was charged more than once"},
            {"id": "card-delivery", "description": "A new bank card has not arrived"},
        ],
    }],
}
with torch.inference_mode():
    answers = model.predict(encoder.batch([request]).to(device))
print(answers)

The first load downloads the pinned base model if needed. Once cached, local_files_only=True enables offline loading. CPU, CUDA, and Apple MPS are supported; float32 base weights need several GB of memory plus inference overhead.

Choice responses contain selected_id, probabilities, and confidence. Score responses contain the expected level and distribution. Noul responses contain probability_true and a true/false distribution. Confidence describes distribution concentration; it is not a calibrated probability of correctness.

Input limits

Limit Value
Questions per record 1
Choice candidates 2โ€“8
Physical tokens 512
Logical positions 512
Dense attention-mask budget 64 MiB

Over-budget requests are rejected rather than truncated. Behavior with larger limits or packed questions is unvalidated.

Measured performance

Choice benchmark Examples Accuracy 95% interval Uniform-choice baseline
BFSI core 800 99.375% 98.750โ€“99.875% 27.589%
Chess core 200 71.500% 65.488โ€“77.500% 27.014%

BFSI covers banking, insurance, and wealth-management support categories. Both benchmarks use supplied candidate subsets containing the correct answer, with errors counted as incorrect. Intervals use source-family bootstrap sampling. These figures do not represent full-taxonomy intent classification, general chess strength, or results on the entire locked test corpus. Machine-readable summaries are in evaluation/.

Limitations

  • Quality outside English support-intent and tactical chess tasks is unmeasured.
  • Chess results measure selecting a tactical solution among supplied legal options, not Elo or full-game strength.
  • Ambiguous evidence, misleading candidate descriptions, omitted answers, and distribution shifts can cause incorrect decisions.
  • Probabilities are uncalibrated. Validate thresholds and error rates on the intended application.
  • Support classification does not establish the accuracy of financial advice or eligibility decisions.
  • Score, Noul, longer contexts, larger candidate sets, and multi-question requests have no measured task-quality guarantees.

Files

  • backbone/: LoRA adapter weights and configuration.
  • pointer.pt: decision-head tensors, loaded with weights_only=True.
  • tokenizer/: tokenizer files.
  • octo.json: model configuration and input contract.
  • octo/ and pyproject.toml: installable inference runtime.
  • evaluation/: model performance summaries.
  • file_checksums.json: release-file SHA-256 checksums.

Use the v1.0 revision to pin this model version.

License

Octo and its bundled runtime are distributed under the MIT license, copyright b1zantine. The separately downloaded Qwen base weights use Apache 2.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for b1zantine/octo

Adapter
(81)
this model