makandal-multiple-v2

Try it in your browser: The Trio Serie Space runs Makandal with the other two models of the series: the Klara tab is a conversation with its eight tools, by voice or text, and another tab compares it with gemma-3-1b-it. More on the Makandal page of thetrio.space.

google/gemma-3-1b-it fine-tuned to be the small, local brain of Klara, a Kreyòl voice assistant that runs on a laptop. It answers in Haitian Creole, French, Spanish and English the way a voice assistant speaks (one to three short sentences, the answer first, no markdown, numbers as digits), holds a conversation, and calls tools: the time, the weather, an encyclopedia, a web search, a calculator, a memory of the person, and the user's own saved knowledge.

What changed from makandal-multiple (v1): v1 saw single questions only, could not call tools, and in a real session repeated its own last answer four turns in a row. v2 adds multi-turn conversations, tool traces and seven knowledge domains to v1's data.

Results

Same held-out tests for every model; nothing here was trained on.

gemma-3-1b-it v1 v2
Belebele Kreyòl / French / Spanish / English (restricted to 1–4) 29.5 / 39.5 / 40.0 / 47.5 34.5 / 46.5 / 42.5 / 49.0 39.0 / 51.0 / 47.5 / 50.5
Answer loss on held-out teacher answers, Kreyòl 5.44 1.10 0.86
Tool requests (300): right tool 10% – 93%
… called a tool when one was needed 4% – 98%
… no tool for small talk 93% – 100%
… arguments valid JSON 79% – 100%
Word problems with calculate (200): accuracy 6–48% – 28–30%
… used calculate 0–1% – 98–100%
Conversations (150, 722 replies): replies repeating an earlier one 6.6% – 0%
Replay of the session v1 looped in: refusals / repeats 0 / 1 looped 0 / 0

Math is GSM8K: English from the test split, Kreyòl, French and Spanish translated from training-split problems held out of training; the base model scores 47.5% on the Spanish ones, which it may have seen in English, and v2 30%. Two-tool requests (time and weather together, say) score 57% on the "right tool" line, which checks only the first call; the misses read were mostly correct calls in another order.

Known limits:

  • Tools after the first turn. Every tool trace in training was a single request, and the multi-turn conversations never called a tool, so later in a conversation it often answers from itself instead of calling one. Klara works around part of this; a v2.1 with multi-turn tool traces is planned.
  • remember is called for less than half of "my name is…" statements in a conversation, for the same reason.
  • Facts without tools are not reliable. A 1B model invents names and places. Do not use it alone for medical, legal or financial answers.

How it was made

Teacher: google/gemma-4-26B-A4B-it (DeepInfra, through the Hugging Face router), judged by google/gemma-4-31B-it; 2.7% of the examples (a top-up of tool traces and conversations) were written and judged by deepseek-v4-pro. Sequence-level distillation: the student was fine-tuned on the teacher's text.

Data (64,932 examples, every one checked by rules and most by the judge):

examples
v1's chat and translation (audited: refusals and false promises removed) 37,093
Domains: math with calculate (GSM8K train), science (ARC train), health and farming (grounded in Kreyòl documents), money and law, technology, homework 11,442
Multi-turn conversations in Klara's style 7,649
Tool traces with Klara's eight tools, run for real (Open-Meteo, Wikipedia, a calculator) 5,690
Replay: reading comprehension in all four languages 3,058

Kreyòl is half the data; French, Spanish and English about a sixth each. Tool results are in Kreyòl, as Klara's tools return them, and the answer follows the person's language.

Training: full fine-tune, loss on every assistant turn (words, tool calls and end of turn), one epoch, lr 1e-5 cosine, 4,069 steps, one A40, 107 minutes.

Tools and the chat template

The template (chat_template.jinja) puts the system prompt and the tool list at the top of the first user turn, writes a call as <tool_call>{"name": …, "arguments": {…}}</tool_call> in the model's turn, and a result as <tool_response>…</tool_response> in the next user turn. The tools it learned are in Klara's local_brain.py; it was trained with the system prompt Ou rele Klara. Ou se yon asistan vwa ki pale kreyòl ayisyen. and works best with it.

from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("jsbeaudry/makandal-multiple-v2")
model = AutoModelForCausalLM.from_pretrained("jsbeaudry/makandal-multiple-v2", dtype="auto")
weather = {"type": "function", "function": {"name": "weather", "description": "The weather now in a place.",
           "parameters": {"type": "object", "properties": {"place": {"type": "string"}}, "required": ["place"]}}}
msgs = [{"role": "system", "content": "Ou rele Klara. Ou se yon asistan vwa ki pale kreyòl ayisyen.\n"},
        {"role": "user", "content": "Ki tan l ap fè Okap jodi a?"}]
ids = tok.apply_chat_template(msgs, tools=[weather], add_generation_prompt=True, return_tensors="pt", return_dict=True)
out = model.generate(**ids, max_new_tokens=80)
print(tok.decode(out[0, ids["input_ids"].shape[1]:], skip_special_tokens=True))
# <tool_call>{"name": "weather", "arguments": {"place": "Okap"}}</tool_call>

For llama.cpp, use jsbeaudry/makandal-multiple-v2-GGUF.

Downloads last month
330
Safetensors
Model size
1.0B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jsbeaudry/makandal-multiple-v2

Finetuned
(559)
this model
Quantizations
2 models

Space using jsbeaudry/makandal-multiple-v2 1