Lumma-0.6B-Instruct

Introduction

Lumma-0.6B-Instruct is a compact, efficient multilingual language model designed for strong performance in resource-constrained environments. It is pre-trained from scratch on 1 trillion tokens and further enhanced through instruction tuning and Direct Preference Optimisation. This is a pre-RL checkpoint. The model supports English and 10 Indic languages.

Benchmark results

We benchmarked Lumma-0.6B-Instruct across multiple benchmarks, with an intentional focus on instruction-following capabilities. Despite its compact size, Lumma-0.6B-Instruct is able to match or outperform similar models up to 3Γ— larger on several instruction-following benchmarks.

While the model also delivers decent performance on mathematics and coding, we believe these capabilities are less critical for the primary real-world use cases targeted by such a small model, where developers typically prioritize efficient and reliable instruction following.

We expect further improvements with the RL-trained version of Lumma-0.6B-Instruct, particularly as we continue optimizing the model for real-world instruction-following tasks.

🌍 Supported Languages

The model is trained on English and a diverse set of Indic languages, including Hindi, Bengali, Tamil, Telugu, Marathi, Gujarati, Kannada, Malayalam, Punjabi, Odia

πŸš€ Usage

!pip install transformers=='5.4.0'

from IPython.display import display, Markdown
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model_name = "FrontiersMind/Lumma-0.6B-Instruct"

device = "cuda" if torch.cuda.is_available() else "cpu"

tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    trust_remote_code=True,
    dtype=torch.bfloat16
).to(device).eval()

prompt = "Explain newton's second law of motion"

messages = [
    {"role": "user", "content": prompt}
]

prompt = tokenizer.apply_chat_template(messages, tokenize=False)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

generated_ids = model.generate(
    **inputs,
    max_new_tokens=500,
    do_sample=True,
    temperature=0.3,
    top_p=0.90,
    top_k=20,
    repetition_penalty=1.1,
)

generated_ids = [
    output_ids[len(input_ids):] for input_ids, output_ids in zip(inputs.input_ids, generated_ids)
]

response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]

#print(response)
Markdown(response)

πŸ“¬ Feedback & Suggestions

We’d love to hear your thoughts, feedback, and ideas!

Downloads last month
432
Safetensors
Model size
0.6B params
Tensor type
F32
Β·
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for FrontiersMind/Lumma-0.6B-Instruct

Finetuned
(6)
this model

Collection including FrontiersMind/Lumma-0.6B-Instruct