Instructions to use Valen-Team/Valen-2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Valen-Team/Valen-2B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Valen-Team/Valen-2B", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Valen-Team/Valen-2B", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Valen-Team/Valen-2B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Valen-Team/Valen-2B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Valen-Team/Valen-2B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Valen-Team/Valen-2B
- SGLang
How to use Valen-Team/Valen-2B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Valen-Team/Valen-2B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Valen-Team/Valen-2B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Valen-Team/Valen-2B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Valen-Team/Valen-2B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Valen-Team/Valen-2B with Docker Model Runner:
docker model run hf.co/Valen-Team/Valen-2B
Valen-2B
Qwen3.5 with a two-layer MLP-Mixer decision head. Supports text, images and videos, and returns Choice, Noul and Score decisions. Multiple questions can share one state encoding with execution="shared_state". Trained on millions of samples.
Inference
Install Python 3.10+, PyTorch, Transformers 5.4.0, torchvision, Pillow and av. Flash Attention 2 is optional with a compatible CUDA build. The model contains its tokenizer, processor and custom inference code; a separate base-model download is unnecessary.
import torch
from transformers import AutoModel
model = AutoModel.from_pretrained(
"Valen-Team/Valen-2B", trust_remote_code=True,
dtype="auto", attn_implementation="sdpa",
).to("cuda").eval()
torch.set_float32_matmul_precision("highest")
torch.backends.cudnn.allow_tf32 = False
print(model.predict({
"state": "A cat is on the sofa.",
"questions": {
"animal": {"type": "choice", "instructions": "Which animal is present?",
"criteria": {"cat": "A cat", "dog": "A dog"}},
"on_sofa": {"type": "noul", "instructions": "The cat is on the sofa."},
},
}, execution="shared_state"))
Use attn_implementation="flash_attention_2" for Flash Attention, or execution="question" for independent questions. Video defaults to 16 frames. Image/video request examples and training instructions are in the Valen repository.
dtype="auto" preserves the trained FP32 parameters. The backbone runs under BF16 autocast and the Mixer stays FP32, matching native checkpoint evaluation. Loading all weights as BF16 rounds the trained parameters and can change outputs.
Evaluation
| Benchmark | Accuracy (%) |
|---|---|
| VisualDecisionBench Image | 80.41 |
| VisualDecisionBench Video | 82.85 |
| JevBench | 83.55 |
- Downloads last month
- 65