Instructions to use upstage/Solar-Open2-250B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use upstage/Solar-Open2-250B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="upstage/Solar-Open2-250B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("upstage/Solar-Open2-250B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use upstage/Solar-Open2-250B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "upstage/Solar-Open2-250B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "upstage/Solar-Open2-250B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/upstage/Solar-Open2-250B
- SGLang
How to use upstage/Solar-Open2-250B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "upstage/Solar-Open2-250B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "upstage/Solar-Open2-250B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "upstage/Solar-Open2-250B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "upstage/Solar-Open2-250B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use upstage/Solar-Open2-250B with Docker Model Runner:
docker model run hf.co/upstage/Solar-Open2-250B
Question on the architecture ablation and 1M-context evaluation
Thank you for releasing Solar Open 2 and its technical report. I found the hybrid KDA–softmax architecture, NoPE design, and support for negative eigenvalues particularly interesting.
I have a question regarding the proxy architecture ablation.
The report uses validation loss, MMLU, and HellaSwag learning curves to show that the hybrid architecture reaches a given level of general capability with fewer training tokens. These results provide useful evidence of improved pre-training sample efficiency.
However, the report also states that the central architectural objective is to provide a usable context window beyond one million tokens and support long-horizon agentic tasks. The motivations for NoPE, hybrid attention, and negative eigenvalues are closely related to long-context extrapolation, information retention, and state tracking.
I recognize that the final model is evaluated on AA-LCR and that multiple metrics are tracked during the length-expansion stage. However, those results reflect the completed model after pre-training, length expansion, checkpoint merging, and post-training, rather than isolating the architectural contribution itself.
Were controlled proxy ablations conducted between the all-softmax baseline and the hybrid architecture on length-dependent long-context tasks, such as retrieval, aggregation, information overwrite, or state tracking? If such experiments were omitted only at the report level but actually conducted, would it be possible to share those results?
My concern is not that MMLU and HellaSwag are unsuitable as low-cost proxy signals. Rather, they demonstrate general sample efficiency but do not directly establish whether the hybrid architecture improves the long-context capabilities it was primarily designed to support. If such controlled long-context experiments were conducted, it would be valuable to see the results.