Self_Correction_v1

Qwen2.5-7B-Instruct fine-tuned with verified math and code correction examples. The failed attempt and objective verifier feedback are context; training loss is computed only on the verified corrected response. This repository contains merged BF16 weights and can be loaded directly by vLLM.

Downloads last month
-
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Kxck/Self_Correction_v1

Base model

Qwen/Qwen2.5-7B
Finetuned
(3029)
this model