Whisper Tiny es

This model is a fine-tuned version of openai/whisper-tiny on the following datasets:

  • deepdml/common_voice_26_0
  • disco-eth/WorldSpeech
  • facebook/voxpopuli
  • facebook/multilingual_librispeech
  • deepdml/voxforge
  • google/fleurs
  • deepdml/basque_parliament_1

It achieves the following results on the evaluation set:

  • Best checkpoint: 44,000
  • Loss: 0.2888
  • Wer: 16.0693
  • Cer: 6.0413

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 1e-05
  • train_batch_size: 128
  • eval_batch_size: 128
  • seed: 42
  • optimizer: Use adamw_torch with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: linear
  • lr_scheduler_warmup_steps: 0.04
  • training_steps: 46000

Training results

Training Loss Epoch Step Validation Loss Wer Raw Cer Raw Wer Cer
0.3663 0.0217 1000 0.5971 31.8722 12.1478 31.6510 12.1036
0.2625 0.0435 2000 0.5140 28.2789 10.5567 28.2345 10.5484
0.2183 0.0652 3000 0.4763 25.3893 9.2745 25.3683 9.2709
0.2320 0.0870 4000 0.4435 24.3033 9.0073 24.2944 9.0057
0.3290 0.1087 5000 0.4363 23.9673 9.1117 23.9590 9.1103
0.2030 0.1304 6000 0.4118 22.9428 8.5741 22.9390 8.5735
0.3456 0.1522 7000 0.4131 22.7190 8.6583 22.7139 8.6574
0.3284 0.1739 8000 0.3996 22.6752 8.4567 22.6727 8.4563
0.2336 0.1957 9000 0.3692 20.9236 7.9288 20.9223 7.9286
0.2950 0.2174 10000 0.3622 21.1436 8.2788 21.1436 8.2788
0.3875 0.2391 11000 0.3575 20.1197 7.5460 20.1197 7.5460
0.1884 1.0020 12000 0.3406 19.3906 7.2308 19.3906 7.2308
0.1370 1.0237 13000 0.3373 19.3215 7.3625 19.3215 7.3625
0.1426 1.0455 14000 0.3364 19.1237 7.1903 19.1237 7.1903
0.1616 1.0672 15000 0.3329 18.5785 6.9592 18.5785 6.9592
0.1560 1.0889 16000 0.3278 19.0952 7.3977 19.0952 7.3977
0.1511 1.1107 17000 0.3308 18.7877 7.2803 18.7877 7.2803
0.1642 1.1324 18000 0.3276 18.3573 6.7823 18.3573 6.7823
0.2495 1.1542 19000 0.3300 18.6108 7.0579 18.6108 7.0579
0.1903 1.1759 20000 0.3232 18.4504 7.0561 18.4504 7.0561
0.1982 1.1976 21000 0.3161 18.3636 7.0710 18.3636 7.0710
0.3891 1.2194 22000 0.3154 18.5240 7.1787 18.5240 7.1787
0.2725 1.2411 23000 0.3155 17.8260 6.6920 17.8260 6.6920
0.1271 2.0040 24000 0.3016 17.1540 6.5646 17.1540 6.5646
0.1171 2.0257 25000 0.3090 17.2700 6.5550 17.2700 6.5550
0.1738 2.0474 26000 0.3037 17.3721 6.6015 17.3721 6.6015
0.1532 2.0692 27000 0.3069 17.1977 6.4330 17.1977 6.4330
0.2220 2.0909 28000 0.3056 17.4139 6.5979 17.4139 6.5979
0.1509 2.1127 29000 0.3010 16.9923 6.5297 16.9923 6.5297
0.1484 2.1344 30000 0.3033 16.7863 6.2084 16.7863 6.2084
0.1558 2.1561 31000 0.3021 17.1407 6.5016 17.1407 6.5016
0.1563 2.1779 32000 0.2979 16.6677 6.2236 16.6677 6.2236
0.2823 2.1996 33000 0.3006 17.1850 6.4912 17.1850 6.4912
0.1463 2.2213 34000 0.2955 16.6018 6.3068 16.6018 6.3068
0.1209 2.2431 35000 0.2916 16.6373 6.3950 16.6373 6.3950
0.1145 3.0059 36000 0.2941 16.8801 6.5048 16.8801 6.5048
0.1545 3.0277 37000 0.2925 16.6208 6.4095 16.6208 6.4095
0.3211 3.0494 38000 0.2905 16.6354 6.4966 16.6354 6.4966
0.1914 3.0712 39000 0.2926 16.2645 6.1487 16.2645 6.1487
0.1672 3.0929 40000 0.2936 16.5517 6.2828 16.5517 6.2828
0.1256 3.1146 41000 0.2910 16.6861 6.5220 16.6861 6.5220
0.1023 3.1364 42000 0.2918 16.4179 6.2376 16.4179 6.2376
0.1597 3.1581 43000 0.2913 16.6151 6.4257 16.6151 6.4257
0.1912 3.1798 44000 0.2888 16.0693 6.0413 16.0693 6.0413
0.1459 3.2016 45000 0.2891 16.2379 6.1107 16.2379 6.1107
0.2485 3.2233 46000 0.2889 16.3450 6.1852 16.3450 6.1852

Framework versions

  • Transformers 5.14.1
  • Pytorch 2.6.0+cu124
  • Datasets 5.0.1
  • Tokenizers 0.22.2

Citation

Please cite the model using the following BibTeX entry:

@misc{deepdml/whisper-tiny-es-mix-norm,
      title={Fine-tuned Whisper tiny ASR model for speech recognition in Spanish},
      author={Jimenez, David},
      howpublished={\url{https://huggingface.co/deepdml/whisper-tiny-es-mix-norm}},
      year={2026}
    }
Downloads last month
8,478
Safetensors
Model size
37.8M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for deepdml/whisper-tiny-es-mix-norm

Finetuned
(1898)
this model
Finetunes
1 model

Datasets used to train deepdml/whisper-tiny-es-mix-norm

Evaluation results