Experiment-1-A

Modded-GPT training experiment on pre-tokenized GPT-2 FineWeb-Edu shards. The repository contains custom PyTorch source and raw safetensors weights; it is not a drop-in Transformers model.

Status

  • Status: complete
  • Processed training tokens: 2,000,158,720
  • Next zero-based training step: 3815
  • Run ID: 16a40141-39c7-4924-a14d-da799d58e24f
  • GPU: NVIDIA A100-SXM4-40GB
  • Training dtype: BF16
  • Target tokens: 2,000,000,000
  • Tokens per optimizer step: 524,288

Repository layout

  • model/model.safetensors: final standalone model weights
  • checkpoints/250M, 500M, 750M, final: retained full checkpoints
  • checkpoints/latest/manifest.json: pointer to the newest valid resumable checkpoint
  • training/: exact source and resolved training configuration
  • runtime/: environment metadata
  • metrics/training_metrics.jsonl: per-step training metrics

FineWeb-Edu dataset shards are not included.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train efe-T/Experiment-1-A