English

TinyStories GPT-124M

A 124M parameter GPT-2 architecture model implemented from scratch (following Sebastian Raschka's Build a Large Language Model from Scratch) and pretrained on the TinyStories dataset.

Details

  • Architecture: 12 layers, 12 heads, 768 embedding dim, 256 context length
  • Vocab: GPT-2 BPE (tiktoken), 50257 tokens
  • Training: ~2 epochs over ~360M tokens, AdamW (lr 3e-4, weight decay 0.1), bf16 mixed precision, on an NVIDIA DGX Spark

Loading

This is not using the prebuilt transformers library— it uses a custom GPTModel class, included as CodingGPTFr.py.

```python import torch from CodingGPTFr import GPTModel

config = {"vocab_size": 50257, "context_length": 256, "emb_dim": 768, "n_heads": 12, "n_layers": 12, "drop_rate": 0.1, "qkv_bias": False} model = GPTModel(config) model.load_state_dict(torch.load("model_final.pth", map_location="cpu")) model.eval() ```

See generate.py for sampling with temperature and top-k.

Limitations

Trained only on synthetic children's stories. It writes fluent, simple narratives but has no factual knowledge, cannot do arithmetic, and is a base model (not instruction-tuned) — it completes text rather than answering questions. Logical consistency across a story degrades at this scale.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train noOneCodes/tinystories-gpt-124m