01 · Engineering Live demo
GPT Story Generation
From-scratch Transformer Language Model
Overview
Built a GPT-style language model from scratch for story generation, implementing a custom training pipeline with GPT-2 BPE tokenization, SwiGLU and RMSNorm. Conducted hyperparameter experiments across learning rate, warmup and weight decay.
Results
31.7M
Parameters
15.2
Test PPL
1st
Competition place
Try the model
Give the model a story beginning and see what it generates. Runs on the deployed checkpoint, not in your browser.
0 / 500
Temperature0.80
Model
Architecture
Decoder-only Transformer
Parameters
31.7M
Layers
7
Attention heads
6
Embedding dim
384
Feed-forward
SwiGLU · hidden 1024 · bias-free
Normalization
RMSNorm
Tokenizer
GPT-2 BPE · vocab 50,304
Context
Causal self-attention · tied embeddings
Approach
ROCStories
↓
GPT-2 BPE
↓
Decoder-only Transformer
↓
Fine-tuning
↓
70% A + 30% B parameter average
↓
Final checkpoint
Experiments
Learning rate
Warmup schedule
Weight decay
Stack
PyTorchTransformersLanguage Modeling
Notes
Markdown body is rendered under Notes on the detail page. Write freely here — what the problem was, what surprised you, what you would do differently.