← Back to work
01 · Engineering Live demo

GPT Story Generation

From-scratch Transformer Language Model

Overview

Built a GPT-style language model from scratch for story generation, implementing a custom training pipeline with GPT-2 BPE tokenization, SwiGLU and RMSNorm. Conducted hyperparameter experiments across learning rate, warmup and weight decay.

Results
31.7M
Parameters
15.2
Test PPL
1st
Competition place
Try the model

Give the model a story beginning and see what it generates. Runs on the deployed checkpoint, not in your browser.

0 / 500
Temperature0.80
Model
Architecture
Decoder-only Transformer
Parameters
31.7M
Layers
7
Attention heads
6
Embedding dim
384
Feed-forward
SwiGLU · hidden 1024 · bias-free
Normalization
RMSNorm
Tokenizer
GPT-2 BPE · vocab 50,304
Context
Causal self-attention · tied embeddings
Approach
ROCStories
GPT-2 BPE
Decoder-only Transformer
Fine-tuning
70% A + 30% B parameter average
Final checkpoint
Experiments
Learning rate
Warmup schedule
Weight decay
Stack
PyTorchTransformersLanguage Modeling
Notes

Markdown body is rendered under Notes on the detail page. Write freely here — what the problem was, what surprised you, what you would do differently.