MiniGPT Studio
Training a language model from scratch — on a laptop.
Ask MiniGPT something.
This is a small model trained on this Mac. It answers offline and can be wrong.
Ready.
An end-to-end pipeline that trains a decoder-only Transformer from random weights on Apple Silicon, evaluates it honestly, and lets you talk to it in a native-feeling macOS app. No pretrained weights. No hosted model APIs.
- parameters
- 29.4M
29,447,168 — 8 layers, 8 heads, width 512
- token corpus
- 445M
305,062 synthetic conversations, leakage-checked
- pretrained weights
- 0
every parameter starts random
- test perplexity
- 18.54
13.8M story model · vs 469.42 unigram, 8,603.77 random
The idea
Most student “AI projects” call someone else’s model. I wanted to own every step — from the first downloaded byte to the last generated token.
01Overview
Version 1.0.0 beta · GPL-3.0 · 38 commits
A whole language-model lab, running offline on one Mac.
MiniGPT Studio downloads pinned public datasets and checks every file against a recorded SHA-256. It groups related and near-duplicate texts so no family of documents can land in both training and test, fits its own byte-level BPE tokenizer on training text only, and trains a GPT-style model on the Apple GPU with crash-safe checkpoints and exact resume.
The model is then scored on the entire held-out test split against unigram and random baselines, and served in a desktop chat window with a native macOS Liquid Glass backdrop. After setup, nothing touches the network.
02Architecture
A hand-written decoder-only Transformer.
Written directly in PyTorch modules — fused QKV projection, manual scaled dot-product attention with a causal mask, pre-norm residual blocks, a GELU MLP and tied input/output embeddings. No FlashAttention, no borrowed model code.
Input · 256-token window
Illustrative tokens · byte-level BPE, 8,000 vocab
Token embedding
8,000 × 512
+ Learned positions
256 × 512
Transformer block × 8
LayerNorm
pre-norm
Causal self-attention
8 heads × 64 · fused QKV
⊕ residual
LayerNorm
pre-norm
MLP
512 → 2,048 → 512 · GELU
⊕ residual
Causal mask
Each position attends only to itself and the past — the future is masked to −∞ before the softmax.
Final LayerNorm
512
LM head
tied to token embedding
Next-token distribution
softmax over 8,000
03The 29.4M foundation
Exact shape of the largest model.
- Parameters
- 29,447,168
- Layers × heads
- 8 × 8 (head width 64)
- Model width / MLP
- 512 / 2,048 · GELU
- Context
- 256 tokens · learned positions
- Tokenizer
- Byte-level BPE · 8,000 vocab · fit on train only
- Normalization
- Pre-norm LayerNorm · tied embeddings
- Optimizer
- AdamW · grad-clip 1.0 · warmup + cosine
- Batch
- 128 windows × 256 = 32,768 positions / update
- Precision / device
- float32 · Apple GPU (PyTorch MPS)
- Run
- 7,836 updates · 34,318 s (≈ 9.5 h) · 7,482 positions/s
The run stopped at its self-imposed wall-clock ceiling, 7,836 of 10,000 planned updates.
04Training
From noise to 1.68 validation loss.
Sampled training and validation loss of the 29.4M foundation, straight from the run log. It starts at 9.06 — right where a uniform guess over 8,000 tokens sits (ln 8000 ≈ 8.99) — and reaches its best checkpoint at update 7,500.
Source: docs/BROAD_RUN.md · sampled over 800 windows × 256 targets.
05The pipeline
Twelve verifiable steps, from a download to a conversation.
Every step from grouping onward has a verify command that rebuilds its output from verified inputs and checks it matches.
- 01
Pinned download
Exact Hugging Face revisions over HTTPS, streamed under byte and time limits. The only network step.
- 02
Verify
Every file re-hashed offline against SHA-256 digests pinned in code.
- 03
Import
Parquet decoded to JSONL with per-row provenance. Nothing is dropped.
- 04
Group
Rows linked by completion family and exact normalized-text hash; test-touching groups flagged.
- 05
Isolate
Near-duplicates found with 5-word-shingle Jaccard ≥ 0.9 and prefix filtering; test-connected rows excluded.
- 06
Split
Whole groups selected, 5% of groups held for validation, every publisher test row kept out.
- 07
Tokenizer
8,000-token byte-level BPE fit on training stories only, then frozen for every corpus.
- 08
Token arrays
uint32 token streams plus a hashed manifest; a replay verifier re-derives every token.
- 09
Train
From random init on the Apple GPU: AdamW, warmup + cosine, atomic checkpoints, exact resume.
- 10
Evaluate
The whole held-out test stream, scored against unigram and random baselines.
- 11
Scale
445M tokens of test-excluded synthetic dialogue train the 29.4M foundation.
- 12
Chat
A PyQt6 window on native Liquid Glass streams replies from a background worker process.
06Evaluation
Scored on the entire held-out test set.
The 13.8M story model, evaluated on all 5,919,541 next-token targets of the 21,371-document test split — never used for training or checkpoint selection.
Perplexity 18.54 means the model is, on average, about as uncertain as choosing between 18–19 equally likely next tokens — out of a vocabulary of 8,000.
Lower is better. Log scale. Source: docs/STORY_EVALUATION.md.
07Engineering
Built like it has to survive a crash at hour nine.
Atomic checkpoints
Write to a partial file, flush, fsync, rename. An interrupted save can never corrupt the last good checkpoint.
Strict recovery
Checkpoints carry optimizer state, schedule position, every RNG state and data identity. Resume refuses anything that doesn’t match. CPU resume is bit-for-bit identical in tests.
Leakage-aware data
Exact, family and near-duplicate grouping keeps any test-connected document out of training — checked by independent replay.
Process isolation
The UI never imports the numerical stack. Each job runs in a child worker that streams JSON-lines progress and returns a validated receipt.
One job at a time
A project-local advisory lock guarantees a single numerical job owns the GPU; a second one fails fast in under 0.1 s.
Bounded everything
Explicit row, byte, token and time limits with no silent truncation — including a hardened checkpoint loader.
08In numbers
- test functions
- 776
across 63 test files
- lines of Python
- 16,303
in the package itself
- pipeline stages
- 12
each with a replay verifier
- positions trained
- 256.8M
on the 29.4M foundation
09The app
Train, compare and chat — without a terminal.
A compact PyQt6 chat window with a native AppKit Liquid Glass backdrop, plus windows for dataset preparation, training with a live loss chart and time estimate, text completion, and side-by-side run comparison with CSV/JSON export.
broad-main-v1
every-shard causal corpus · from random weights
Model
8 × 512 · 29.4M
Device
Apple GPU (MPS)
Update
7,836 / 10,000
Best val loss
1.677 @ 7,500
Throughput
7,482 pos/s
Elapsed
9 h 32 m
10From the source
Small, readable, deliberate.
def forward(self, hidden: torch.Tensor) -> torch.Tensor:
batch, time, width = hidden.shape
qkv = self.qkv(hidden).reshape(batch, time, 3, self.heads, self.head_width)
query, key, value = qkv.permute(2, 0, 3, 1, 4).unbind(0)
scores = query @ key.transpose(-2, -1) / math.sqrt(self.head_width)
scores = scores.masked_fill(self.future_mask[:time, :time], float("-inf"))
probabilities = self.attention_dropout(scores.softmax(dim=-1))
attended = (probabilities @ value).transpose(1, 2).contiguous().reshape(batch, time, width)
return self.output_dropout(self.projection(attended))def _atomic_save(payload: dict, path: Path) -> None:
"""A failed replacement leaves an existing complete checkpoint intact."""
temporary = path.with_suffix(".partial")
try:
with temporary.open("wb") as handle:
torch.save(payload, handle)
handle.flush()
os.fsync(handle.fileno())
temporary.replace(path)
finally:
temporary.unlink(missing_ok=True)11Roadmap
What’s next.
- 01
Scale the foundation — a larger model trained for longer is the measured path to better answers.
- 02
Longer context — grow beyond the 256-token window toward full multi-turn conversations.
- 03
Sharper replies — continue reply-focused adaptation so answers stay on topic and follow instructions.
Next: MiniGPT Studio is a work in progress — version 1.0 beta, with new runs documented in the repository as they land.
Next project
HYPE