Skip to content
Yuval Gerber
01AI / Deep Learning

MiniGPT Studio

Training a language model from scratch — on a laptop.

Role
Solo — research, engineering, design
Year
2026
Stack
Python 3.12 · PyTorch (MPS) · NumPy · HF tokenizers · PyArrow · PyQt6 · Objective-C / AppKit · pytest · uv
MiniGPT
MiniGPTcausal-assistant-v1 · 500 updatesNew chat

Ask MiniGPT something.

This is a small model trained on this Mac. It answers offline and can be wrong.

Ready.

Message MiniGPT…
Send
MiniGPT

An end-to-end pipeline that trains a decoder-only Transformer from random weights on Apple Silicon, evaluates it honestly, and lets you talk to it in a native-feeling macOS app. No pretrained weights. No hosted model APIs.

parameters
29.4M

29,447,168 — 8 layers, 8 heads, width 512

token corpus
445M

305,062 synthetic conversations, leakage-checked

pretrained weights
0

every parameter starts random

test perplexity
18.54

13.8M story model · vs 469.42 unigram, 8,603.77 random

The idea

Most student “AI projects” call someone else’s model. I wanted to own every step — from the first downloaded byte to the last generated token.

01Overview

Version 1.0.0 beta · GPL-3.0 · 38 commits

A whole language-model lab, running offline on one Mac.

MiniGPT Studio downloads pinned public datasets and checks every file against a recorded SHA-256. It groups related and near-duplicate texts so no family of documents can land in both training and test, fits its own byte-level BPE tokenizer on training text only, and trains a GPT-style model on the Apple GPU with crash-safe checkpoints and exact resume.

The model is then scored on the entire held-out test split against unigram and random baselines, and served in a desktop chat window with a native macOS Liquid Glass backdrop. After setup, nothing touches the network.

02Architecture

A hand-written decoder-only Transformer.

Written directly in PyTorch modules — fused QKV projection, manual scaled dot-product attention with a causal mask, pre-norm residual blocks, a GELU MLP and tied input/output embeddings. No FlashAttention, no borrowed model code.

Input · 256-token window

Once␣upon␣a␣time,␣a␣girl…

Illustrative tokens · byte-level BPE, 8,000 vocab

Token embedding

8,000 × 512

+ Learned positions

256 × 512

Transformer block × 8

LayerNorm

pre-norm

Causal self-attention

8 heads × 64 · fused QKV

⊕ residual

LayerNorm

pre-norm

MLP

512 → 2,048 → 512 · GELU

⊕ residual

× 8

Causal mask

Each position attends only to itself and the past — the future is masked to −∞ before the softmax.

Final LayerNorm

512

LM head

tied to token embedding

Next-token distribution

softmax over 8,000

03The 29.4M foundation

Exact shape of the largest model.

Parameters
29,447,168
Layers × heads
8 × 8 (head width 64)
Model width / MLP
512 / 2,048 · GELU
Context
256 tokens · learned positions
Tokenizer
Byte-level BPE · 8,000 vocab · fit on train only
Normalization
Pre-norm LayerNorm · tied embeddings
Optimizer
AdamW · grad-clip 1.0 · warmup + cosine
Batch
128 windows × 256 = 32,768 positions / update
Precision / device
float32 · Apple GPU (PyTorch MPS)
Run
7,836 updates · 34,318 s (≈ 9.5 h) · 7,482 positions/s

The run stopped at its self-imposed wall-clock ceiling, 7,836 of 10,000 planned updates.

04Training

From noise to 1.68 validation loss.

Sampled training and validation loss of the 29.4M foundation, straight from the run log. It starts at 9.06 — right where a uniform guess over 8,000 tokens sits (ln 8000 ≈ 8.99) — and reaches its best checkpoint at update 7,500.

Training and validation loss of the 29.4M-parameter foundationValidation loss falls from 9.06 at update 1 to a best of 1.677 at update 7,500; training stopped at update 7,836 of 10,000.2.03.04.06.09.002k4k6k8k10k9.06 · random initbest 1.677 @ 7,500stopped at 7,836wall-clock ceilingUPDATES
Validation loss Training lossPlanned, not run

Source: docs/BROAD_RUN.md · sampled over 800 windows × 256 targets.

05The pipeline

Twelve verifiable steps, from a download to a conversation.

Every step from grouping onward has a verify command that rebuilds its output from verified inputs and checks it matches.

  1. 01

    Pinned download

    Exact Hugging Face revisions over HTTPS, streamed under byte and time limits. The only network step.

  2. 02

    Verify

    Every file re-hashed offline against SHA-256 digests pinned in code.

  3. 03

    Import

    Parquet decoded to JSONL with per-row provenance. Nothing is dropped.

  4. 04

    Group

    Rows linked by completion family and exact normalized-text hash; test-touching groups flagged.

  5. 05

    Isolate

    Near-duplicates found with 5-word-shingle Jaccard ≥ 0.9 and prefix filtering; test-connected rows excluded.

  6. 06

    Split

    Whole groups selected, 5% of groups held for validation, every publisher test row kept out.

  7. 07

    Tokenizer

    8,000-token byte-level BPE fit on training stories only, then frozen for every corpus.

  8. 08

    Token arrays

    uint32 token streams plus a hashed manifest; a replay verifier re-derives every token.

  9. 09

    Train

    From random init on the Apple GPU: AdamW, warmup + cosine, atomic checkpoints, exact resume.

  10. 10

    Evaluate

    The whole held-out test stream, scored against unigram and random baselines.

  11. 11

    Scale

    445M tokens of test-excluded synthetic dialogue train the 29.4M foundation.

  12. 12

    Chat

    A PyQt6 window on native Liquid Glass streams replies from a background worker process.

06Evaluation

Scored on the entire held-out test set.

The 13.8M story model, evaluated on all 5,919,541 next-token targets of the 21,371-document test split — never used for training or checkpoint selection.

Random model (same shape)8,603.77
Train-only unigram469.42
MiniGPT story model18.54
1101001,00010,000

Perplexity 18.54 means the model is, on average, about as uncertain as choosing between 18–19 equally likely next tokens — out of a vocabulary of 8,000.

Lower is better. Log scale. Source: docs/STORY_EVALUATION.md.

07Engineering

Built like it has to survive a crash at hour nine.

08In numbers

test functions
776

across 63 test files

lines of Python
16,303

in the package itself

pipeline stages
12

each with a replay verifier

positions trained
256.8M

on the 29.4M foundation

09The app

Train, compare and chat — without a terminal.

A compact PyQt6 chat window with a native AppKit Liquid Glass backdrop, plus windows for dataset preparation, training with a live loss chart and time estimate, text completion, and side-by-side run comparison with CSV/JSON export.

MiniGPT — Train

broad-main-v1

every-shard causal corpus · from random weights

Stop & save

Model

8 × 512 · 29.4M

Device

Apple GPU (MPS)

Update

7,836 / 10,000

Best val loss

1.677 @ 7,500

Throughput

7,482 pos/s

Elapsed

9 h 32 m

Training and validation loss of the 29.4M-parameter foundationValidation loss falls from 9.06 at update 1 to a best of 1.677 at update 7,500; training stopped at update 7,836 of 10,000.2.03.04.06.09.002k4k6k8k10kUPDATES

10From the source

Small, readable, deliberate.

src/minigpt_studio/model.pyview on GitHub ↗
def forward(self, hidden: torch.Tensor) -> torch.Tensor:
    batch, time, width = hidden.shape
    qkv = self.qkv(hidden).reshape(batch, time, 3, self.heads, self.head_width)
    query, key, value = qkv.permute(2, 0, 3, 1, 4).unbind(0)
    scores = query @ key.transpose(-2, -1) / math.sqrt(self.head_width)
    scores = scores.masked_fill(self.future_mask[:time, :time], float("-inf"))
    probabilities = self.attention_dropout(scores.softmax(dim=-1))
    attended = (probabilities @ value).transpose(1, 2).contiguous().reshape(batch, time, width)
    return self.output_dropout(self.projection(attended))
src/minigpt_studio/training.pyview on GitHub ↗
def _atomic_save(payload: dict, path: Path) -> None:
    """A failed replacement leaves an existing complete checkpoint intact."""
    temporary = path.with_suffix(".partial")
    try:
        with temporary.open("wb") as handle:
            torch.save(payload, handle)
            handle.flush()
            os.fsync(handle.fileno())
        temporary.replace(path)
    finally:
        temporary.unlink(missing_ok=True)

11Roadmap

What’s next.

  • 01

    Scale the foundation — a larger model trained for longer is the measured path to better answers.

  • 02

    Longer context — grow beyond the 256-token window toward full multi-turn conversations.

  • 03

    Sharper replies — continue reply-focused adaptation so answers stay on topic and follow instructions.

Next: MiniGPT Studio is a work in progress — version 1.0 beta, with new runs documented in the repository as they land.

Next project

HYPE