Train a language model from nothing, on your own machine.
⚠️ Status: pre-alpha. The whole path works from the command line: train a tokenizer, pack a corpus, pretrain from random weights, continue someone else's checkpoint, fine-tune on conversations or preference pairs, score the result and export it for llama.cpp. The web UI queues most of that but cannot yet prepare conversation or preference data — see Not built yet. What "pre-alpha" means beyond that is that only Metal, CPU and one AMD card have had a real model trained on them; see Hardware.
bloomery demoTrains a language model from random weights on your machine in about a minute, then talks to it. No download, no GPU required.
1/4 generated 6,000 synthetic documents
2/4 trained a 486-token vocabulary
3/4 packed 125,482 training tokens
4/4 training 3M params on mps (bfloat16)
step 300 loss 0.680 15,828 tok/s
trained in 614,400 tokens · final loss 0.680 · val 0.668
samples (the model has never seen these prompts)
Ana found a shy ball in the river. The bear wanted it too. They shared it and felt calm.
One morning Ben walked to the forest. A clever bear was waiting there. Hugo felt proud.
The sleepy cat lost its key. Ben looked in the forest and found it. The dog was excited.
That is a 3-million-parameter model that did not exist a minute earlier. It is a toy trained on a synthetic grammar — note that it loses track of who the story is about, which is exactly what 3M parameters buys you — but every step of the path is the real one.
A bloomery is the earliest kind of iron furnace — the small, buildable one that turned raw ore into usable metal long before industrial blast furnaces existed. Its output is called a bloom: a rough mass of iron, ready to be refined.
That's the idea here. Point it at a folder of text and get back a language model that didn't exist before. Then keep going — continued pretraining, instruction tuning, export — all from the same place, all on hardware you own.
There are good local fine-tuning GUIs. There are good from-scratch pretraining libraries. There is nothing that is both.
| From scratch | Continued pretraining | SFT / RL | Local GUI | |
|---|---|---|---|---|
| Transformer Lab | ❌ | ~ | ✅ | ✅ |
| Unsloth Studio | ❌ | ❌ | ✅ | ✅ |
| LLaMA-Factory | ❌ | ✅ | ✅ | ✅ |
| Oumi | ✅ | ✅ | ✅ | ❌ |
| nanochat | ✅ | ✅ | ✅ | ❌ |
| bloomery | ✅ | ✅ | ✅ | ✅ |
That column was a tilde until preference optimization existed. Supervised
fine-tuning and DPO both work now; DPO is LoRA-only, which is what the hardware
this targets can actually run, and --method full says so rather than failing
at step one.
The tools with the capability are CLI and YAML. The tools with the interface start from someone else's weights. Bloomery is from-scratch first, with fine-tuning as the natural next step rather than the headline.
- Pretrain from random init — train your own byte-level BPE tokenizer, pack a corpus, pick a size with one dial, watch loss come down, then generate from the result. Checkpoints are standard Hugging Face directories.
- Honest pre-flight estimates — VRAM and token budget computed from your
measured hardware, before you start rather than at hour six.
trainrefuses a configuration that will not fit and names what to change. - Measured throughput —
benchruns real training steps instead of guessing from a peak-FLOPS table. - CPU thread caps —
--coresworks identically on every platform. - Replay mixtures — weighted, versioned dataset blends with per-component forgetting detection.
- Export for llama.cpp and Ollama —
exportwrites a GGUF and a Modelfile, at f16, q8_0 or q4_0. LoRA adapters are folded in first, since GGUF has no notion of an adapter. The K-quants need llama.cpp'sllama-quantize, and the command says so rather than pretending otherwise. - Score a checkpoint —
evalruns any checkpoint against any prepared dataset, so two of them can be compared. A training run's validation loss says whether that run was still improving; it cannot say whether this checkpoint is better than that one, because two runs report losses over their own splits. A text corpus is scored by loss and perplexity, a preference corpus by how often the model ranks the better answer higher — reported both as summed log-probability, which is what DPO optimises, and per token, which is what removes the bias toward whichever answer is shorter. No HellaSwag or MMLU: models this size score at chance on them, so the numbers would look like rigour and carry nothing. - Teach it which answer is better —
prepare --preferencepacks{prompt, chosen, rejected}records andadapttrains DPO on them, against a reference that is the same model with its adapters switched off, so no second copy of the weights is needed. The prompt is stored once, because the comparison only means anything if both answers follow a token-identical one. Reported per side rather than as a margin alone: the usual way a DPO run goes wrong is both rewards falling while the margin still rises, which a margin reads as healthy. - Fine-tune on conversations —
prepare --chatpacks a corpus of chat or prompt/completion records and masks the prompts, so the model is scored only on what it was meant to produce. There is no separate command: the dataset carries the objective, soadapttrains it. Blend it with a plain corpus and the plain half is trained on in full, which is replay against forgetting. - Continue an existing model —
adaptcarries on training a checkpoint of yours or anyone else's, every weight or LoRA adapters against a frozen base. It refuses a corpus packed with a tokenizer the model cannot read, which is otherwise a silent failure. Pair it with a replay mixture and the per-component evaluation names whatever starts getting worse. - Web UI and job queue —
bloomery serveopens a page that queues prepare, train, adapt, export, eval and bench jobs, streams their state live, tails their logs and cancels them. Per-job CPU, memory and GPU limits. It binds to localhost by default. The form offers exactly the flags the runner accepts, and a test fails when the two drift apart.
- Preference optimization beyond DPO — IPO, KTO, SimPO, PPO. Not scheduled.
- Preference training from the web UI.
prepare --preferenceand--chatare both absent from the job queue's form, which only offers the flags the runner declares. Both are reachable from the CLI. Fixing one should fix both, so it is a change of its own rather than a half-measure here.
| Platform | Status |
|---|---|
| macOS (Apple Silicon) | training verified on an M1 via Metal, small models only |
| CPU only | training verified; toy models only |
| Linux + NVIDIA | detection covered by tests; training not yet run on real hardware |
| Linux + AMD (ROCm) | training verified on an RX 6700 XT via the gfx1030 override |
| Windows (WSL2) | detection covered by tests; supported path on Windows |
| Windows (native) | job logs bounded live, except across a supervisor restart; training not run on real hardware |
Being precise rather than optimistic: the GPU probe is exercised against captured
vendor output on every CI run across three operating systems, and Metal, CPU and
one AMD card have had a real model trained on them. NVIDIA has not. If you have
NVIDIA hardware, bloomery doctor --json in an issue is genuinely useful.
The AMD run was on Ubuntu 26.04 with archive-packaged ROCm 7.1 and an RX 6700 XT
— a card ROCm does not officially support, which needs
HSA_OVERRIDE_GFX_VERSION=10.3.0. Without that variable the card answers yes to
every question PyTorch can ask and then segfaults on its first matmul, so
doctor reports it as an error and train refuses rather than starting a run
that cannot finish. With it, training runs at roughly 2.5× the speed of this
machine's CPU. An officially supported card (RDNA3 or later) needs none of this.
| Milestone | Scope | Status |
|---|---|---|
| M0 | bloomery doctor — hardware probe, capability estimator, installers |
done |
| M1 | Pretrain from scratch end to end, then generate from it | done |
| M2 | Replay mixtures — weighted, versioned dataset blends | done |
| M3 | Web UI, job queue, cancel/resume, resource limits | done |
| M4 | Continued pretraining on existing models, full or LoRA | done |
| M4.1 | Supervised fine-tuning on conversations | done |
| M5 | Export — GGUF, quantization, Ollama | done |
| M6 | Preference optimization — DPO on chosen/rejected pairs | done |
| M7 | AMD hardening on real hardware, Windows, evaluation | done |
M2 came before the web UI because weighted replay is what makes "keep adding datasets" work instead of quietly degrading the model — and it is the part no comparable tool has.
Needs uv. The script installs it if it's missing.
git clone https://github.com/finsicle/bloomery.git
cd bloomery
./scripts/install.shOn Windows, .\scripts\install.ps1 — though WSL2 is the supported path.
This installs the core package only, which is a few megabytes, then immediately reports what your machine can do. Adding PyTorch is a separate, much larger step, and there is no point spending that download before you know which backend you need:
uv pip install --torch-backend=auto -e ".[train]"--torch-backend=auto inspects your CUDA driver, AMD GPU version or Intel GPU
and resolves the matching wheel index by itself. That one flag is most of what
made dropping Docker viable.
bloomery prepare --name mine --source ./my-text --vocab 8192
bloomery train --data mine --depth 8 --steps 5000
bloomery chat --run run1prepare trains a byte-level BPE tokenizer on your corpus and packs it into
memory-mapped token shards. Do it once per corpus; every model you train on it
reuses the result.
train starts from random weights. --depth is the only shape knob you need
— width, head count, MLP size and learning rate are all derived from it, so
there are not twelve numbers to get wrong. --size d12 picks a named preset
instead.
Checkpoints are ordinary Hugging Face directories. AutoModelForCausalLM.from_pretrained
loads them, which is the whole reason bloomery emits a Llama-architecture model
rather than inventing one — vLLM and other runtimes that read Llama-architecture
Hugging Face checkpoints load them directly. (A LoRA run writes adapters rather
than a whole model; export folds those in.)
GGUF is the exception, and export writes it here rather than shelling out.
llama.cpp's converter identifies a tokenizer by hashing its output against a
table of known models, and a tokenizer bloomery trained is a new hash by
construction — so the stock path refuses precisely the models this project
exists to produce. Refusing is right for a general tool, since the hash picks
the pre-tokenizer and the wrong one yields a model that loads and talks
nonsense. Bloomery already knows the answer, so it fills the field in itself.
train estimates memory before it starts and refuses a configuration that
cannot fit, because an out-of-memory error arrives whenever the allocator
happens to hit the ceiling — which can be well into a run, after the tokenizer,
the packing and the model build have all succeeded.
Reproduce the refusal below on any machine (the numbers on the right depend on your own memory, so yours will differ):
bloomery prepare --name mine --synthetic 2000 --vocab 8192
bloomery train --data mine --depth 12 --batch 8 --seq 512 --steps 1000 --device cpumemory ~3.5 GiB needed of 1.2 GiB system RAM — 2 GiB available, CPU only
error this configuration needs about 3.5 GiB but only 1.2 GiB is available
(system RAM — 2 GiB available, CPU only).
try one of:
--grad-checkpoint (saves ~1.1 GiB, costs ~30% speed)
--batch 4 (~2.8 GiB, still short)
--seq 256 (~2.8 GiB, still short)
--depth 5 (the largest depth that fits as configured)
Every suggestion is re-estimated against your actual budget before being
offered, which is why some are labelled still short — a suggestion that does
not help is worse than none. The estimate is deliberately conservative;
--force starts anyway.
Gradient checkpointing is off by default and the estimate accounts for that, so turning it on is usually the largest single saving available.
bloomery bench --size d12Measures real training throughput on your machine and tells you what a compute-optimal run would actually cost in hours.
Training on corpus A, then B, then C makes the model worse at A. That is catastrophic forgetting, and it is why "just keep adding data" does not work literally. The defence is replay: mix a share of the earlier data back into every later run.
Replay only helps if the blend is a real object you can name, version and reuse. So it is one:
bloomery mix create --name blend --add new:0.8 --replay old:0.2
bloomery train --mix blend --name run1 --depth 8 --steps 5000Weights are raw numbers, normalised on use — 80/20 and 0.8/0.2 are the same
blend. Components you mark --replay are the ones that exist to stop forgetting,
and that share gets reported.
Every component is evaluated on its own held-out split. This is the part that matters. A single aggregate validation loss is dominated by whichever component carries the most weight, so it can fall while the oldest corpus in the blend quietly degrades:
step 150 val loss 0.7004 ppl 2.01 new 0.696 old 0.742
step 200 val loss 0.6913 ppl 2.00 new 0.687 old 0.731
forgetting old is 0.0180 above its best
That warning names the corpus that is getting worse while there is still time to raise its weight. Doing so creates the next version rather than editing the current one:
bloomery mix add blend --replay old:0.3 --note "raise replay after old regressed"
bloomery mix show blend # weights, replay share, and full lineageVersions are immutable on disk. mix show prints the whole chain, so "what was
run 7 actually trained on" has an answer.
Two guards worth knowing about:
- Mismatched tokenizers are refused. Each corpus is packed with its own tokenizer, so blending two of them means token id 4,211 refers to different symbols in each. That would not error — it would train, converge to nothing, and give no clue why. Components must agree on vocabulary size and tokenizer hash.
- A blend with no replay component gets a warning, because that is the configuration that forgets.
The first thing that works, and the thing to paste into a bug report.
╭─ gpus ───────────────────────────────────────────────────────────────────────╮
│ # vendor device memory arch source │
│ 0 amd Radeon RX 7900 XTX 24 GiB gfx1100 amd-smi │
╰──────────────────────────────────────────────────────────────────────────────╯
╭─ what this machine can train ────────────────────────────────────────────────╮
│ budget Radeon RX 7900 XTX — 24 GiB VRAM │
│ │
│ pretrain from scratch up to 1B (1.07B) · needs ~21.5B tokens │
│ fine-tune with QLoRA up to 7B (6.98B) │
│ │
│ model params tokens step Full LoRA QLoRA │
│ tiny 5M 105M 32×512 1.6 GiB 1.5 GiB 1.5 GiB │
│ GPT-2 small class 124M 2.5B 8×1024 4.7 GiB 3.0 GiB 2.8 GiB │
│ nanochat d26 595M 11.9B 4×2048 13 GiB 5.1 GiB 4.2 GiB │
│ 1B 1.07B 21.5B 2×2048 21 GiB 5.8 GiB 4.2 GiB │
│ 7B 6.98B 139.6B 1×4096 117 GiB 20 GiB 9.2 GiB │
╰──────────────────────────────────────────────────────────────────────────────╯
It detects OS (including WSL2), CPU, RAM, free disk, and GPUs across NVIDIA, AMD, Apple and Intel — then estimates, from your measured VRAM, which model sizes are actually trainable. Memory estimates are conservative and assume AdamW, bf16 and gradient checkpointing.
Two things it does that are easy to skip and expensive to omit:
- Reads the PCI bus directly. So "you have no GPU" and "you have a GPU whose driver isn't loaded" produce different messages. The second is the state most new AMD users are actually in.
- Checks the installed PyTorch against the detected hardware. A CPU-only wheel on a machine with two H100s is a silent, expensive mistake.
--json gives machine-readable output. Exit status is non-zero when something
would stop a training run, so it works as a check in a script.
uv pip install -e ".[dev]"
pytest && ruff check . && ruff format --check . && mypy srcThe GPU vendor parsers are fixture-driven, against captured nvidia-smi,
amd-smi and rocm-smi output plus synthetic sysfs trees. That is deliberate:
the AMD path has to be testable and CI-covered on machines with no AMD hardware,
which is most of them.
GNU Affero General Public License v3.0 or later — see LICENSE.
In plain terms: use it, modify it, run it, sell services around it. If you distribute a modified version — or offer one to users over a network — you have to make your source available under the same terms. You cannot build a proprietary product on top of bloomery and keep it closed.
Contributions are accepted under a Contributor License Agreement, which keeps open the possibility of offering commercial licenses to organisations that cannot use AGPL software. You keep the copyright in your work.
Whatever license bloomery ends up under governs the tool only. It does not govern models you produce with it. If you continue-pretrain or fine-tune an existing base model, that base model's terms (Llama Community License, Gemma Terms of Use, and so on) flow through to your result independently. Check them.
Bloomery stands on:
- nanochat (MIT) — the from-scratch reference this project's pretraining path is modelled on
- PyTorch (BSD-3)
- transformers, peft, trl, accelerate (Apache-2.0)
- llama.cpp (MIT) — GGUF export