Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 5 additions & 2 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -64,7 +64,7 @@ jobs:

- name: install
shell: bash
run: uv pip install --torch-backend=auto -e ".[dev,train,serve,adapt]"
run: uv pip install --torch-backend=auto -e ".[dev,train,serve,adapt,export]"

# Assert the interpreter and the training stack directly. This replaces an
# earlier trick of failing the build whenever pytest reported any skip,
Expand All @@ -82,6 +82,9 @@ jobs:
# Same reasoning for the adapter tests: they importorskip on peft,
# so a broken adapt install would skip every LoRA test silently.
python -c "import peft; print('adapt extra ok', peft.__version__)"
# And the export tests, for the same reason: a broken gguf install
# would skip every round-trip test rather than failing.
python -c "import gguf; print('export extra ok')"

# No retry. One was added here on the theory that the illegal-instruction
# crashes on Windows were an intermittent fault in the PyTorch wheel that
Expand Down Expand Up @@ -147,7 +150,7 @@ jobs:

# The train extra is needed here too: mypy resolves numpy's bundled stubs,
# and dropping it would turn a type error into an unanalysable import.
- run: uv pip install --torch-backend=auto -e ".[dev,train,serve,adapt]"
- run: uv pip install --torch-backend=auto -e ".[dev,train,serve,adapt,export]"

- name: ruff check
run: ruff check .
Expand Down
22 changes: 17 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,6 +84,10 @@ a milestone are not built yet — see [Roadmap](#roadmap).
- **CPU thread caps** — `--cores` works identically on every platform.
- **Replay mixtures** — weighted, versioned dataset blends with per-component
forgetting detection.
- **Export for llama.cpp and Ollama** — `export` writes a GGUF and a Modelfile,
at f16, q8_0 or q4_0. LoRA adapters are folded in first, since GGUF has no
notion of an adapter. The K-quants need llama.cpp's `llama-quantize`, and the
command says so rather than pretending otherwise.
- **Fine-tune on conversations** — `prepare --chat` packs a corpus of chat or
prompt/completion records and masks the prompts, so the model is scored only
on what it was meant to produce. There is no separate command: the dataset
Expand All @@ -99,8 +103,7 @@ a milestone are not built yet — see [Roadmap](#roadmap).
them. Per-job CPU, memory and GPU limits. It binds to localhost by default.

### Not built yet
- **Preference optimization** — DPO and friends (M5+).
- **Export to GGUF, Ollama, MLX** (M5).
- **Preference optimization** — DPO and friends. Not scheduled.

## Hardware

Expand Down Expand Up @@ -128,7 +131,7 @@ hardware, `bloomery doctor --json` in an issue is genuinely useful.
| **M3** | Web UI, job queue, cancel/resume, resource limits | done |
| **M4** | Continued pretraining on existing models, full or LoRA | done |
| **M4.1** | Supervised fine-tuning on conversations | done |
| **M5** | Export — GGUF, quantization, Ollama | |
| **M5** | Export — GGUF, quantization, Ollama | done |
| **M6** | AMD hardening on real hardware, Windows, evaluation | |

M2 came before the web UI because weighted replay is what makes "keep adding
Expand Down Expand Up @@ -179,8 +182,17 @@ instead.

Checkpoints are ordinary Hugging Face directories. `AutoModelForCausalLM.from_pretrained`
loads them, which is the whole reason bloomery emits a Llama-architecture model
rather than inventing one — GGUF conversion, vLLM and Ollama all work without a
bespoke converter.
rather than inventing one — vLLM and other runtimes that read Llama-architecture
Hugging Face checkpoints load them directly. (A LoRA run writes adapters rather
than a whole model; `export` folds those in.)

GGUF is the exception, and `export` writes it here rather than shelling out.
llama.cpp's converter identifies a tokenizer by hashing its output against a
table of known models, and a tokenizer bloomery trained is a new hash by
construction — so the stock path refuses precisely the models this project
exists to produce. Refusing is right for a general tool, since the hash picks
the pre-tokenizer and the wrong one yields a model that loads and talks
nonsense. Bloomery already knows the answer, so it fills the field in itself.

`train` estimates memory before it starts and refuses a configuration that
cannot fit, because an out-of-memory error arrives whenever the allocator
Expand Down
6 changes: 6 additions & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -67,6 +67,12 @@ serve = [
"pydantic>=2.9",
"websockets>=13.0",
]
# Writing GGUF. The official library from the llama.cpp repository: pure Python,
# and it ships its own type information. Kept out of core, which stays at three.
export = [
"gguf>=0.10",
"numpy>=1.26",
]
# Quantized training. CUDA and ROCm only — no macOS wheels exist.
quant = [
"bitsandbytes>=0.44; sys_platform != 'darwin'",
Expand Down
143 changes: 143 additions & 0 deletions src/bloomery/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@
import json
import os
import sys
from dataclasses import replace
from pathlib import Path
from typing import Any, NoReturn

Expand All @@ -15,6 +16,7 @@

from bloomery import __version__, paths
from bloomery.capability import LADDER_BY_KEY, Method, assess, format_params
from bloomery.export import QUANTIZATIONS
from bloomery.probe import probe_host_report
from bloomery.probe.types import GIB
from bloomery.render import render_report
Expand Down Expand Up @@ -1119,6 +1121,147 @@ def chat(
console.print()


# --------------------------------------------------------------------------- #
# export
# --------------------------------------------------------------------------- #


@app.command()
def export(
run: str | None = typer.Option(None, "--run", "-r", help="Run name to export."),
checkpoint: Path | None = typer.Option(
None, "--checkpoint", "-c", help="Path to a checkpoint directory."
),
name: str | None = typer.Option(
None, "--name", "-n", help="Name for the export. Defaults to the run's."
),
quantize: str = typer.Option(
"f16", "--quantize", "-q", help=f"One of: {', '.join(QUANTIZATIONS)}."
),
as_json: bool = typer.Option(False, "--json"),
) -> None:
"""Write a checkpoint as GGUF, for llama.cpp and Ollama.

Written here rather than by llama.cpp's converter, which cannot read the
models this project exists to produce: it identifies a tokenizer by hashing
its output against a table of known models, and a tokenizer bloomery trained
is a new hash by construction.

LoRA adapters are folded into the weights first, since GGUF has no notion of
an adapter and a runtime would otherwise read the untouched base.
"""
_quiet_transformers()
import shutil

from bloomery.export import GGUF_NAME as gguf_name
from bloomery.export import ExportError, to_gguf, write_modelfile
from bloomery.train import checkpoint as ckpt
from bloomery.train.loop import ModelLoadError, load_model, merge_adapters

if bool(run) == bool(checkpoint):
_die("give exactly one of --run or --checkpoint")
if quantize not in QUANTIZATIONS:
_die(f"unknown quantization {quantize!r}; choose one of: {', '.join(QUANTIZATIONS)}")

target = checkpoint if checkpoint else ckpt.checkpoint_dir(paths.run_dir(run or ""))
# A run saving right now may have been killed mid-promotion, leaving the
# checkpoint beside this path rather than at it. Export runs concurrently
# with training by design, so it is the command most likely to arrive then —
# and for the same reason it must only read the aside copy, never move it:
# a save in mid-promotion looks exactly like an interrupted one.
target = ckpt.resolve(target)
if not target.is_dir():
_die(f"{target} does not exist")

out_name = name or run or target.parent.name
destination = paths.export_dir(out_name)

with console.status(f"loading {target}"):
try:
from bloomery.data import load_tokenizer

model = merge_adapters(load_model(target))
tokenizer = load_tokenizer(target)
except ModelLoadError as exc:
_die(str(exc))
except Exception as exc: # noqa: BLE001 - tokenizer loading raises many types
_die(f"could not read {target}: {exc}")

# Staged then renamed, the same way a checkpoint is written. A half-written
# GGUF looks loadable and is not, and export can run while the checkpoint it
# is reading is being replaced by a training step.
staging = destination.with_name(destination.name + ".tmp")
if staging.exists():
shutil.rmtree(staging)
staging.mkdir(parents=True)

# Cleared on every exit, not only on ExportError. A full disk raises OSError
# and a long quantization pass can take a KeyboardInterrupt, and either would
# otherwise leave a partial GGUF sitting in exports/<name>.tmp.
placed = False
try:
with console.status(f"writing {quantize} gguf"):
result = to_gguf(model, tokenizer, staging / gguf_name, quantization=quantize)
write_modelfile(staging, tokenizer)

# The old export is moved aside rather than deleted, so there is never a
# moment with no export at all — a kill between the two, or a rename that
# fails across filesystems, would otherwise take the previous good GGUF
# and put nothing in its place.
previous = destination.with_name(destination.name + ".previous")
if previous.exists():
shutil.rmtree(previous)
if destination.exists():
destination.rename(previous)
try:
staging.rename(destination)
except OSError:
if previous.exists() and not destination.exists():
previous.rename(destination)
raise
placed = True
shutil.rmtree(previous, ignore_errors=True)
except ExportError as exc:
_die(str(exc))
Comment thread
coderabbitai[bot] marked this conversation as resolved.
except OSError as exc:
_die(f"could not write the export: {exc}")
finally:
if not placed:
shutil.rmtree(staging, ignore_errors=True)

result = replace(result, path=destination / gguf_name)

if as_json:
console.print_json(json.dumps(result.to_dict()))
return

console.print(
f"model [bold]{result.architecture}[/bold] {format_params(result.parameters)} params"
)
console.print(f"format [bold]{result.quantization}[/bold] {result.tensors} tensors")
console.print(f"size [bold]{result.bytes_written / 1e6:,.1f} MB[/bold] → {result.path}")
console.print(f"vocab {result.vocab_size:,} tokens")
# The number most likely to surprise: it is the sequence length the model was
# trained at, and a runtime will treat it as a hard limit.
console.print(
f"context {result.context_length:,} tokens [dim]the length it was trained at[/dim]"
)
if result.unquantized:
console.print(
f"[yellow] {len(result.unquantized)} tensor(s) kept at f16; their shape "
"cannot be blocked for this format[/yellow]"
)
console.print()
console.print(
f"next: [bold]ollama create {paths.slug(out_name)} -f {destination / 'Modelfile'}[/bold]"
)
if quantize != "q4_0":
console.print(
"[dim] smaller still with llama.cpp's llama-quantize, which has the "
"K-quants this cannot write[/dim]"
)


# --------------------------------------------------------------------------- #
# bench
# --------------------------------------------------------------------------- #
Expand Down
Loading