fuellabs

A sparse model,
built from scratch.

Fuel Labs documents how HobbyLM is built, trained, and evaluated. Articles publish when their evidence and review are complete.

Explore the work

A model record and its research record, presented together.

All research

Model

The current Fuel Labs language-model project.

HobbyLMSparse mixture of experts

Research record

Technical accounts publish as evidence and review are completed.

Pretraining HobbyLM from scratchPublished
HobbyLM architectureIn preparation
Post-training and SFTIn preparation
Evaluation methodology and resultsEvidence pending

Selective routing

For each token, HobbyLM activates a selected subset of computational paths and combines their outputs before continuing through the model.

A conceptual sparse mixture-of-experts routing trace showing a token entering a router, selected routes, a shared expert, and their combined output. No measured router scores or expert roles are shown.
A conceptual view of selective computation inside HobbyLM.

Inside a
sparse model

HobbyLM combines one dense decoder layer with 19 sparse mixture-of-experts layers. Each sparse layer selects eight of 64 routed experts per token and evaluates one shared expert on every token.

One dense decoder layer followed by 19 sparse MoE layers.

Run HobbyLM
locally

HobbyLM runs locally on Apple silicon through MLX, with sparse selected-expert computation and KV-cached decoding.

Fuel — zsh
fuel ~ % git clone https://github.com/fuellabs/hobbyLM.git
Cloning into 'hobbyLM'...
fuel ~ % cd hobbyLM && hobbylm-mlx \
    --prompt "Explain sparse routing in one sentence."
HobbyLM
Sparse routing activates only the experts most relevant to each token, reducing computation while preserving model capacity.