DiffuRefill-1B (base)
A 1.08B-parameter masked-diffusion language model, pretrained from scratch for $0 on free, preemptible notebook GPUs.
- 200,000 steps, 92.8B tokens, 21 days (2026-09-10 → 2026-10-01)
- ~200 worker launches on free RTX PRO 6000 notebooks that are reclaimed without warning; the whole fleet was lost and rebuilt many times
- The only shared state was this Hub: workers train 150 steps alone, upload, and a restartable merger averages them (DiLoCo)
- 📄 Technical report:
paper/DiffuRefill-1B-technical-report.pdf
This is a base model: no instruction tuning, no safety tuning. It writes fluent English and often gets facts wrong.
What it is
DiffuRefill is a bidirectional transformer trained with the masked-diffusion (MDLM) objective: random tokens are replaced by [MASK] and the model learns to fill them in from both sides. Instead of writing text strictly left to right, it starts from a row of masks and fills it in over several passes, and it can commit several tokens per pass.
| Parameters | 1.08B (tied embeddings) |
| Layers / width / heads | 18 / 2048 / 16 |
| FFN | SwiGLU, 5632 |
| Positions / norm | RoPE / RMSNorm (pre-norm) |
| Context | 2048 |
| Tokenizer | openbmb/MiniCPM4-0.5B |
| Weights | uniform average of the last 130 merged models (steps ~181k–200k) |
Usage
# pip install torch transformers safetensors huggingface_hub
from huggingface_hub import hf_hub_download
import importlib.util, torch
path = hf_hub_download("Asilarkness/DiffuRefill-1B-base", "modeling_diffurefill.py")
spec = importlib.util.spec_from_file_location("modeling_diffurefill", path)
mod = importlib.util.module_from_spec(spec); spec.loader.exec_module(mod)
model = mod.DiffuRefill.from_pretrained("Asilarkness/DiffuRefill-1B-base").cuda()
# one token per pass: slowest, most coherent
print(model.generate("The theory of relativity states that", max_new_tokens=64))
# 4 tokens per pass: ~4x fewer passes, some quality loss
print(model.generate("The theory of relativity states that", max_new_tokens=64, per_pass=4))
# scoring: log p(continuation | context), the scorer used for the benchmarks
print(model.loglikelihood("The capital of France is", " Paris.")) # -2.03
print(model.loglikelihood("The capital of France is", " Berlin.")) # -8.11
Benchmarks
Zero-shot, lm-evaluation-harness, every model run by us with the same harness and settings. Our model is plugged into the harness with a left-to-right chain-rule scorer (eval/dlm_lmeval.py). A Monte-Carlo ELBO estimator, as used by LLaDA, scored slightly lower (HellaSwag 40.8 vs 42.5 on 1000 items), so the scorer does not understate the model.
| Model | Type | Params | Tokens | Train FLOPs | HellaSwag | ARC-e | ARC-c | PIQA | WinoGr. | OBQA | SciQ | LAMBADA | MMLU |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| DiffuRefill-1B | diffusion | 1.08B | 0.09T | 6.0e20 | 34.8 | 52.4 | 25.0 | 62.7 | 50.6 | 27.4 | 89.5 | 34.0 | 26.9 |
| Pythia-1B | AR | 1.0B | 0.3T | 1.8e21 | 47.2 | 56.9 | 27.0 | 69.4 | 54.1 | 31.2 | 83.4 | 55.8 | 23.1 |
| SmolLM2-360M | AR | 0.36B | 4T | 8.6e21 | 56.3 | 70.3 | 38.1 | 71.5 | 58.9 | 37.8 | 91.1 | 53.4 | 25.5 |
| Qwen2.5-0.5B | AR | 0.49B | 18T | 5.3e22 | 52.2 | 64.9 | 31.8 | 69.9 | 56.1 | 35.0 | 93.0 | 52.0 | 47.5 |
| SmolLM2-1.7B | AR | 1.7B | 11T | 1.1e23 | 71.2 | 77.8 | 47.1 | 77.6 | 65.9 | 44.0 | 93.3 | 67.5 | 48.5 |
| Qwen2.5-1.5B | AR | 1.5B | 18T | 1.7e23 | 67.8 | 75.5 | 45.0 | 75.9 | 63.5 | 40.6 | 94.3 | 62.2 | 59.7 |
HellaSwag, ARC-c, PIQA, OBQA: length-normalised accuracy; the rest: accuracy. FLOPs ≈ 6·params·tokens.
Read this honestly. DiffuRefill-1B is clearly behind autoregressive models of its size. Every one of them used 3× (Pythia) to 280× (Qwen2.5-1.5B) more training compute, and masked diffusion is known to need much more compute than autoregression for the same likelihood (≈16× by Nie et al., 2025). It is relatively strong on SciQ (science facts, matching the fact-heavy anneal) and weakest on LAMBADA and WinoGrande (at chance). TinyLlama and OLMo-1B were also run but load incorrectly under the current transformers (e.g. OLMo LAMBADA 0%), so they are left out. All raw numbers are in eval/results.json.
Speed
Greedy generation of 256 tokens after a 32-token prompt, one RTX PRO 6000 (96 GB), bf16. Autoregressive models use Hugging Face generate() with a KV cache; DiffuRefill uses no cache and re-reads the whole row each pass.
| tokens / s | DiffuRefill-1B | SmolLM2-1.7B | Qwen2.5-1.5B |
|---|---|---|---|
| batch 1, 1 token per pass | 193 | 101 | 82 |
| batch 1, 8 tokens per pass | 1536 | – | – |
| batch 32, 1 token per pass | 431 | 3119 | 2612 |
| batch 32, 8 tokens per pass | 3449 | – | – |
At batch 1 (latency) diffusion wins: 1.9× at the same quality setting and 15× at 8 tokens per pass. At batch 32 (throughput) the KV cache wins unless diffusion commits 8 tokens per pass. More tokens per pass costs quality, and optimised AR servers such as vLLM are much faster than HF generate().
How it was trained
| Objective | MDLM masked diffusion, independent token masking; document-masked attention; salient span masking on 25% of rows; 10% of micro-batches as 128-token rows |
| Optimiser | AdamW (0.9, 0.95), wd 0.1, clip 1.0; per-worker batch 64×2048 = 131k tokens |
| Distributed | DiLoCo over the Hub: 150 local steps, plain averaging (outer lr 1, no momentum), 2–8 workers |
| LR | warm-up 500 → 3e-4 (halved to 1.5e-4 early), cosine to 10%, then linear to zero from step 180k |
| Data, steps 0–158.8k | Ultra-FineWeb-L3 (English Multi-Style and QA synthetic), later with 5% generated fact sentences |
| Data, steps 158.8k–200k | anneal mix: Dolmino (OLMo 3 stage 2) 20%, Nemotron Fact-Seeking 15%, Wiki-Rewrite 15%, Ultra-FineWeb-L3 18%, Nemotron-CC-Math 8%, UltraData-Code L3 7%, OpenCoder annealing 8%, UltraData-Math exercises 4%, generated facts 5% |
| Hardware | free RTX PRO 6000 Blackwell (96 GB) notebooks, ~35k tokens/s per GPU |
Lessons (details in the report)
- ✅ Adopt the merged model by replacement. Keeping local drift made 60 of 63 merges worse.
- ✅ Halve the LR for the merge interval. Local drift fell 6× and a 1.3B-token plateau ended.
- ✅ Document masking and 10% short rows helped. Uniform weight averaging was the best final checkpoint.
- ❌ Outer Nesterov momentum, a mid-run switch to Muon, span/PMI masking, phrase tokens and cautious AdamW did not help this run.
- 🔬 On a small test model, writing a chain of thought first and then filling the answer in one parallel pass lifted a multi-step arithmetic task from 45% to 100%. That is the plan for the reasoning stage.
Limitations
This is one run, and interventions were chosen on small probes and a 38M test model. The fact probe overlaps with the Wikipedia-derived anneal data. The model is English only and has no safety tuning, so do not use it for anything that matters without further training.
Citation
@misc{diffurefill2026,
title = {DiffuRefill-1B: Pretraining a Masked Diffusion Language Model on Free, Preemptible Notebook GPUs},
author = {Asilarkness},
year = {2026},
url = {https://huggingface.co/Asilarkness/DiffuRefill-1B-base}
}
Compute was provided free by molab notebooks. Engineering was done with Claude Code (Anthropic) as an assistant. The training run's working repository, with intermediate checkpoints and logs, is Asilarkness/DiffuRefill-1B.
- Downloads last month
- 34