DiffuRefill-1B (base)

A 1.08B-parameter masked-diffusion language model, pretrained from scratch for $0 on free, preemptible notebook GPUs.

  • 200,000 steps, 92.8B tokens, 21 days (2026-09-10 → 2026-10-01)
  • ~200 worker launches on free RTX PRO 6000 notebooks that are reclaimed without warning; the whole fleet was lost and rebuilt many times
  • The only shared state was this Hub: workers train 150 steps alone, upload, and a restartable merger averages them (DiLoCo)
  • 📄 Technical report: paper/DiffuRefill-1B-technical-report.pdf

This is a base model: no instruction tuning, no safety tuning. It writes fluent English and often gets facts wrong.

What it is

DiffuRefill is a bidirectional transformer trained with the masked-diffusion (MDLM) objective: random tokens are replaced by [MASK] and the model learns to fill them in from both sides. Instead of writing text strictly left to right, it starts from a row of masks and fills it in over several passes, and it can commit several tokens per pass.

Parameters 1.08B (tied embeddings)
Layers / width / heads 18 / 2048 / 16
FFN SwiGLU, 5632
Positions / norm RoPE / RMSNorm (pre-norm)
Context 2048
Tokenizer openbmb/MiniCPM4-0.5B
Weights uniform average of the last 130 merged models (steps ~181k–200k)

Usage

# pip install torch transformers safetensors huggingface_hub
from huggingface_hub import hf_hub_download
import importlib.util, torch

path = hf_hub_download("Asilarkness/DiffuRefill-1B-base", "modeling_diffurefill.py")
spec = importlib.util.spec_from_file_location("modeling_diffurefill", path)
mod = importlib.util.module_from_spec(spec); spec.loader.exec_module(mod)

model = mod.DiffuRefill.from_pretrained("Asilarkness/DiffuRefill-1B-base").cuda()

# one token per pass: slowest, most coherent
print(model.generate("The theory of relativity states that", max_new_tokens=64))
# 4 tokens per pass: ~4x fewer passes, some quality loss
print(model.generate("The theory of relativity states that", max_new_tokens=64, per_pass=4))

# scoring: log p(continuation | context), the scorer used for the benchmarks
print(model.loglikelihood("The capital of France is", " Paris."))   # -2.03
print(model.loglikelihood("The capital of France is", " Berlin."))  # -8.11

Benchmarks

Zero-shot, lm-evaluation-harness, every model run by us with the same harness and settings. Our model is plugged into the harness with a left-to-right chain-rule scorer (eval/dlm_lmeval.py). A Monte-Carlo ELBO estimator, as used by LLaDA, scored slightly lower (HellaSwag 40.8 vs 42.5 on 1000 items), so the scorer does not understate the model.

Model Type Params Tokens Train FLOPs HellaSwag ARC-e ARC-c PIQA WinoGr. OBQA SciQ LAMBADA MMLU
DiffuRefill-1B diffusion 1.08B 0.09T 6.0e20 34.8 52.4 25.0 62.7 50.6 27.4 89.5 34.0 26.9
Pythia-1B AR 1.0B 0.3T 1.8e21 47.2 56.9 27.0 69.4 54.1 31.2 83.4 55.8 23.1
SmolLM2-360M AR 0.36B 4T 8.6e21 56.3 70.3 38.1 71.5 58.9 37.8 91.1 53.4 25.5
Qwen2.5-0.5B AR 0.49B 18T 5.3e22 52.2 64.9 31.8 69.9 56.1 35.0 93.0 52.0 47.5
SmolLM2-1.7B AR 1.7B 11T 1.1e23 71.2 77.8 47.1 77.6 65.9 44.0 93.3 67.5 48.5
Qwen2.5-1.5B AR 1.5B 18T 1.7e23 67.8 75.5 45.0 75.9 63.5 40.6 94.3 62.2 59.7

HellaSwag, ARC-c, PIQA, OBQA: length-normalised accuracy; the rest: accuracy. FLOPs ≈ 6·params·tokens.

Read this honestly. DiffuRefill-1B is clearly behind autoregressive models of its size. Every one of them used 3× (Pythia) to 280× (Qwen2.5-1.5B) more training compute, and masked diffusion is known to need much more compute than autoregression for the same likelihood (≈16× by Nie et al., 2025). It is relatively strong on SciQ (science facts, matching the fact-heavy anneal) and weakest on LAMBADA and WinoGrande (at chance). TinyLlama and OLMo-1B were also run but load incorrectly under the current transformers (e.g. OLMo LAMBADA 0%), so they are left out. All raw numbers are in eval/results.json.

Speed

Greedy generation of 256 tokens after a 32-token prompt, one RTX PRO 6000 (96 GB), bf16. Autoregressive models use Hugging Face generate() with a KV cache; DiffuRefill uses no cache and re-reads the whole row each pass.

tokens / s DiffuRefill-1B SmolLM2-1.7B Qwen2.5-1.5B
batch 1, 1 token per pass 193 101 82
batch 1, 8 tokens per pass 1536 – –
batch 32, 1 token per pass 431 3119 2612
batch 32, 8 tokens per pass 3449 – –

At batch 1 (latency) diffusion wins: 1.9× at the same quality setting and 15× at 8 tokens per pass. At batch 32 (throughput) the KV cache wins unless diffusion commits 8 tokens per pass. More tokens per pass costs quality, and optimised AR servers such as vLLM are much faster than HF generate().

How it was trained

Objective MDLM masked diffusion, independent token masking; document-masked attention; salient span masking on 25% of rows; 10% of micro-batches as 128-token rows
Optimiser AdamW (0.9, 0.95), wd 0.1, clip 1.0; per-worker batch 64×2048 = 131k tokens
Distributed DiLoCo over the Hub: 150 local steps, plain averaging (outer lr 1, no momentum), 2–8 workers
LR warm-up 500 → 3e-4 (halved to 1.5e-4 early), cosine to 10%, then linear to zero from step 180k
Data, steps 0–158.8k Ultra-FineWeb-L3 (English Multi-Style and QA synthetic), later with 5% generated fact sentences
Data, steps 158.8k–200k anneal mix: Dolmino (OLMo 3 stage 2) 20%, Nemotron Fact-Seeking 15%, Wiki-Rewrite 15%, Ultra-FineWeb-L3 18%, Nemotron-CC-Math 8%, UltraData-Code L3 7%, OpenCoder annealing 8%, UltraData-Math exercises 4%, generated facts 5%
Hardware free RTX PRO 6000 Blackwell (96 GB) notebooks, ~35k tokens/s per GPU

Lessons (details in the report)

  • ✅ Adopt the merged model by replacement. Keeping local drift made 60 of 63 merges worse.
  • ✅ Halve the LR for the merge interval. Local drift fell 6× and a 1.3B-token plateau ended.
  • ✅ Document masking and 10% short rows helped. Uniform weight averaging was the best final checkpoint.
  • ❌ Outer Nesterov momentum, a mid-run switch to Muon, span/PMI masking, phrase tokens and cautious AdamW did not help this run.
  • 🔬 On a small test model, writing a chain of thought first and then filling the answer in one parallel pass lifted a multi-step arithmetic task from 45% to 100%. That is the plan for the reasoning stage.

Limitations

This is one run, and interventions were chosen on small probes and a 38M test model. The fact probe overlaps with the Wikipedia-derived anneal data. The model is English only and has no safety tuning, so do not use it for anything that matters without further training.

Citation

@misc{diffurefill2026,
  title  = {DiffuRefill-1B: Pretraining a Masked Diffusion Language Model on Free, Preemptible Notebook GPUs},
  author = {Asilarkness},
  year   = {2026},
  url    = {https://huggingface.co/Asilarkness/DiffuRefill-1B-base}
}

Compute was provided free by molab notebooks. Engineering was done with Claude Code (Anthropic) as an assistant. The training run's working repository, with intermediate checkpoints and logs, is Asilarkness/DiffuRefill-1B.

Downloads last month
34
Safetensors
Model size
1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train Asilarkness/DiffuRefill-1B-base

Paper for Asilarkness/DiffuRefill-1B-base