Text Generation
Transformers
Safetensors
English
metadiffusion
diffusion
diffusion-lm
ar-to-diffusion
custom_code
Instructions to use CodeSoft/MetaDiffusion-600M-ChatBase with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use CodeSoft/MetaDiffusion-600M-ChatBase with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="CodeSoft/MetaDiffusion-600M-ChatBase", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("CodeSoft/MetaDiffusion-600M-ChatBase", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use CodeSoft/MetaDiffusion-600M-ChatBase with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "CodeSoft/MetaDiffusion-600M-ChatBase" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "CodeSoft/MetaDiffusion-600M-ChatBase", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/CodeSoft/MetaDiffusion-600M-ChatBase
- SGLang
How to use CodeSoft/MetaDiffusion-600M-ChatBase with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "CodeSoft/MetaDiffusion-600M-ChatBase" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "CodeSoft/MetaDiffusion-600M-ChatBase", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "CodeSoft/MetaDiffusion-600M-ChatBase" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "CodeSoft/MetaDiffusion-600M-ChatBase", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use CodeSoft/MetaDiffusion-600M-ChatBase with Docker Model Runner:
docker model run hf.co/CodeSoft/MetaDiffusion-600M-ChatBase
| license: apache-2.0 | |
| datasets: | |
| - HuggingFaceTB/smol-smoltalk | |
| - HuggingFaceH4/no_robots | |
| - nvidia/OpenMathInstruct-2 | |
| language: | |
| - en | |
| base_model: | |
| - Qwen/Qwen3-0.6B | |
| pipeline_tag: text-generation | |
| library_name: transformers | |
| tags: | |
| - metadiffusion | |
| - diffusion | |
| - diffusion-lm | |
| - ar-to-diffusion | |
| # MetaDiffusion-600M-ChatBase | |
| Experimental bidirectional masked-diffusion chat model converted from Qwen3-0.6B via AR-to-diffusion model surgery (28L x 1024W, ~0.82B params, untied head, bf16, 40K-token context (RoPE base 1e6), Apache-2.0). Intended as a base for further SFT, not a production chatbot. | |
| ## What this is | |
| The AR checkpoint becomes the initialization (weights copied, timestep modules zero-init, the [MASK] and seven auxiliary "rainbow" padding rows are mean-initialized); diffusion behavior is learned throughout training. Trained using smol-smoltalk, no_robots, and OpenMathInstruct-2. | |
| ## Architecture | |
| - Blocks: 28 transformer layers, hidden dim 1024, SwiGLU MLP with intermediate 3072, pre-norm RMSNorm (eps 1e-6), QK-norm on. Timestep conditioning is a sinusoidal MLP embedding (1024) feeding per-block adaLN-style scale+shift modulation. | |
| - Attention: GQA with 16 query heads / 8 KV heads, head_dim 128. Bidirectional self-attention with no causal mask. | |
| - Context: 40,960 tokens max (RoPE, base theta 1e6). | |
| - Params: 0.82B total with untied embeddings: embed_tokens 151,677 x 1024 and a separate lm_head of the same size. | |
| - Vocab / IO: 151,677 rows = Qwen3's 151,669 + [MASK] (id 151669) + 7 rainbow padding tokens (151670-151676); pad_token_id is <|endoftext|> (151643), eos is <|im_end|> (151645). bf16 weights, 371 tensors in model.safetensors. | |
| ## Architecture graph | |
| <a href="https://hfviewer.com/CodeSoft/MetaDiffusion-600M-ChatBase?utm_source=huggingface&utm_medium=embedded_model_card&utm_campaign=CodeSoft_MetaDiffusion-600M-ChatBase_card" target="_blank" rel="noopener"> | |
| <img | |
| src="https://hfviewer.com/api/card.svg?source=CodeSoft%2FMetaDiffusion-600M-ChatBase&granularity=0" | |
| alt="Architecture graph for CodeSoft/MetaDiffusion-600M-ChatBase. Open in hfviewer" | |
| width="100%" | |
| /> | |
| </a> | |
| ## Use with Transformers | |
| ```python | |
| import torch | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| repo = "CodeSoft/MetaDiffusion-600M-ChatBase" | |
| m = AutoModelForCausalLM.from_pretrained( | |
| repo, | |
| trust_remote_code=True, | |
| dtype=torch.bfloat16, | |
| ).to("cuda") | |
| tok = AutoTokenizer.from_pretrained( | |
| repo, | |
| subfolder="tokenizer", | |
| trust_remote_code=True, | |
| ) | |
| prompt = tok.apply_chat_template( | |
| [{"role": "user", "content": "hi"}], | |
| tokenize=False, | |
| add_generation_prompt=True, | |
| ) | |
| inputs = tok(prompt, return_tensors="pt").to("cuda") | |
| with torch.inference_mode(): | |
| out = m.generate( | |
| **inputs, | |
| max_new_tokens=100, | |
| ) | |
| print(tok.decode(out[0], skip_special_tokens=True)) | |
| ``` | |
| # Chat with it (chat.py) | |
| ```bash | |
| python chat.py \ | |
| --model-path model.safetensors \ | |
| --tokenizer ./tokenizer \ | |
| --im-end-bias 2.0 --im-end-bias-t 0.3 --watch | |
| ``` | |
| ## Fine-tune (train.py) | |
| ```bash | |
| # 1. Init: convert the AR model to a diffusion init | |
| python convert.py --source Qwen/Qwen3-0.6B \ | |
| --output init/metadiffusion-600M-instruct.pt \ | |
| --tokenizer-out data/tokenizer | |
| # 2. Corpus: smol, opc, math and no_robots, or a local --jsonl of {"messages": [...]} rows. | |
| # --val-fraction holds out a disjoint val set for early stopping. | |
| python prepare_data.py --datasets smol,math --out data \ | |
| --val-fraction 0.05 | |
| # 3. Train (defaults: lr 5e-5, bf16, seq 512, batch auto-detected) | |
| python train.py --init-checkpoint init/metadiffusion-600M-instruct.pt \ | |
| --data-dir data --output-dir checkpoints --max-steps 30000 | |
| # 4. Continue a run: checkpoints carry model + optimizer + scheduler | |
| # state, so --resume-from picks up LR position and momentum exactly | |
| python train.py --init-checkpoint init/metadiffusion-600M-instruct.pt \ | |
| --data-dir data --output-dir checkpoints \ | |
| --resume-from checkpoints_p2/step_20000.pt --max-steps 16000 | |
| # 5. Test, then ship | |
| python chat.py --model-path checkpoints_/step_30000.pt \ | |
| --tokenizer data/tokenizer --watch | |
| python export_hf.py --checkpoint checkpoints/step_30000.pt \ | |
| --tokenizer data/tokenizer --output MetaDiffusion-600M-ChatBase | |
| ``` | |
| ## Limitations | |
| This model is an experimental research checkpoint intended for further fine-tuning and experimentation. It is not optimized for instruction-following, factuality, safety, or production deployment. Behavior may differ substantially from the original Qwen3-0.6B-Instruct model. | |
| ## License | |
| Apache-2.0 | |