GLM-4.7-Flash-0.82B-MTP-FP8-Dynamic

FP8_DYNAMIC version of inference-optimization/GLM-4.7-Flash-0.82B-MTP.

Quantization recipe

Uses Transformers 5.17.0 and LLM Compressor PR #3225.

import torch
from compressed_tensors.offload import set_onload_device
from compressed_tensors.quantization import preset_name_to_scheme
from transformers import AutoTokenizer, Glm4MoeLiteForCausalLM

from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier
from llmcompressor.utils import load_context

MODEL_ID = "inference-optimization/GLM-4.7-Flash-0.82B-MTP"
SAVE_DIR = "GLM-4.7-Flash-0.82B-MTP-FP8-Dynamic"

with load_context(Glm4MoeLiteForCausalLM, load_mtp=True):
    model = Glm4MoeLiteForCausalLM.from_pretrained(
        MODEL_ID, dtype=torch.bfloat16, device_map="cpu",
    )
set_onload_device(model, "cuda")

recipe = QuantizationModifier(
    config_groups={
        "mtp": preset_name_to_scheme("FP8_DYNAMIC", targets=[r"re:^mtp\.layers\."]),
        "backbone": preset_name_to_scheme("FP8_DYNAMIC", targets=["Linear"]),
    },
    ignore=["lm_head", r"re:.*\.eh_proj$", r"re:.*\.indexer\..*"],
)

oneshot(model=model, recipe=recipe)
model.save_pretrained(SAVE_DIR)
AutoTokenizer.from_pretrained(MODEL_ID).save_pretrained(SAVE_DIR)

Architecture and tokenizer: zai-org/GLM-4.7-Flash.

Downloads last month
4
Safetensors
Model size
0.8B params
Tensor type
F32
·
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for inference-optimization/GLM-4.7-Flash-0.82B-MTP-FP8-Dynamic

Quantized
(1)
this model

Collections including inference-optimization/GLM-4.7-Flash-0.82B-MTP-FP8-Dynamic