Instructions to use xingxm/DesignCoder with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use xingxm/DesignCoder with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="xingxm/DesignCoder")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("xingxm/DesignCoder", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use xingxm/DesignCoder with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "xingxm/DesignCoder" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "xingxm/DesignCoder", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/xingxm/DesignCoder
- SGLang
How to use xingxm/DesignCoder with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "xingxm/DesignCoder" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "xingxm/DesignCoder", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "xingxm/DesignCoder" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "xingxm/DesignCoder", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use xingxm/DesignCoder with Docker Model Runner:
docker model run hf.co/xingxm/DesignCoder
Download eval/api_baselines_200/README.md from xingxm/DesignCoder: direct link, hf CLI and curl.
- Browser
- Download file 7.5 kB
-
https://huggingface.co/xingxm/DesignCoder/resolve/main/eval/api_baselines_200/README.md
- Command line
-
hf download hf://xingxm/DesignCoder/eval/api_baselines_200/README.md
-
curl -L -o README.md https://huggingface.co/xingxm/DesignCoder/resolve/main/eval/api_baselines_200/README.md
DesignCoder UI Bench 200 — API 模型对比
7 个 API 模型在 DesignCoder 200 题 UI 生成基准上的产物与评分。 每条记录包含:任务 prompt、模型生成的单文件 HTML、渲染截图,以及三族 rubric 的逐条判定。
评测日期 2026-09-26 · judge = deepseek-v4.1-flash-expires-on-0910 · 生成状态:完整:每个模型 200 题全部生成并评分。
分数
按各模型已完成题目平均(覆盖不同时不可横向比较):
| 模型 | 网关 id | 已生成 | Overall | Overall(no-VSD) | Prompt Fit | Frozen | Landing | Dashboard |
|---|---|---|---|---|---|---|---|---|
| GPT-5.2 | api_azure_openai_gpt-5.2 |
200 | 89.90 | 88.38 | 88.33 | 86.37 | 92.09 | 85.83 |
| DeepSeek-V4 Flash | deepseek-v4-flash |
200 | 89.17 | 87.40 | 74.17 | 71.77 | 86.52 | 94.10 |
| DeepSeek-V4 Pro | deepseek-v4-pro |
200 | 88.89 | 87.19 | 78.03 | 78.52 | 87.45 | 91.56 |
| Claude Opus 4.6 | claude-opus-4-6-v1 |
200 | 86.91 | 84.90 | 73.38 | 73.01 | 82.86 | 94.42 |
| Claude Sonnet 4.6 | claude-sonnet-4.6 |
200 | 86.12 | 83.98 | 68.90 | 67.66 | 82.62 | 92.61 |
| GLM-5.1 | glm-5.1 |
200 | 79.66 | 76.77 | 59.69 | 60.70 | 75.23 | 87.91 |
| Kimi-K2.6 | kimi-k2.6 |
200 | 79.58 | 76.55 | 57.61 | 56.94 | 74.99 | 88.10 |
同题对比(200 题,7 个模型均已完成)
覆盖偏差已排除,这组数字可以横向比较:
| 模型 | Overall | Overall(no-VSD) | Prompt Fit | Frozen |
|---|---|---|---|---|
| GPT-5.2 | 89.90 | 88.38 | 88.33 | 86.37 |
| DeepSeek-V4 Flash | 89.17 | 87.40 | 74.17 | 71.77 |
| DeepSeek-V4 Pro | 88.89 | 87.19 | 78.03 | 78.52 |
| Claude Opus 4.6 | 86.91 | 84.90 | 73.38 | 73.01 |
| Claude Sonnet 4.6 | 86.12 | 83.98 | 68.90 | 67.66 |
| GLM-5.1 | 79.66 | 76.77 | 59.69 | 60.70 |
| Kimi-K2.6 | 79.58 | 76.55 | 57.61 | 56.94 |
覆盖偏差为什么重要:dashboard 题的全模型均分比 landing 高约 8 分。题目按文件顺序生成, 所以跑得慢的模型会集中在靠前的 dashboard 题上而显得偏高。跑满 200 题后该偏差消失。
目录结构
eval/api_baselines_200/metadata.jsonl 每行一个 模型×题目 单元:prompt、分数、产物路径
eval/api_baselines_200/scores_summary.json 按模型/surface/track 的聚合
eval/api_baselines_200/per_case_scores.csv 逐题分数
eval/api_baselines_200/raw/index.jsonl 生成日志(耗时、token、失败归因)
eval/api_baselines_200/raw/shots.jsonl 截图日志(页高、子资源失败、像素判据)
eval/api_baselines_200/raw/judge.jsonl **每条 rubric 的原子判定与理由**
eval/api_baselines_200/pages/<model>/<case_id>.html 模型产物
本仓库未包含截图。 分数是由截图评出的,但 PNG 体积约 750MB 故未上传。
raw/shots.jsonl保留了每张截图的页高、ink(非背景像素占比)与子资源失败数; 用pages/里的 HTML 按下方「评测方法」第 2 步的参数即可重新渲染复现。
用法
import json
from huggingface_hub import hf_hub_download
meta = hf_hub_download("xingxm/DesignCoder", "eval/api_baselines_200/metadata.jsonl",
repo_type="model")
rows = [json.loads(l) for l in open(meta, encoding="utf-8")]
rows[0]["prompt"], rows[0]["overall"], rows[0]["page"]
# 取某一格的产物(路径相对本目录)
page = hf_hub_download("xingxm/DesignCoder", "eval/api_baselines_200/" + rows[0]["page"], repo_type="model")
想换评分口径,不需要重新调用任何模型——raw/judge.jsonl 里是逐条判定,重新聚合即可。
评测方法
- 生成:单轮调用,要求输出单个自包含 HTML(CSS/JS 内联)。dashboard 题额外提供
ECharts / Chart.js / Lucide 的 CDN 地址,以免图表类任务因缺绘图库而失分。
temperature= 1.0,max_completion_tokens= 128000, 不设reasoning_effort(各模型用各自默认档)。 - 截图:headless Chromium,视口 1440×900,整页,高度上限 8000px。
- 评分:
deepseek-v4.1-flash-expires-on-0910作为视觉 judge,分三族 rubric 独立提问,只看截图。 - 聚合:按
rubrics-fixed.json的scoring规范,顶层槽位等权平均。
两种 Overall
规范本身有一处分歧:rubrics-fixed.json 说顶层槽位含「每个生效的 Static 维度」,而
DesignCoder 的 bench_data/README.zh-CN.md 描述 V5.9.2 历史聚合时写的是「除
Visual System Design 外」。两种都给出:overall 按前者,overall_no_vsd 按后者。
⚠️ 与其它口径对照时的两处差异
- judge 不同(最主要的差异)。官方
run_unified_judge.py默认用 Codex CLI 调gpt-5.6-sol;本次用deepseek-v4.1-flash-expires-on-0910,因为运行环境没有 Node / Codex CLI。 跨 judge 的绝对分不可比。 - 生成契约不同。本次是单轮、单文件 HTML;DesignCoder 本体是多轮 agent,带
design_search/websearch工具,可检索设计 DNA 与图片素材。
Frozen 族的条数不是差异:本次每题全量发送 23–25 条 check,与本仓库
eval/rubric_stats.json 记录的口径一致(200 题共 4983 条、每题均值 24.91)。
DesignEvaluator-Skill 的 README 提到 runner 会截断到 10 条,但那不是这批结果的口径。
这 7 个模型之间内部完全可比(同契约、同 judge、同 rubric、同截图流程)。
judge 的已知局限
- 同厂偏好未校验:被测模型中有 DeepSeek 系,judge 也是 DeepSeek 系。
- judge id 带下线日期:
expires-on-0910是下线日期而非版本号,该模型下线后这批分数 无法用同一把尺子复现。每条判定都记了 judge id 与时间。 - judge 默认档位会把思考预算耗尽并截断 JSON,故以
reasoning_effort=low运行。
网关上缺失的模型
原始对比表里有、但网关上调不通的(2026-09-26 实测):
| 模型 | 情况 |
|---|---|
| GPT-5.4 | 两个网关都不可用(2026-09-25 实测)。AI Hub:401 PlatformConfigError 1030「该模型下所有厂商均不支持标准协议」,/standard/v1/ 与 /openapi/v2/ × 带不带 provider=taiji 四种组合全试过。model-eval:api_azure_openai_gpt-5.4-response-async 返回 403「Key based authentication is disabled for this resource」,api_azure_openai_gpt-5.4 返回 500「构建候选集失败: no available account」。 |
| Qwen3.5-4B/9B/27B | 401 UserModelNotFound;试过 qwen3.5-{4b,9b,27b}[-instruct] 等写法。开源权重模型,网关不提供,要本地推理。 |
产物质量守恒
生成成功不等于产物可用。本次发现并作废重跑了 4 类渲染故障(全部集中在一个模型):
HTML 未闭合、输出退化为重复串、布局塌陷(body 高度 0 但 DOM 有上千词)、
loading 遮罩未隐藏。判据以截图像素多样性为准——judge 看的就是像素。
raw/shots.jsonl 里保留了每张截图的 ink(非背景像素占比)、painted_height 与
子资源失败数,便于把低分归因到产物损坏而非设计能力。
来源
题目与 rubric 来自 DesignCoder examples/designcoder/bench_data
(test-prompts-200.jsonl / test-rubrics-200.jsonl / rubrics-fixed.json)。
产物由各模型生成,版权与使用条款遵循各模型提供方的规定。