UniMate

One Unified Model to Animate Diverse Skeletons (SIGGRAPH Asia 2026)

Project Page · Paper · Video · Code · Dataset · Interactive Demo

UniMate teaser

Pretrained checkpoints of UniMate, a text-conditioned flow-matching model that generates motion for skeletons of any topology: animals, humanoids and rigged objects.

Models

Model Architecture Training data Joints Steps Params¹ Status
unimate_uniml3d_f60_v3_preview graph attention, AdaLN text UniML3D d0f19d9 5–70 100k 74.1M Recommended
unimate_uniml3d_f60_v2 graph attention, AdaLN text UniML3D faaa817 5–70 100k 74.1M Previous version
unimate_uniml3d_f60_v2_cross_attn full attention, cross-attention text UniML3D faaa817 5–70 100k 66.2M Variant
unimate_mixamo_f60_v2 graph attention, AdaLN text UniML3D faaa817, Mixamo 22 (one rig) 120k 47.8M Recommended for Mixamo
unimate_truebones_f60_v2 graph attention, AdaLN text UniML3D faaa817, Truebones 5–100 80k 47.8M Recommended for Truebones
unimate_uniml3d_f60_v1_preview graph attention, AdaLN text UniML3D pre-release build 5–60 120k 74.1M Superseded

¹ Denoiser only; the frozen google/flan-t5-base text encoder is downloaded on first use.

The Mixamo model animates the 22-joint Mixamo rig only. The Truebones model covers 73 of the 74 species (all but Dragon), including Bear, Centipede and Monkey, which exceed the UniML3D models' limit. Use each model's latest checkpoint.

Model details

Architecture Transformer denoiser, width 512, 8 heads, 10 layers (6 for the Mixamo and Truebones models); flow matching, linear path, velocity prediction
Inputs English motion description; the skeleton's T-pose, hierarchy and cleaned joint labels
Output 60 frames at 30 fps; per joint: position (3), rotation relative to the T-pose (6), velocity (3)
Text Frozen google/flan-t5-base; classifier-free guidance (10% caption dropout, default scale 3.0)

Quick start

From the root of the code repository, in its unimate environment:

# 1. download a model (for another model, change the folder and the step)
hf download Linzhan/UniMate --local-dir outputs \
    --include "unimate_uniml3d_f60_v3_preview/*.json" --include "unimate_uniml3d_f60_v3_preview/*.npy" \
    --include "unimate_uniml3d_f60_v3_preview/checkpoints/checkpoint_step_100000.pt"

# 2. prepare a rigged GLB / glTF / FBX, animated or not: label its joints, check REVIEW.md, then build
python -m data_process.rig_preprocess run --input robot.glb --output_dir outputs/rig/robot
python -m data_process.rig_preprocess run --input robot.glb --output_dir outputs/rig/robot \
    --annotation outputs/rig/robot/annotation.json

# 3. generate: every prompt on every asset (dataset skeletons by name, e.g. truebones:Horse)
python -m unimate.inference.sample --exp_dir outputs/unimate_uniml3d_f60_v3_preview \
    --asset outputs/rig/robot truebones:Horse --prompt "An object walks forward." --num_repetitions 3

# 4. drive each asset's mesh with its motions (one GLB / FBX per motion)
bash scripts/run_animate_motion.sh outputs/unimate_uniml3d_f60_v3_preview/samples

A dataset skeleton needs only its source's conditioning, plus canonical meshes to drive:

hf download Linzhan/UniML3D --repo-type dataset --local-dir dataset --include "features/*/cond.npy" --include "canonical_assets/mixamo/*"

Test-case files, in-betweening, joint editing and motion expansion are documented in the code README.

Training

AdamW (lr 1e-4, betas 0.9 / 0.99, weight decay 1e-5, clipping 1.0), 3% linear warmup then cosine decay to 5%, EMA 0.9999 (used at inference). Loss: masked L2 flow matching + 0.5 × geodesic rotation + 0.1 × velocity smoothness. 60-frame windows at 30 fps, batch 16 per GPU, with joint addition / removal, pooling and perturbation augmentations (Mixamo model: batch 32, no augmentation).

Model GPUs Time
unimate_uniml3d_f60_v3_preview 8× H100 ≈ 23 h
unimate_uniml3d_f60_v2 8× H100 ≈ 22 h
unimate_uniml3d_f60_v2_cross_attn 8× H100 ≈ 36 h
unimate_mixamo_f60_v2 4× H100 ≈ 12 h
unimate_truebones_f60_v2 4× H100 ≈ 15 h
unimate_uniml3d_f60_v1_preview 6× H100 ≈ 23 h

Each model's config.json is its full training configuration. The code repository's configs/uniml3d_60frames_graph_adaln.json follows the v3_preview recipe, and configs/{mixamo,truebones}_60frames_graph_adaln.json those of the Mixamo and Truebones models:

accelerate launch --num_processes 8 -m unimate.training.train --config configs/uniml3d_60frames_graph_adaln.json

Repository layout

config.json                    index of the models
<model>/
  config.json                  model and data configuration, read by inference
  dataset_stats.npy            normalization statistics
  checkpoints/                 checkpoint_step_<N>.pt every 10k steps (model, EMA, optimizer, scheduler)
  logs/                        TensorBoard event files
  samples/step_<NNNNNN>/       skeleton renders of motions sampled during training

unimate_uniml3d_f60_v3_preview holds its latest three checkpoints, without logs or samples.

Versions

  • v3_preview (2026-10-05), over v2: newer data (UniML3D d0f19d9: re-reviewed annotations; 480 Objaverse-XL rigs authored lying down or upside down stood upright instead of filtered); three captions per clip (generic 35%, detailed 25%, normal 40%); fixed per-dataset sampling weights; one normalization pool for all datasets; logit-normal flow time.
  • v2 (2026-09-28), over v1_preview: UniML3D faaa817; joint limit 70 instead of 60 (98.1% of the clips instead of 86.9%); datasets balanced before object types (Mixamo from under 1% to about 24% of the samples); 100k steps. Includes the cross-attention variant and the Mixamo and Truebones models.
  • v1_preview (2026-09-27): first release, on a pre-release build of UniML3D.

Limitations

  • Contact. No unified contact model: contact-rich motions can slide, drift, hover or penetrate the ground; foot locking or IK post-processing helps where contacts are well defined.
  • Rare topologies and motions. The data is long-tailed toward humanoids and common locomotion; rare skeletons and unusual motions can come out static, jittery or off-prompt.
  • Format. Fixed 60-frame (2 s) samples (longer motion via expansion); a skeleton must fit the model's joint limit; prompts should describe the motion only, as every training caption starts with "An object".

See the paper (Section 6, Appendix F) for details.

License

Checkpoints: CC BY-NC 4.0. Code: MIT. This covers only our rights in the weights; the training data keep their source terms:

Model Training data
unimate_uniml3d_f60_v3_preview, _v2, _v2_cross_attn, _v1_preview Truebones ZOO, Mixamo, Objaverse-XL
unimate_mixamo_f60_v2 Mixamo
unimate_truebones_f60_v2 Truebones ZOO
  • Truebones ZOO: a commercial asset pack by Truebones. Third parties allege that some of its animal assets come from commercial games (among them Resident Evil, Skyrim, Bless Online and Conan Exiles). We have not verified these claims; until they are resolved, treat the provenance of the Truebones-trained models as unclear.
  • Mixamo: Adobe's Mixamo terms of use.
  • Objaverse-XL: the license of each source object; some are non-commercial.

We make no warranty that the checkpoints, or motions generated with them, are free of third-party rights. Rights holders with a concern can open a discussion on this repository or contact the authors.

Citation

@article{mou2026unimate,
  title   = {UniMate: One Unified Model to Animate Diverse Skeletons},
  author  = {Mou, Linzhan and Lei, Jiahui and Dou, Zhiyang and Cai, Chenyue and Song, Chaoyue and Finkelstein, Adam and Rusinkiewicz, Szymon},
  journal = {arXiv preprint arXiv:2609.05415},
  year    = {2026}
}
Downloads last month
340
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 1 Ask for provider support

Dataset used to train Linzhan/UniMate

Space using Linzhan/UniMate 1

Collection including Linzhan/UniMate

Paper for Linzhan/UniMate