audio-cpp commited on
Commit
96367c9
·
verified ·
1 Parent(s): 4afa508

Upload folder using huggingface_hub

Browse files
.gitattributes CHANGED
@@ -71,3 +71,15 @@ HeartMuLa-GGUF/heartmula-f16.gguf filter=lfs diff=lfs merge=lfs -text
71
  HeartMuLa-GGUF/heartmula-q8_0.gguf filter=lfs diff=lfs merge=lfs -text
72
  Qwen3-TTS-12Hz-1.7B-Base-GGUF/qwen3-tts-12hz-1.7b-base-q8_0_v2.gguf filter=lfs diff=lfs merge=lfs -text
73
  Voxtral-Mini-4B-Realtime-2602-GGUF/voxtral-mini-4b-realtime-2602-q4_k.gguf filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
 
 
 
 
 
 
 
 
71
  HeartMuLa-GGUF/heartmula-q8_0.gguf filter=lfs diff=lfs merge=lfs -text
72
  Qwen3-TTS-12Hz-1.7B-Base-GGUF/qwen3-tts-12hz-1.7b-base-q8_0_v2.gguf filter=lfs diff=lfs merge=lfs -text
73
  Voxtral-Mini-4B-Realtime-2602-GGUF/voxtral-mini-4b-realtime-2602-q4_k.gguf filter=lfs diff=lfs merge=lfs -text
74
+ BS-RoFormer-ep368-GGUF/bs-roformer-ep368-q8_0.gguf filter=lfs diff=lfs merge=lfs -text
75
+ Confucius4-TTS-GGUF/confucius4-tts-orig.gguf filter=lfs diff=lfs merge=lfs -text
76
+ DotTTS-SOAR-GGUF/dots-tts-soar-bf16.gguf filter=lfs diff=lfs merge=lfs -text
77
+ DotTTS-SOAR-GGUF/dots-tts-soar-orig.gguf filter=lfs diff=lfs merge=lfs -text
78
+ DramaBox-GGUF/dramabox-q8_0.gguf filter=lfs diff=lfs merge=lfs -text
79
+ Fun-ASR-Nano-2512-GGUF/fun-asr-nano-2512-f16.gguf filter=lfs diff=lfs merge=lfs -text
80
+ Fun-ASR-Nano-2512-GGUF/fun-asr-nano-2512-q8_0.gguf filter=lfs diff=lfs merge=lfs -text
81
+ Inflect-Micro-v2-GGUF/inflect-micro-v2-orig.gguf filter=lfs diff=lfs merge=lfs -text
82
+ Kroko-ASR-GGUF/kroko-en-community-64-l-q8_0.gguf filter=lfs diff=lfs merge=lfs -text
83
+ Parakeet-TDT-0.6B-v3-GGUF/parakeet-tdt-0.6b-v3-f16.gguf filter=lfs diff=lfs merge=lfs -text
84
+ Parakeet-TDT-0.6B-v3-GGUF/parakeet-tdt-0.6b-v3-q8_0.gguf filter=lfs diff=lfs merge=lfs -text
85
+ RVC-GGUF/rvc-f16.gguf filter=lfs diff=lfs merge=lfs -text
BS-RoFormer-ep368-GGUF/bs-roformer-ep368-q8_0.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9a55a8cad369d00f6e0fb208bb0cd87e30e25430772b8491e20a4eace6423ad2
3
+ size 172532256
Confucius4-TTS-GGUF/confucius4-tts-orig.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:eec4ab3fae3cda1e8ec2b96599442def30dbc980ea817f42b94e073158536244
3
+ size 8192757760
DotTTS-SOAR-GGUF/dots-tts-soar-bf16.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:5292089b053ffa55dddaa3eb5e888c7adefee75006f8d0e6c717a13c4933a41c
3
+ size 4788783360
DotTTS-SOAR-GGUF/dots-tts-soar-orig.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d4406cc394b9c25f5df92e3c053f64013cad1e2286271bc4509b6db795d13e67
3
+ size 5165040864
DramaBox-GGUF/dramabox-q8_0.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:75e7e80fc748defb188cb902c34c62bc12539a7bba477215dccf59a7218a451e
3
+ size 18942803808
Fun-ASR-Nano-2512-GGUF/MANIFEST.json ADDED
@@ -0,0 +1,19 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "source": {
3
+ "repo": "FunAudioLLM/Fun-ASR-Nano-2512-hf",
4
+ "revision": "854d88f94205cd17d2afdb24332130d86fbe654a",
5
+ "model_safetensors_sha256": "335ca3e74917f1156690400e2c344350112950165789cf78ce3d0a367affd821"
6
+ },
7
+ "artifacts": [
8
+ {
9
+ "file": "fun-asr-nano-2512-q8_0.gguf",
10
+ "bytes": 1045334432,
11
+ "sha256": "4d727357574b079b7f43336b2930f39da086ca02f5d8d50872090b4c1c3d5e0a"
12
+ },
13
+ {
14
+ "file": "fun-asr-nano-2512-f16.gguf",
15
+ "bytes": 1675708832,
16
+ "sha256": "3d906c3ccfed07efef88ff53d6cc94b788b9d2edf1492a5679d041b43e98c5be"
17
+ }
18
+ ]
19
+ }
Fun-ASR-Nano-2512-GGUF/README.md ADDED
@@ -0,0 +1,73 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ library_name: audio.cpp
3
+ license: other
4
+ license_name: funasr-model-license-1.1
5
+ license_link: https://huggingface.co/FunAudioLLM/Fun-ASR-Nano-2512-hf/blob/main/README.md
6
+ base_model: FunAudioLLM/Fun-ASR-Nano-2512-hf
7
+ pipeline_tag: automatic-speech-recognition
8
+ tags:
9
+ - audio.cpp
10
+ - gguf
11
+ - speech-recognition
12
+ - multilingual
13
+ - funasr
14
+ ---
15
+
16
+ # Fun-ASR-Nano-2512 GGUF
17
+
18
+ Standalone audio.cpp GGUF builds of
19
+ [FunAudioLLM/Fun-ASR-Nano-2512-hf](https://huggingface.co/FunAudioLLM/Fun-ASR-Nano-2512-hf).
20
+ Each file embeds the model configuration, processor configuration, tokenizer,
21
+ chat template, and the audio.cpp model package specification.
22
+
23
+ ## Files
24
+
25
+ | File | Size | SHA256 |
26
+ | --- | ---: | --- |
27
+ | `fun-asr-nano-2512-q8_0.gguf` | 1,045,334,432 bytes | `4d727357574b079b7f43336b2930f39da086ca02f5d8d50872090b4c1c3d5e0a` |
28
+ | `fun-asr-nano-2512-f16.gguf` | 1,675,708,832 bytes | `3d906c3ccfed07efef88ff53d6cc94b788b9d2edf1492a5679d041b43e98c5be` |
29
+
30
+ The source checkpoint is pinned to revision
31
+ `854d88f94205cd17d2afdb24332130d86fbe654a`. The source
32
+ `model.safetensors` SHA256 is
33
+ `335ca3e74917f1156690400e2c344350112950165789cf78ce3d0a367affd821`.
34
+
35
+ ## audio.cpp
36
+
37
+ ```bash
38
+ audiocpp_cli \
39
+ --task asr \
40
+ --family fun_asr_nano \
41
+ --model fun-asr-nano-2512-q8_0.gguf \
42
+ --backend cuda \
43
+ --audio speech.wav
44
+ ```
45
+
46
+ Fun-ASR-Nano currently provides offline multilingual ASR. It does not expose
47
+ streaming or timestamp output. On CUDA, audio.cpp keeps the Q8_0 encoder and
48
+ adaptor weights native and loads decoder weights as BF16 by default for stable
49
+ logits. An explicit `fun_asr_nano.decoder_weight_type` session option overrides
50
+ that default.
51
+
52
+ ## Reproducibility
53
+
54
+ The files were generated with audio.cpp's `audiocpp_gguf` converter:
55
+
56
+ ```bash
57
+ audiocpp_gguf \
58
+ --input model.safetensors \
59
+ --root /path/to/Fun-ASR-Nano-2512-hf \
60
+ --output fun-asr-nano-2512-q8_0.gguf \
61
+ --type q8_0 \
62
+ --family fun_asr_nano \
63
+ --model-spec model_specs/fun_asr_nano.json
64
+ ```
65
+
66
+ Both formats were checked with `audiocpp_gguf --inspect` and full reference
67
+ audio transcription on CPU and NVIDIA H100 CUDA.
68
+
69
+ ## License
70
+
71
+ The original model and these converted weights are governed by the
72
+ FunASR Model Open Source License Agreement v1.1 distributed with the source
73
+ model. Review that agreement before using or redistributing the files.
Fun-ASR-Nano-2512-GGUF/fun-asr-nano-2512-f16.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3d906c3ccfed07efef88ff53d6cc94b788b9d2edf1492a5679d041b43e98c5be
3
+ size 1675708832
Fun-ASR-Nano-2512-GGUF/fun-asr-nano-2512-q8_0.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4d727357574b079b7f43336b2930f39da086ca02f5d8d50872090b4c1c3d5e0a
3
+ size 1045334432
Inflect-Micro-v2-GGUF/inflect-micro-v2-orig.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d4af1cb6a92cdd8be550e8e7c25805ece222ec0f8e75daf26fc00b4e04ef4b03
3
+ size 72082176
Kroko-ASR-GGUF/kroko-en-community-64-l-q8_0.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:596c749fcb6c582f2d7aaa9aefcff2929e405867a4d04864a55b2c9c2963baa5
3
+ size 167756928
Parakeet-TDT-0.6B-v3-GGUF/parakeet-tdt-0.6b-v3-f16.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:83622c44275694039e543e135ac91b40d94f48fc80a1bf187b5a1cbcd5dd8649
3
+ size 1255384320
Parakeet-TDT-0.6B-v3-GGUF/parakeet-tdt-0.6b-v3-q8_0.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:074e61ac1abd3d3efcfc10798d17bf9a975b31768b466fc2578364d206dde64c
3
+ size 915733744
README.md CHANGED
@@ -23,7 +23,10 @@ base_model:
23
  - Aratako/MioCodec-25Hz-44.1kHz-v2
24
  - Aratako/MioTTS-1.7B
25
  - Aratako/Semantic-DACVAE-Japanese-32dim
 
 
26
  - fishaudio/s2-pro
 
27
  - HeartMuLa/HeartCodec-oss-20260123
28
  - HeartMuLa/HeartMuLa-oss-3B
29
  - HeartMuLa/HeartMuLaGen
@@ -49,6 +52,7 @@ base_model:
49
  - microsoft/VibeVoice-1.5B
50
  - microsoft/VibeVoice-ASR
51
  - mistralai/Voxtral-Mini-4B-Realtime-2602
 
52
  - mlx-community/SeedVC-MLX
53
  - mlx-community/index-tts2-mlx
54
  - mlx-community/mel-roformer-mlx
@@ -56,6 +60,8 @@ base_model:
56
  - mlx-community/wavlm-base-plus-mlx
57
  - nvidia/diar_sortformer_4spk-v1
58
  - nvidia/nemotron-3.5-asr-streaming-0.6b
 
 
59
  - stabilityai/stable-audio-3-medium
60
  - stabilityai/stable-audio-3-small-music
61
  - stabilityai/stable-audio-3-small-sfx
@@ -80,16 +86,23 @@ for the full matrix and drift notes.
80
  | Directory | Files | audio.cpp family | Tested | Original model license |
81
  |---|---|---|---|---|
82
  | `ACE-Step1.5-GGUF` | `base/ace-step-1.5-base-bf16.gguf`, `base/ace-step-1.5-base-q8_0.gguf`, `turbo/ace-step-1.5-turbo-bf16.gguf`, `turbo/ace-step-1.5-turbo-q8_0.gguf` | `ace_step` | 16-bit + Q8 drift | See original model package |
 
83
  | `Chatterbox-GGUF` | `chatterbox-f16.gguf`, `chatterbox-q8_0.gguf` | `chatterbox` | 16-bit + Q8 ASR-match drift | MIT |
84
  | `Citrinet-ASR-GGUF` | `citrinet-asr-q8_0.gguf` | `citrinet_asr` | Q8 pass | CC-BY-4.0 |
 
 
 
85
  | `Fish-Audio-S2-Pro-GGUF` | `fish-audio-s2-pro-bf16.gguf`, `fish-audio-s2-pro-q8_0.gguf` | `fish_audio` | 16-bit + Q8 pass | See original model package |
 
86
  | `HeartMuLa-GGUF` | `heartmula-f16.gguf`, `heartmula-q8_0.gguf` | `heartmula` | 16-bit + Q8 drift | Apache-2.0 |
87
  | `HTDemucs-GGUF` | `htdemucs-f16.gguf`, `htdemucs-q8_0.gguf` | `htdemucs` | 16-bit pass, Q8 drift | See original model package |
88
  | `Higgs-Audio-v3-STT-GGUF` | `higgs-audio-v3-stt-f16.gguf`, `higgs-audio-v3-stt-q8_0.gguf` | `higgs_audio_stt` | 16-bit + Q8 pass | Apache-2.0 |
89
  | `Higgs-Audio-v3-TTS-4B-GGUF` | `higgs-audio-v3-tts-4b-bf16.gguf`, `higgs-audio-v3-tts-4b-q8_0.gguf` | `higgs_audio_tts` | 16-bit + Q8 pass | See original model package |
90
  | `IndexTTS2-GGUF` | `index-tts2-orig.gguf`, `index-tts2-f16.gguf`, `index-tts2-q8_0.gguf` | `index_tts2` | orig + 16-bit pass/drift, Q8 ASR-match drift | bilibili Model Use License Agreement |
 
91
  | `Irodori-TTS-500M-v3-GGUF` | `irodori-tts-500m-v3-f16.gguf`, `irodori-tts-500m-v3-q8_0.gguf` | `irodori_tts` | 16-bit pass, Q8 drift | MIT |
92
  | `Irodori-TTS-600M-v3-VoiceDesign-GGUF` | `irodori-tts-600m-v3-voicedesign-f16.gguf`, `irodori-tts-600m-v3-voicedesign-q8_0.gguf` | `irodori_tts` | 16-bit pass, Q8 drift | MIT |
 
93
  | `MOSS-TTS-Local-v1.5-GGUF` | `moss-tts-local-v1.5-bf16.gguf`, `moss-tts-local-v1.5-q8_0.gguf` | `moss_tts_local` | 16-bit pass, Q8 ASR-match drift | Apache-2.0 |
94
  | `MOSS-TTS-Nano-100M-GGUF` | `moss-tts-nano-100m-bf16.gguf`, `moss-tts-nano-100m-q8_0.gguf` | `moss_tts_nano` | 16-bit pass, Q8 ASR-match drift | Apache-2.0 |
95
  | `Mel-Band-RoFormer-GGUF` | `mel-band-roformer-f16.gguf`, `mel-band-roformer-q8_0.gguf` | `mel_band_roformer` | 16-bit + Q8 drift | MIT |
@@ -97,13 +110,15 @@ for the full matrix and drift notes.
97
  | `MioTTS-1.7B-GGUF` | `miotts-1.7b-orig.gguf`, `miotts-1.7b-bf16.gguf`, `miotts-1.7b-q8_0.gguf` | `miotts` | orig pass, 16-bit drift, Q8 ASR-match drift | Apache-2.0 |
98
  | `Nemotron-3.5-ASR-Streaming-0.6B-GGUF` | `nemotron-3.5-asr-streaming-0.6b-f16.gguf`, `nemotron-3.5-asr-streaming-0.6b-q8_0.gguf` | `nemotron_asr` | 16-bit pass, Q8 minor filler drift | OpenMDW-1.1 |
99
  | `OmniVoice-GGUF` | `omnivoice-bf16.gguf`, `omnivoice-f16.gguf`, `omnivoice-q8_0.gguf` | `omnivoice` | 16-bit + Q8 drift | Apache-2.0 |
 
100
  | `PocketTTS-GGUF` | `english/`, `german/`, `italian/`, `portuguese/`, `spanish/` each contain `bf16` and `q8_0` GGUFs | `pocket_tts` | 16-bit pass, Q8 drift | See original model package |
101
  | `Qwen3-ASR-0.6B-GGUF` | `qwen3-asr-0.6b-f16.gguf`, `qwen3-asr-0.6b-q8_0.gguf` | `qwen3_asr` | 16-bit + Q8 pass | Apache-2.0 |
102
  | `Qwen3-ASR-1.7B-GGUF` | `qwen3-asr-1.7b-f16.gguf`, `qwen3-asr-1.7b-q8_0.gguf` | `qwen3_asr` | 16-bit + Q8 pass | Apache-2.0 |
103
  | `Qwen3-ForcedAligner-0.6B-GGUF` | `qwen3-forced-aligner-0.6b-f16.gguf`, `qwen3-forced-aligner-0.6b-q8_0.gguf` | `qwen3_forced_aligner` | 16-bit + Q8 pass | Apache-2.0 |
104
- | `Qwen3-TTS-12Hz-1.7B-Base-GGUF` | `qwen3-tts-12hz-1.7b-base-orig.gguf`, `qwen3-tts-12hz-1.7b-base-bf16.gguf`, `qwen3-tts-12hz-1.7b-base-q8_0.gguf` | `qwen3_tts` | orig pass, 16-bit + Q8 ASR-match drift | Apache-2.0 |
105
  | `Qwen3-TTS-12Hz-1.7B-CustomVoice-GGUF` | `qwen3-tts-12hz-1.7b-customvoice-bf16.gguf`, `qwen3-tts-12hz-1.7b-customvoice-q8_0.gguf` | `qwen3_tts` | 16-bit + Q8 ASR-match drift | Apache-2.0 |
106
  | `Qwen3-TTS-12Hz-1.7B-VoiceDesign-GGUF` | `qwen3-tts-12hz-1.7b-voicedesign-bf16.gguf`, `qwen3-tts-12hz-1.7b-voicedesign-q8_0.gguf` | `qwen3_tts` | 16-bit + Q8 ASR-match drift | Apache-2.0 |
 
107
  | `SeedVC-MLX-GGUF` | `seed-vc-mlx-orig.gguf`, `seed-vc-mlx-f16.gguf`, `seed-vc-mlx-q8_0.gguf` | `seed_vc` | 16-bit + Q8 drift | GPL-3.0 |
108
  | `Sortformer-Diar-4spk-v1-GGUF` | `sortformer-diar-4spk-v1-f16.gguf`, `sortformer-diar-4spk-v1-q8_0.gguf` | `sortformer_diar` | 16-bit + Q8 pass | CC-BY-NC-4.0 |
109
  | `Stable-Audio-3-Medium-GGUF` | `stable-audio-3-medium-f16.gguf`, `stable-audio-3-medium-q8_0.gguf` | `stable_audio` | 16-bit + Q8 drift | Stability AI Community License |
 
23
  - Aratako/MioCodec-25Hz-44.1kHz-v2
24
  - Aratako/MioTTS-1.7B
25
  - Aratako/Semantic-DACVAE-Japanese-32dim
26
+ - Banafo/Kroko-ASR
27
+ - dots-studio/dots.tts-soar
28
  - fishaudio/s2-pro
29
+ - FunAudioLLM/Fun-ASR-Nano-2512-hf
30
  - HeartMuLa/HeartCodec-oss-20260123
31
  - HeartMuLa/HeartMuLa-oss-3B
32
  - HeartMuLa/HeartMuLaGen
 
52
  - microsoft/VibeVoice-1.5B
53
  - microsoft/VibeVoice-ASR
54
  - mistralai/Voxtral-Mini-4B-Realtime-2602
55
+ - mirek190/audio.cpp
56
  - mlx-community/SeedVC-MLX
57
  - mlx-community/index-tts2-mlx
58
  - mlx-community/mel-roformer-mlx
 
60
  - mlx-community/wavlm-base-plus-mlx
61
  - nvidia/diar_sortformer_4spk-v1
62
  - nvidia/nemotron-3.5-asr-streaming-0.6b
63
+ - nvidia/parakeet-tdt-0.6b-v3
64
+ - owensong/Inflect-Micro-v2
65
  - stabilityai/stable-audio-3-medium
66
  - stabilityai/stable-audio-3-small-music
67
  - stabilityai/stable-audio-3-small-sfx
 
86
  | Directory | Files | audio.cpp family | Tested | Original model license |
87
  |---|---|---|---|---|
88
  | `ACE-Step1.5-GGUF` | `base/ace-step-1.5-base-bf16.gguf`, `base/ace-step-1.5-base-q8_0.gguf`, `turbo/ace-step-1.5-turbo-bf16.gguf`, `turbo/ace-step-1.5-turbo-q8_0.gguf` | `ace_step` | 16-bit + Q8 drift | See original model package |
89
+ | `BS-RoFormer-ep368-GGUF` | `bs-roformer-ep368-q8_0.gguf` | `bs_roformer` | Q8 pass | See original model package |
90
  | `Chatterbox-GGUF` | `chatterbox-f16.gguf`, `chatterbox-q8_0.gguf` | `chatterbox` | 16-bit + Q8 ASR-match drift | MIT |
91
  | `Citrinet-ASR-GGUF` | `citrinet-asr-q8_0.gguf` | `citrinet_asr` | Q8 pass | CC-BY-4.0 |
92
+ | `Confucius4-TTS-GGUF` | `confucius4-tts-orig.gguf` | `confucius4_tts` | orig pass | See original model package |
93
+ | `DotTTS-SOAR-GGUF` | `dots-tts-soar-orig.gguf`, `dots-tts-soar-bf16.gguf` | `dots_tts` | experimental | See original model package |
94
+ | `DramaBox-GGUF` | `dramabox-q8_0.gguf` | `dramabox` | Q8 pass | See original model package |
95
  | `Fish-Audio-S2-Pro-GGUF` | `fish-audio-s2-pro-bf16.gguf`, `fish-audio-s2-pro-q8_0.gguf` | `fish_audio` | 16-bit + Q8 pass | See original model package |
96
+ | `Fun-ASR-Nano-2512-GGUF` | `fun-asr-nano-2512-f16.gguf`, `fun-asr-nano-2512-q8_0.gguf` | `fun_asr_nano` | 16-bit + Q8 pass | FunASR Model Open Source License Agreement v1.1 |
97
  | `HeartMuLa-GGUF` | `heartmula-f16.gguf`, `heartmula-q8_0.gguf` | `heartmula` | 16-bit + Q8 drift | Apache-2.0 |
98
  | `HTDemucs-GGUF` | `htdemucs-f16.gguf`, `htdemucs-q8_0.gguf` | `htdemucs` | 16-bit pass, Q8 drift | See original model package |
99
  | `Higgs-Audio-v3-STT-GGUF` | `higgs-audio-v3-stt-f16.gguf`, `higgs-audio-v3-stt-q8_0.gguf` | `higgs_audio_stt` | 16-bit + Q8 pass | Apache-2.0 |
100
  | `Higgs-Audio-v3-TTS-4B-GGUF` | `higgs-audio-v3-tts-4b-bf16.gguf`, `higgs-audio-v3-tts-4b-q8_0.gguf` | `higgs_audio_tts` | 16-bit + Q8 pass | See original model package |
101
  | `IndexTTS2-GGUF` | `index-tts2-orig.gguf`, `index-tts2-f16.gguf`, `index-tts2-q8_0.gguf` | `index_tts2` | orig + 16-bit pass/drift, Q8 ASR-match drift | bilibili Model Use License Agreement |
102
+ | `Inflect-Micro-v2-GGUF` | `inflect-micro-v2-orig.gguf` | `inflect_v2` | orig pass | See original model package |
103
  | `Irodori-TTS-500M-v3-GGUF` | `irodori-tts-500m-v3-f16.gguf`, `irodori-tts-500m-v3-q8_0.gguf` | `irodori_tts` | 16-bit pass, Q8 drift | MIT |
104
  | `Irodori-TTS-600M-v3-VoiceDesign-GGUF` | `irodori-tts-600m-v3-voicedesign-f16.gguf`, `irodori-tts-600m-v3-voicedesign-q8_0.gguf` | `irodori_tts` | 16-bit pass, Q8 drift | MIT |
105
+ | `Kroko-ASR-GGUF` | `kroko-en-community-64-l-q8_0.gguf` | `kroko_asr` | Q8 pass | See original model package |
106
  | `MOSS-TTS-Local-v1.5-GGUF` | `moss-tts-local-v1.5-bf16.gguf`, `moss-tts-local-v1.5-q8_0.gguf` | `moss_tts_local` | 16-bit pass, Q8 ASR-match drift | Apache-2.0 |
107
  | `MOSS-TTS-Nano-100M-GGUF` | `moss-tts-nano-100m-bf16.gguf`, `moss-tts-nano-100m-q8_0.gguf` | `moss_tts_nano` | 16-bit pass, Q8 ASR-match drift | Apache-2.0 |
108
  | `Mel-Band-RoFormer-GGUF` | `mel-band-roformer-f16.gguf`, `mel-band-roformer-q8_0.gguf` | `mel_band_roformer` | 16-bit + Q8 drift | MIT |
 
110
  | `MioTTS-1.7B-GGUF` | `miotts-1.7b-orig.gguf`, `miotts-1.7b-bf16.gguf`, `miotts-1.7b-q8_0.gguf` | `miotts` | orig pass, 16-bit drift, Q8 ASR-match drift | Apache-2.0 |
111
  | `Nemotron-3.5-ASR-Streaming-0.6B-GGUF` | `nemotron-3.5-asr-streaming-0.6b-f16.gguf`, `nemotron-3.5-asr-streaming-0.6b-q8_0.gguf` | `nemotron_asr` | 16-bit pass, Q8 minor filler drift | OpenMDW-1.1 |
112
  | `OmniVoice-GGUF` | `omnivoice-bf16.gguf`, `omnivoice-f16.gguf`, `omnivoice-q8_0.gguf` | `omnivoice` | 16-bit + Q8 drift | Apache-2.0 |
113
+ | `Parakeet-TDT-0.6B-v3-GGUF` | `parakeet-tdt-0.6b-v3-f16.gguf`, `parakeet-tdt-0.6b-v3-q8_0.gguf` | `parakeet_tdt` | 16-bit + Q8 pass | See original model package |
114
  | `PocketTTS-GGUF` | `english/`, `german/`, `italian/`, `portuguese/`, `spanish/` each contain `bf16` and `q8_0` GGUFs | `pocket_tts` | 16-bit pass, Q8 drift | See original model package |
115
  | `Qwen3-ASR-0.6B-GGUF` | `qwen3-asr-0.6b-f16.gguf`, `qwen3-asr-0.6b-q8_0.gguf` | `qwen3_asr` | 16-bit + Q8 pass | Apache-2.0 |
116
  | `Qwen3-ASR-1.7B-GGUF` | `qwen3-asr-1.7b-f16.gguf`, `qwen3-asr-1.7b-q8_0.gguf` | `qwen3_asr` | 16-bit + Q8 pass | Apache-2.0 |
117
  | `Qwen3-ForcedAligner-0.6B-GGUF` | `qwen3-forced-aligner-0.6b-f16.gguf`, `qwen3-forced-aligner-0.6b-q8_0.gguf` | `qwen3_forced_aligner` | 16-bit + Q8 pass | Apache-2.0 |
118
+ | `Qwen3-TTS-12Hz-1.7B-Base-GGUF` | `qwen3-tts-12hz-1.7b-base-orig.gguf`, `qwen3-tts-12hz-1.7b-base-bf16.gguf`, `qwen3-tts-12hz-1.7b-base-q8_0_v2.gguf` | `qwen3_tts` | orig pass, 16-bit + Q8 ASR-match drift | Apache-2.0 |
119
  | `Qwen3-TTS-12Hz-1.7B-CustomVoice-GGUF` | `qwen3-tts-12hz-1.7b-customvoice-bf16.gguf`, `qwen3-tts-12hz-1.7b-customvoice-q8_0.gguf` | `qwen3_tts` | 16-bit + Q8 ASR-match drift | Apache-2.0 |
120
  | `Qwen3-TTS-12Hz-1.7B-VoiceDesign-GGUF` | `qwen3-tts-12hz-1.7b-voicedesign-bf16.gguf`, `qwen3-tts-12hz-1.7b-voicedesign-q8_0.gguf` | `qwen3_tts` | 16-bit + Q8 ASR-match drift | Apache-2.0 |
121
+ | `RVC-GGUF` | `rvc-f16.gguf` | `rvc` | F16 pass | See original model package |
122
  | `SeedVC-MLX-GGUF` | `seed-vc-mlx-orig.gguf`, `seed-vc-mlx-f16.gguf`, `seed-vc-mlx-q8_0.gguf` | `seed_vc` | 16-bit + Q8 drift | GPL-3.0 |
123
  | `Sortformer-Diar-4spk-v1-GGUF` | `sortformer-diar-4spk-v1-f16.gguf`, `sortformer-diar-4spk-v1-q8_0.gguf` | `sortformer_diar` | 16-bit + Q8 pass | CC-BY-NC-4.0 |
124
  | `Stable-Audio-3-Medium-GGUF` | `stable-audio-3-medium-f16.gguf`, `stable-audio-3-medium-q8_0.gguf` | `stable_audio` | 16-bit + Q8 drift | Stability AI Community License |
RVC-GGUF/rvc-f16.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:5ddfc32ab63199bcca4bcd95a007a8add35ccff7e2c20e1272f19b9617410b57
3
+ size 1260188416
Voxtral-Mini-4B-Realtime-2602-GGUF/README.md ADDED
@@ -0,0 +1,172 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ library_name: audio.cpp
3
+ pipeline_tag: automatic-speech-recognition
4
+ tags:
5
+ - audio.cpp
6
+ - gguf
7
+ - voxtral
8
+ - asr
9
+ - streaming-asr
10
+ ---
11
+
12
+ # Voxtral Mini 4B Realtime GGUF for audio.cpp
13
+
14
+ This repository contains quantized standalone GGUF checkpoints for running Voxtral Mini 4B Realtime ASR with [audio.cpp](https://github.com/0xShug0/audio.cpp). The GGUF files embed the audio.cpp model spec and required sidecars, so the model can be used directly from the checkpoint path without a separate local model-spec directory.
15
+
16
+ ## What audio.cpp does
17
+
18
+ audio.cpp is a young C++/GGML audio inference framework focusing on CUDA performance. It runs speech and audio models locally with CLI and server interfaces, including ASR, TTS, voice conversion, source separation, diarization, VAD, and audio generation models. The project currently tracks 35+ model families and is growing quickly. For Voxtral Realtime, audio.cpp provides offline and streaming speech recognition from a single GGUF checkpoint.
19
+
20
+ ## Files
21
+
22
+ | File | Quantization | Size | Recommended use |
23
+ | --- | --- | ---: | --- |
24
+ | `voxtral-mini-4b-realtime-2602-q8_0.gguf` | Q8_0 | 4.8 GiB | Default balanced checkpoint for high-quality ASR with lower memory than BF16. |
25
+ | `voxtral-mini-4b-realtime-2602-q4_k.gguf` | Q4_K | 2.9 GiB | Lower-memory and faster checkpoint for CUDA testing and deployment. Validate output quality for your domain. |
26
+
27
+ ## Performance
28
+
29
+ Measurements below are from audio.cpp CUDA validation runs. Results vary by GPU, driver, backend, audio length, and decode settings.
30
+
31
+ ### Q8_0 vs BF16 reference
32
+
33
+ BF16 was used only as a local reference baseline for validation. This repository publishes quantized GGUF checkpoints only.
34
+
35
+ | Mode | BF16 reference | Q8_0 | Q8_0 improvement |
36
+ | --- | ---: | ---: | ---: |
37
+ | Offline ASR speed | 11.1x-12.5x realtime | 14.7x-16.7x realtime | 1.31x-1.38x faster |
38
+ | Offline ASR peak VRAM | 10,909 MiB | 7,754 MiB | 3,155 MiB lower |
39
+ | Streaming server TTFT | 207.308 ms | 179.896 ms | 27.412 ms lower |
40
+ | Streaming client TTFT | 550.526 ms | 530.558 ms | 19.968 ms lower |
41
+ | Streaming speed | 4.7x realtime | 5.4x realtime | 1.15x faster |
42
+ | Streaming peak VRAM | 12,616 MiB | 8,972 MiB | 3,644 MiB lower |
43
+
44
+ ### Q4_K quick check
45
+
46
+ | Route | Q8_0 RTF | Q4_K RTF | Q4_K vs Q8_0 |
47
+ | --- | ---: | ---: | ---: |
48
+ | Offline short | 0.0862 | 0.0629 | 1.37x faster |
49
+ | Offline medium | 0.0643 | 0.0476 | 1.35x faster |
50
+ | Offline longer | 0.0576 | 0.0439 | 1.31x faster |
51
+ | Offline sampled | 0.0630 | 0.0500 | 1.26x faster |
52
+ | Streaming path | 0.1036 | 0.0904 | 1.15x faster |
53
+
54
+ In the quick validation set, Q4_K transcripts matched Q8_0 except for one capitalization-only difference.
55
+
56
+ ## Use with audio.cpp
57
+
58
+ Build audio.cpp with the helper script for your platform, then use the generated `audiocpp_cli` binary. The project provides build paths for Linux, Windows, and macOS; see the [audio.cpp README](https://github.com/0xShug0/audio.cpp#build) for the current build matrix and detailed requirements.
59
+
60
+ ```bash
61
+ git clone https://github.com/0xShug0/audio.cpp
62
+ cd audio.cpp
63
+
64
+ # Linux: CUDA, Vulkan, or CPU
65
+ scripts/build_linux.sh --backend cuda --target audiocpp_cli
66
+
67
+ # Windows: CUDA or CPU presets
68
+ powershell.exe -NoProfile -ExecutionPolicy Bypass -File .\scripts\build_windows.ps1 -Preset windows-cuda-release -Target audiocpp_cli
69
+
70
+ # macOS: Metal
71
+ scripts/build_metal.sh --target audiocpp_cli
72
+ ```
73
+
74
+ The examples below assume `audiocpp_cli` is on your `PATH`. You can also replace it with the built binary path for your platform, such as `build/linux-cuda-release/bin/audiocpp_cli`.
75
+
76
+ Run offline ASR with Q8_0:
77
+
78
+ ```bash
79
+ MODEL=/path/to/Voxtral-Mini-4B-Realtime-2602-GGUF/voxtral-mini-4b-realtime-2602-q8_0.gguf
80
+
81
+ audiocpp_cli \
82
+ --task asr \
83
+ --family voxtral_realtime \
84
+ --model "$MODEL" \
85
+ --backend cuda \
86
+ --threads 8 \
87
+ --audio input.wav \
88
+ --text-out transcript.txt
89
+ ```
90
+
91
+ Run the Q4_K checkpoint by changing the model path:
92
+
93
+ ```bash
94
+ MODEL=/path/to/Voxtral-Mini-4B-Realtime-2602-GGUF/voxtral-mini-4b-realtime-2602-q4_k.gguf
95
+ ```
96
+
97
+ ## Streaming ASR
98
+
99
+ Streaming from an audio file:
100
+
101
+ ```bash
102
+ audiocpp_cli \
103
+ --task asr \
104
+ --family voxtral_realtime \
105
+ --model "$MODEL" \
106
+ --backend cuda \
107
+ --threads 8 \
108
+ --mode streaming \
109
+ --audio input.wav \
110
+ --text-out transcript.txt
111
+ ```
112
+
113
+ Streaming raw 16 kHz mono PCM from `ffmpeg`:
114
+
115
+ ```bash
116
+ ffmpeg -i input.mp3 -ar 16000 -ac 1 -f s16le - \
117
+ | audiocpp_cli \
118
+ --task asr \
119
+ --family voxtral_realtime \
120
+ --model "$MODEL" \
121
+ --backend cuda \
122
+ --threads 8 \
123
+ --mode streaming \
124
+ --audio -
125
+ ```
126
+
127
+ For better streaming throughput, batch a few decode steps:
128
+
129
+ ```bash
130
+ audiocpp_cli \
131
+ --task asr \
132
+ --family voxtral_realtime \
133
+ --model "$MODEL" \
134
+ --backend cuda \
135
+ --threads 8 \
136
+ --mode streaming \
137
+ --audio input.wav \
138
+ --session-option voxtral_realtime.stream_batch_tokens=4
139
+ ```
140
+
141
+ `stream_batch_tokens=4` improves throughput by amortizing encoder work, with up to roughly `4 * 80 ms` of additional buffering delay.
142
+
143
+ ## Prompting and decoding
144
+
145
+ audio.cpp exposes Voxtral Realtime as an ASR model. Pass audio to the CLI and audio.cpp builds the model's transcription prompt internally; no free-form text prompt is required for normal transcription.
146
+
147
+ Decode options can be controlled from the request:
148
+
149
+ ```bash
150
+ audiocpp_cli \
151
+ --task asr \
152
+ --family voxtral_realtime \
153
+ --model "$MODEL" \
154
+ --backend cuda \
155
+ --threads 8 \
156
+ --audio input.wav \
157
+ --text-out transcript.txt \
158
+ --request-option max_new_tokens=256 \
159
+ --do-sample false \
160
+ --temperature 1.0 \
161
+ --top-p 1.0 \
162
+ --top-k 50 \
163
+ --seed 1234
164
+ ```
165
+
166
+ ## Notes
167
+
168
+ - Task: `asr`
169
+ - Family: `voxtral_realtime`
170
+ - Supported modes: offline and streaming
171
+ - Timestamp output is not currently exposed by audio.cpp for this model
172
+ - The GGUF package is intended to be standalone: model weights, sidecars, and audio.cpp model spec are embedded in the checkpoint