Best open-source Text-to-Speech (TTS) models — SOTA neural voice synthesis, zero-shot cloning, multilingual & expressive speech generation.
Sinapsis AI
community
AI & ML interests
Agentic Platform That helps non-technical people develop AI solutions easily and cares about privacy
Best Open Source models for Audio Classification (emotion, music genre, language ID, etc.)
-
MIT/ast-finetuned-audioset-10-10-0.4593
Audio Classification • 86.6M • Updated • 1.29M • 364 -
speechbrain/emotion-recognition-wav2vec2-IEMOCAP
Audio Classification • Updated • 48.5k • 188 -
laion/clap-htsat-fused
Audio Classification • 0.2B • Updated • 7.15M • 125 -
m-a-p/MERT-v1-330M
Audio Classification • Updated • 569k • 94
-
BAAI/bge-m3
Sentence Similarity • Updated • 34.1M • • 3.39k -
sentence-transformers/all-MiniLM-L6-v2
Sentence Similarity • 22.7M • Updated • 250M • • 5.2k -
google/embeddinggemma-300m
Sentence Similarity • 0.3B • Updated • 2.42M • • 1.83k -
Qwen/Qwen3-Embedding-8B
Feature Extraction • 8B • Updated • 2.83M • • 773
Image to Image
Image To Video
Music & Sound Generation - Best Open Source models (MusicGen, Stable Audio, etc.)
Speech to Text (ASR) - Best Open Source models
-
openai/whisper-large-v3
Automatic Speech Recognition • 2B • Updated • 4.95M • • 6.14k -
openai/whisper-large-v3-turbo
Automatic Speech Recognition • 0.8B • Updated • 7.79M • • 3.24k -
nvidia/parakeet-tdt-0.6b-v2
Automatic Speech Recognition • Updated • 597k • 1.53k -
facebook/seamless-m4t-v2-large
Automatic Speech Recognition • 2B • Updated • 324k • 1.01k
Text to Image
-
Tongyi-MAI/Z-Image-Turbo
Text-to-Image • 6B • Updated • 940k • • 5.11k -
dx8152/Qwen-Edit-2509-Multiple-angles
Image-to-Image • Updated • 236k • • 965 -
black-forest-labs/FLUX.1-dev
Text-to-Image • 12B • Updated • 494k • • 14.1k -
stabilityai/stable-diffusion-xl-base-1.0
Text-to-Image • 3B • Updated • 1.45M • • 8.04k
Text to Video
The idea of this Collection is to gather those interesting models that are Open Source and I can use them in the webpage
-
moonshotai/Kimi-K2-Thinking
Text Generation • 1.1T • Updated • 47.5k • • 1.71k -
meta-llama/Llama-3.1-8B-Instruct
Text Generation • 8B • Updated • 7.55M • • 6.59k -
allenai/Olmo-3-32B-Think
Text Generation • 32B • Updated • 14.6k • 174 -
allenai/Olmo-3-7B-Instruct
Text Generation • 7B • Updated • 450k • • 141
-
Qwen/Qwen3-VL-235B-A22B-Thinking
Image-Text-to-Text • 236B • Updated • 11.7k • • 400 -
Qwen/Qwen3-235B-A22B-Instruct-2507
Text Generation • 235B • Updated • 42.4k • • 793 -
Qwen/Qwen-Image-Edit-2509
Image-to-Image • 20B • Updated • 430k • • 1.23k -
Qwen/Qwen3-VL-8B-Instruct
Image-Text-to-Text • 9B • Updated • 4.75M • • 1.04k
Best open-source Text-to-Speech (TTS) models — SOTA neural voice synthesis, zero-shot cloning, multilingual & expressive speech generation.
Music & Sound Generation - Best Open Source models (MusicGen, Stable Audio, etc.)
Best Open Source models for Audio Classification (emotion, music genre, language ID, etc.)
-
MIT/ast-finetuned-audioset-10-10-0.4593
Audio Classification • 86.6M • Updated • 1.29M • 364 -
speechbrain/emotion-recognition-wav2vec2-IEMOCAP
Audio Classification • Updated • 48.5k • 188 -
laion/clap-htsat-fused
Audio Classification • 0.2B • Updated • 7.15M • 125 -
m-a-p/MERT-v1-330M
Audio Classification • Updated • 569k • 94
Speech to Text (ASR) - Best Open Source models
-
openai/whisper-large-v3
Automatic Speech Recognition • 2B • Updated • 4.95M • • 6.14k -
openai/whisper-large-v3-turbo
Automatic Speech Recognition • 0.8B • Updated • 7.79M • • 3.24k -
nvidia/parakeet-tdt-0.6b-v2
Automatic Speech Recognition • Updated • 597k • 1.53k -
facebook/seamless-m4t-v2-large
Automatic Speech Recognition • 2B • Updated • 324k • 1.01k
-
BAAI/bge-m3
Sentence Similarity • Updated • 34.1M • • 3.39k -
sentence-transformers/all-MiniLM-L6-v2
Sentence Similarity • 22.7M • Updated • 250M • • 5.2k -
google/embeddinggemma-300m
Sentence Similarity • 0.3B • Updated • 2.42M • • 1.83k -
Qwen/Qwen3-Embedding-8B
Feature Extraction • 8B • Updated • 2.83M • • 773
Text to Image
-
Tongyi-MAI/Z-Image-Turbo
Text-to-Image • 6B • Updated • 940k • • 5.11k -
dx8152/Qwen-Edit-2509-Multiple-angles
Image-to-Image • Updated • 236k • • 965 -
black-forest-labs/FLUX.1-dev
Text-to-Image • 12B • Updated • 494k • • 14.1k -
stabilityai/stable-diffusion-xl-base-1.0
Text-to-Image • 3B • Updated • 1.45M • • 8.04k
Image to Image
Text to Video
Image To Video
The idea of this Collection is to gather those interesting models that are Open Source and I can use them in the webpage
-
moonshotai/Kimi-K2-Thinking
Text Generation • 1.1T • Updated • 47.5k • • 1.71k -
meta-llama/Llama-3.1-8B-Instruct
Text Generation • 8B • Updated • 7.55M • • 6.59k -
allenai/Olmo-3-32B-Think
Text Generation • 32B • Updated • 14.6k • 174 -
allenai/Olmo-3-7B-Instruct
Text Generation • 7B • Updated • 450k • • 141
-
Qwen/Qwen3-VL-235B-A22B-Thinking
Image-Text-to-Text • 236B • Updated • 11.7k • • 400 -
Qwen/Qwen3-235B-A22B-Instruct-2507
Text Generation • 235B • Updated • 42.4k • • 793 -
Qwen/Qwen-Image-Edit-2509
Image-to-Image • 20B • Updated • 430k • • 1.23k -
Qwen/Qwen3-VL-8B-Instruct
Image-Text-to-Text • 9B • Updated • 4.75M • • 1.04k