catalog v1
Models and engines
Every adapter the desktop app and the local API know about. Status reflects the engine implementation, not whether weights are on your disk.
- TTS
Kokoro 82M
kokoro
Small, fast TTS with a curated voice pack. Runs in real time on CPU.
- TTS
Chatterbox
chatterbox
Zero-shot voice cloning with an emotion exaggeration control and built-in watermarking.
- TTS
XTTS v2
xtts-v2
Multilingual voice cloning from a six second reference clip. Seventeen languages.
- TTS
F5-TTS
f5-tts
Flow-matching TTS with strong zero-shot cloning and natural prosody.
- TTS
Orpheus 3B
orpheus
Llama-based speech LLM with paralinguistic tags such as laughs and sighs.
- TTS
CosyVoice 2
cosyvoice2
Streaming zero-shot TTS with instruction control over style and dialect.
- TTS
Dia 1.6B
dia
Dialogue generation with speaker tags and nonverbal cues in a single pass.
- TTS
Fish Speech 1.5
fish-speech
Multilingual TTS with few-shot cloning across thirteen languages.
- TTS
Parler-TTS Mini
parler-tts
Voice design from a text description: gender, pace, pitch, recording quality.
- TTS
Bark
bark
Generative audio model that produces speech, music snippets and sound effects.
- TTS
Piper
piper
Very fast ONNX voices designed for low-power devices. Large curated voice set.
- TTS
MeloTTS
melotts
CPU real-time multilingual TTS with mixed Chinese and English support.
- TTS
Sesame CSM 1B
csm-1b
Conversational speech model conditioned on prior dialogue turns.
- TTS
Zonos v0.1
zonos
Expressive TTS with explicit controls for pitch, speaking rate and emotion vectors.
- TTS
SpeechT5 TTS
speecht5
Lightweight encoder-decoder TTS using x-vector speaker embeddings.
- TTS
KittenTTS Nano
kitten-tts
Fifteen million parameter TTS that runs anywhere a CPU does.
- TTS
VoxCPM 0.5B
voxcpm
Tokenizer-free TTS with context-aware prosody and voice cloning.
- TTS
IndexTTS 2
indextts2
Duration-controllable zero-shot TTS with disentangled timbre and emotion.
- TTS
Higgs Audio v2
higgs-audio-v2
Audio foundation model for expressive multi-speaker generation.
Connection: localplannedhf.co/bosonai/higgs-audio-v2-generation-3B-basecudarocmmpscpu3B · custom (see model card) - ASR
Whisper large-v3
whisper-large-v3
Reference multilingual transcription and translation model.
- ASR
Whisper large-v3 turbo
whisper-large-v3-turbo
Pruned decoder variant of large-v3 with several times faster decoding.
- ASR
Faster-Whisper
faster-whisper
CTranslate2 port of Whisper with int8 quantization and word timestamps.
- ASR
Distil-Whisper large-v3
distil-whisper
Distilled English Whisper. Roughly six times faster with near-identical WER.
- ASR
Parakeet TDT 0.6B v2
parakeet-tdt
FastConformer TDT model with punctuation and word-level timestamps.
- ASR
Canary 1B Flash
canary
Multitask ASR and speech translation across four languages.
- ASR
Moonshine Base
moonshine
Edge-focused ASR whose compute scales with input length. Good for dictation.
- ASR
Voxtral Mini 3B
voxtral-mini
Speech-language model for transcription plus audio question answering.
- ASR
Kyutai STT 1B
kyutai-stt
Streaming speech-to-text with semantic voice activity detection.
- ASR
SenseVoice Small
sensevoice
ASR with emotion and audio event tags in the transcript.
- ASR
wav2vec 2.0 Base 960h
wav2vec2
CTC baseline for English. Useful for forced alignment and quick checks.
- VC
Seed-VC
seed-vc
Zero-shot voice conversion and singing voice conversion.
- ENH
Resemble Enhance
resemble-enhance
Two-stage denoiser and enhancer for cleaning up reference clips before cloning.
- TTS
OpenAI-compatible TTS
openai-compat-tts
Route speech generation to any server exposing /v1/audio/speech.
Connection: remoteplannedcudarocmmpscpu— · n/a - ASR
OpenAI-compatible ASR
openai-compat-asr
Route transcription to any server exposing /v1/audio/transcriptions.
Connection: remoteplannedcudarocmmpscpu— · n/a
Catalog snapshot from packages/registry/adapters.json, validated in CI. Listing an adapter does not promise it is installed, compatible with your hardware, or licensed for your use case. Check the linked model card.