voice-labs

catalog v1

Models and engines

Every adapter the desktop app and the local API know about. Status reflects the engine implementation, not whether weights are on your disk.

34 adapters
  • Kokoro 82M

    kokoro

    TTS

    Small, fast TTS with a curated voice pack. Runs in real time on CPU.

    Connection: localexperimental
    cudarocmmpscpu
    82M · Apache-2.0
    hf.co/hexgrad/Kokoro-82M
  • Chatterbox

    chatterbox

    TTS

    Zero-shot voice cloning with an emotion exaggeration control and built-in watermarking.

    Connection: localplanned
    cudarocmmpscpu
    0.5B · MIT
    hf.co/ResembleAI/chatterbox
  • XTTS v2

    xtts-v2

    TTS

    Multilingual voice cloning from a six second reference clip. Seventeen languages.

    Connection: localplanned
    cudarocmmpscpu
    467M · CPML (non-commercial)
    hf.co/coqui/XTTS-v2
  • F5-TTS

    f5-tts

    TTS

    Flow-matching TTS with strong zero-shot cloning and natural prosody.

    Connection: localplanned
    cudarocmmpscpu
    336M · CC-BY-NC-4.0
    hf.co/SWivid/F5-TTS
  • Orpheus 3B

    orpheus

    TTS

    Llama-based speech LLM with paralinguistic tags such as laughs and sighs.

    Connection: localplanned
    cudarocmmpscpu
    3B · Apache-2.0
    hf.co/canopylabs/orpheus-3b-0.1-ft
  • CosyVoice 2

    cosyvoice2

    TTS

    Streaming zero-shot TTS with instruction control over style and dialect.

    Connection: localplanned
    cudarocmmpscpu
    0.5B · Apache-2.0
    hf.co/FunAudioLLM/CosyVoice2-0.5B
  • Dia 1.6B

    dia

    TTS

    Dialogue generation with speaker tags and nonverbal cues in a single pass.

    Connection: localplanned
    cudarocmmpscpu
    1.6B · Apache-2.0
    hf.co/nari-labs/Dia-1.6B
  • Fish Speech 1.5

    fish-speech

    TTS

    Multilingual TTS with few-shot cloning across thirteen languages.

    Connection: localplanned
    cudarocmmpscpu
    · CC-BY-NC-SA-4.0
    hf.co/fishaudio/fish-speech-1.5
  • Parler-TTS Mini

    parler-tts

    TTS

    Voice design from a text description: gender, pace, pitch, recording quality.

    Connection: localplanned
    cudarocmmpscpu
    880M · Apache-2.0
    hf.co/parler-tts/parler-tts-mini-v1
  • Bark

    bark

    TTS

    Generative audio model that produces speech, music snippets and sound effects.

    Connection: localplanned
    cudarocmmpscpu
    · MIT
    hf.co/suno/bark
  • Piper

    piper

    TTS

    Very fast ONNX voices designed for low-power devices. Large curated voice set.

    Connection: localplanned
    cudarocmmpscpu
    · MIT
    hf.co/rhasspy/piper-voices
  • MeloTTS

    melotts

    TTS

    CPU real-time multilingual TTS with mixed Chinese and English support.

    Connection: localplanned
    cudarocmmpscpu
    · MIT
    hf.co/myshell-ai/MeloTTS-English
  • Sesame CSM 1B

    csm-1b

    TTS

    Conversational speech model conditioned on prior dialogue turns.

    Connection: localplanned
    cudarocmmpscpu
    1B · Apache-2.0
    hf.co/sesame/csm-1b
  • Zonos v0.1

    zonos

    TTS

    Expressive TTS with explicit controls for pitch, speaking rate and emotion vectors.

    Connection: localplanned
    cudarocmmpscpu
    1.6B · Apache-2.0
    hf.co/Zyphra/Zonos-v0.1-transformer
  • SpeechT5 TTS

    speecht5

    TTS

    Lightweight encoder-decoder TTS using x-vector speaker embeddings.

    Connection: localplanned
    cudarocmmpscpu
    144M · MIT
    hf.co/microsoft/speecht5_tts
  • KittenTTS Nano

    kitten-tts

    TTS

    Fifteen million parameter TTS that runs anywhere a CPU does.

    Connection: localplanned
    cudarocmmpscpu
    15M · Apache-2.0
    hf.co/KittenML/kitten-tts-nano-0.1
  • VoxCPM 0.5B

    voxcpm

    TTS

    Tokenizer-free TTS with context-aware prosody and voice cloning.

    Connection: localplanned
    cudarocmmpscpu
    0.5B · Apache-2.0
    hf.co/openbmb/VoxCPM-0.5B
  • IndexTTS 2

    indextts2

    TTS

    Duration-controllable zero-shot TTS with disentangled timbre and emotion.

    Connection: localplanned
    cudarocmmpscpu
    · custom (see model card)
    hf.co/IndexTeam/IndexTTS-2
  • Higgs Audio v2

    higgs-audio-v2

    TTS

    Audio foundation model for expressive multi-speaker generation.

    Connection: localplanned
    cudarocmmpscpu
    3B · custom (see model card)
    hf.co/bosonai/higgs-audio-v2-generation-3B-base
  • Whisper large-v3

    whisper-large-v3

    ASR

    Reference multilingual transcription and translation model.

    Connection: localplanned
    cudarocmmpscpu
    1.55B · Apache-2.0
    hf.co/openai/whisper-large-v3
  • Whisper large-v3 turbo

    whisper-large-v3-turbo

    ASR

    Pruned decoder variant of large-v3 with several times faster decoding.

    Connection: localplanned
    cudarocmmpscpu
    809M · MIT
    hf.co/openai/whisper-large-v3-turbo
  • Faster-Whisper

    faster-whisper

    ASR

    CTranslate2 port of Whisper with int8 quantization and word timestamps.

    Connection: localexperimental
    cudarocmmpscpu
    1.55B · MIT
    hf.co/Systran/faster-whisper-large-v3
  • Distil-Whisper large-v3

    distil-whisper

    ASR

    Distilled English Whisper. Roughly six times faster with near-identical WER.

    Connection: localplanned
    cudarocmmpscpu
    756M · MIT
    hf.co/distil-whisper/distil-large-v3
  • Parakeet TDT 0.6B v2

    parakeet-tdt

    ASR

    FastConformer TDT model with punctuation and word-level timestamps.

    Connection: localplanned
    cudarocmmpscpu
    0.6B · CC-BY-4.0
    hf.co/nvidia/parakeet-tdt-0.6b-v2
  • Canary 1B Flash

    canary

    ASR

    Multitask ASR and speech translation across four languages.

    Connection: localplanned
    cudarocmmpscpu
    1B · CC-BY-4.0
    hf.co/nvidia/canary-1b-flash
  • Moonshine Base

    moonshine

    ASR

    Edge-focused ASR whose compute scales with input length. Good for dictation.

    Connection: localplanned
    cudarocmmpscpu
    61M · MIT
    hf.co/UsefulSensors/moonshine-base
  • Voxtral Mini 3B

    voxtral-mini

    ASR

    Speech-language model for transcription plus audio question answering.

    Connection: localplanned
    cudarocmmpscpu
    3B · Apache-2.0
    hf.co/mistralai/Voxtral-Mini-3B-2507
  • Kyutai STT 1B

    kyutai-stt

    ASR

    Streaming speech-to-text with semantic voice activity detection.

    Connection: localplanned
    cudarocmmpscpu
    1B · CC-BY-4.0
    hf.co/kyutai/stt-1b-en_fr
  • SenseVoice Small

    sensevoice

    ASR

    ASR with emotion and audio event tags in the transcript.

    Connection: localplanned
    cudarocmmpscpu
    · custom (see model card)
    hf.co/FunAudioLLM/SenseVoiceSmall
  • wav2vec 2.0 Base 960h

    wav2vec2

    ASR

    CTC baseline for English. Useful for forced alignment and quick checks.

    Connection: localplanned
    cudarocmmpscpu
    95M · Apache-2.0
    hf.co/facebook/wav2vec2-base-960h
  • Seed-VC

    seed-vc

    VC

    Zero-shot voice conversion and singing voice conversion.

    Connection: localplanned
    cudarocmmpscpu
    · GPL-3.0
    hf.co/Plachta/Seed-VC
  • Resemble Enhance

    resemble-enhance

    ENH

    Two-stage denoiser and enhancer for cleaning up reference clips before cloning.

    Connection: localplanned
    cudarocmmpscpu
    · MIT
    hf.co/ResembleAI/resemble-enhance
  • OpenAI-compatible TTS

    openai-compat-tts

    TTS

    Route speech generation to any server exposing /v1/audio/speech.

    Connection: remoteplanned
    cudarocmmpscpu
    · n/a
  • OpenAI-compatible ASR

    openai-compat-asr

    ASR

    Route transcription to any server exposing /v1/audio/transcriptions.

    Connection: remoteplanned
    cudarocmmpscpu
    · n/a

Catalog snapshot from packages/registry/adapters.json, validated in CI. Listing an adapter does not promise it is installed, compatible with your hardware, or licensed for your use case. Check the linked model card.