Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
adarshzolekar 's Collections
Reasoning & Agentic Models
Embeddings & Retrieval Models (RAG)
Multimodal AI Models
Audio & Speech Models
Vision Models (Image & Video)
Text & Code Models (NLP)

Audio & Speech Models

updated 24 days ago

Purpose: Speech recognition, text-to-speech, music, audio analysis.

Upvote
2

  • openai/whisper-large-v3

    Automatic Speech Recognition • 2B • Updated Aug 12, 2024 • 4.73M • • 6.32k

  • openai/whisper-large-v3-turbo

    Automatic Speech Recognition • 0.8B • Updated Oct 4, 2024 • 6.77M • • 3.37k

  • hexgrad/Kokoro-82M

    Text-to-Speech • Updated Apr 10, 2025 • 11.6M • • 6.96k

  • Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice

    Text-to-Speech • 2B • Updated Jan 29 • 2.58M • 1.98k
Upvote
2
  • Collection guide
  • Browse collections
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs