Skip to content

VoiceNeuphonic

Neuphonic

Ultra-low-latency text-to-speech that runs on-device.

Category
Voice
Pricing
FREEMIUM
Source
Open core
Hosting
Hybrid
Platforms
WebAPImacOSWindowsLinux
Models
Self-contained (on-device)
Verified
Jun 15, 2026

Neuphonic is a voice-AI company building text-to-speech and voice cloning that run locally with very low latency. Its cloud API targets real-time voice agents, and in October 2025 it open-sourced NeuTTS Air, a 748M-parameter speech language model that runs on CPU via llama.cpp and clones a voice from a few seconds of audio. Aimed at private, offline, and voice-agent use cases.

Capabilities 2

What it actually does — grouped by capability family.

  • Speech synthesis (TTS) (primary capability)
  • Voice cloning (secondary capability)

Pros & cons

  • On-device, CPU-only synthesis
  • Instant voice cloning from a short sample
  • Self-hostable open model (NeuTTS Air)
  • Very low latency for voice agents
  • Cloud API pricing not clearly published
  • Young company (founded 2024, pre-seed)
  • Open model is 748M — smaller than top cloud voices

Tags

Further reading

View all Voice
  • View Cartesia details
    VoiceFREEMIUM

    Cartesia

    Cartesia

    Low-latency streaming text-to-speech for real-time voice.

    Streaming-first speech synthesis built around the Sonic family of state-space models. Aims at real-time agent voices where latency between turns is the product. Strong choice for sub-200ms voice loops.

    Streaming over WebSocket for fast first audio
    Long-form expressive texture trails ElevenLabs
    • tts
    • streaming
    • low-latency
    • real-time
  • View ElevenLabs details
    VoiceFREEMIUM

    ElevenLabs

    ElevenLabs

    Text-to-speech, voice cloning, and multilingual dubbing.

    Hosted speech synthesis at near-human quality — TTS, voice cloning, multilingual dubbing, and conversational voice agents. Default choice when you need a voice that sounds like a person, not a robot.

    Best-in-class voice realism
    Pricier than commodity TTS at scale
    • tts
    • voice-cloning
    • dubbing
    • multilingual
  • View Rime details
    VoiceFREEMIUM

    Rime

    Rime

    Enterprise text-to-speech built for real-time voice agents.

    Rime builds AI voice models for high-stakes business conversations like IVRs, contact centers, and AI phone agents. Its Arcana and Mist models target ultra-low latency and natural, conversational delivery, with deterministic pronunciation control so terms are spoken consistently without retraining. Rime can be deployed on-prem, in a VPC, or via cloud API, and is offered directly or through voice-AI partner platforms.

    Very low latency for real-time voice agents
    Enterprise-focused, not a consumer tool
    • text-to-speech
    • voice-ai
    • tts
    • contact-center
    • +1
  • View Resemble AI details
    VoiceFREEMIUM

    Resemble AI

    Resemble AI

    Voice cloning, audio watermarking, and deepfake detection in one platform.

    Resemble AI spans both sides of synthetic voice: generating it and policing it. The platform offers voice cloning and text-to-speech built on its Chatterbox models, real-time audio watermarking, and Detect, a multimodal deepfake detector covering audio, image, and video. It deploys in the cloud or fully on-premises for regulated environments.

    Generation + detection in one
    Limited free tier
    • voice-cloning
    • deepfake-detection
    • watermarking
    • tts
    • +1