Boson AI’s speech-to-speech voice model bypasses text to rival OpenAI

Something quiet is happening in voice AI that deserves more attention. Boson AI, a Santa Clara startup founded in 2023, is pushing beyond the standard text-to-speech pipeline with a new speech-to-speech voice model called Higgs RealTime — and the technical shift it represents is more significant than it might first appear.
Key takeaways
- Boson AI’s Higgs RealTime eliminates intermediate text conversion in voice processing, reducing latency and preserving vocal nuances in real time.
- Higgs TTS 3, released June 4, 2026, supports expressive speech in over 100 languages with zero-shot voice cloning and inline emotion control.
- The Higgs Avatar API, launched in June 2026, generates real-time talking-head video from a single still image plus audio or text input.
- Earlier Higgs TTS 2 models were open-sourced on Hugging Face in May 2025 after training on over 10 million hours of audio data.
- Boson AI competes in a market alongside OpenAI, Google, and ElevenLabs, differentiating through low-latency, production-grade focus and an open-source developer strategy.
Advancing Beyond Text-to-Speech With a Speech-to-Speech Model
Most voice AI systems today follow the same basic architecture: convert incoming speech to text, process the text, then synthesize a spoken response. It works — but every step in that chain adds delay and strips out the natural texture of human speech. Tone, pacing, hesitation, emotional coloring: all of it gets flattened by the time audio is reconstructed on the other end.
… Continue reading the full article at the original source below.


