Menu

Menu

Menu

[Voice AI Weekly #6] Third Week of August 2026

[Voice AI Weekly #6] Third Week of August 2026

[Voice AI Weekly #6] Third Week of August 2026

Technology

Technology

Voice AI Weekly #6 - Third Week of August 2026

Over the past week in the overseas voice AI industry, the key topics were how naturally to handle situations where someone cuts into the middle of a conversation, and how deeply to integrate voice technology into actual business systems. We have summarized five related news items.


1. Deepgram Switches Conversational TTS 'Flux' to Default Agent Voice

On August 12, Deepgram officially released Flux TTS, a voice synthesis model specialized for conversation. It has been changed to the default voice engine in the Voice Agent API without any additional configuration. New features have been added, such as precisely indicating how much of the consultation was delivered and where it should resume from when a user interrupts while speaking, and immediately changing the speaking speed without needing to reconnect.

In contact centers, it is common for customers to interrupt during a consultation. How smoothly that moment is handled determines the actual call quality satisfaction. Therefore, more and more TTS engines that reflect the context of the conversation as it is are appearing.

Source Flux TTS is generally available, and is now the default agent voice


2. Cartesia Takes 1st Place Simultaneously in Both Voice and Transcription Categories with Sonic-3.6

On August 18, Cartesia unveiled Sonic-3.6, a new voice synthesis model. Along with the transcription model Ink-2 released last June, they announced that they took 1st place in both the Speech and Transcription Arenas of Artificial Analysis. The TTS latency is below 90 milliseconds, the transcription latency is around 100 milliseconds, and it also features a built-in turn detection function.

The competition to reduce latency by combining voice synthesis and speech recognition into a single stack is narrowing down to the millisecond level.

Source Cartesia | Introducing Sonic-3.6 and Ink-2


3. PolyAI Direct-Integrates with Hospital System Epic to Automate Booking Calls

On August 18, PolyAI announced a direct integration with Epic, an electronic health record system. PDS Health, which operates more than 1,100 dental and medical practices, was the first to adopt it. Since it directly connects to each hospital's Epic environment without any intermediate middleware, latency is reduced, and root causes can be identified much faster even when failures occur.

Healthcare contact centers have many calls that are repetitive yet must be accurate, such as appointment changes or prescription verifications. Direct integration with core business systems like Epic is increasingly becoming the standard in this domain.

Source PolyAI launches direct integration with Epic, with PDS Health among the first to deploy at scale


4. NVIDIA Releases Open-Source Simultaneous Two-Way Voice Model 'NemotronLabs VoiceChat'

On August 9, NVIDIA released 'NemotronLabs VoiceChat 11B', an open-source full-duplex (simultaneous two-way) voice model. Instead of separately linking Speech-to-Text (STT), Language Models (LLM), and Text-to-Speech (TTS) as before, a single model handles listening and speaking at the same time.

The turn-taking latency is around 448ms, and the response speed when a user interrupts and cuts in is also quite fast, within 480ms. It also includes a feature that naturally fills the silence caused by calling external tools mid-conversation with pre-determined prompt announcements.

However, only the weights have been released and restricted to 'research purposes', it requires an 80GB-class GPU, and there is no commercial API yet. Although it is difficult to apply to services immediately, it can be a great reference for teams building voice agents with their own infrastructure in that it demonstrated full-duplex processing and live tool calling with a single open-source model.

Source NVIDIA Releases NemotronLabs VoiceChat 11B: An Open Full-Duplex Speech-to-Speech Model with ~450 ms Turn-Taking and Live Tool Calling


5. Sarvam AI Raises Additional $75 Million Led by NVIDIA

Indian voice AI company Sarvam AI announced on August 4 that it secured $75 million in a Series B extension round led by NVIDIA. This is the first time NVIDIA has participated as a strategic investor in an Indian AI firm. Sarvam has previously collected data from 17 million farmers for the Indian Ministry of Agriculture using multilingual voice agents, and has also conducted policy renewal announcements for an insurance company on a scale of 45 million people.

To apply voice AI to language regions with relatively few resources, fine-tuning tailored to that language and domain is ultimately required. The trend of large investments pouring into multilingual, low-resource language voice AI seems to be a signal that this demand is continuously growing.

Source Sarvam AI Onboards Nvidia As Strategic Investor In $75 Million Series B Extension, First Indian Firm To Do So


Whether open-source or commercial, the trend is similar. The competition seems to be head-to-head on how much they can reduce the response speed and how well they can integrate with existing systems.

If you have any questions or want to know more about MAGO's emotion recognition and on-premise solutions, please feel free to contact us at contact@holamago.com. We will be back with new updates next week.