Menu

Menu

Menu

[Voice AI Weekly #9] Second Week of September 2026

[Voice AI Weekly #9] Second Week of September 2026

[Voice AI Weekly #9] Second Week of September 2026

Technology

Technology

Voice AI Weekly #9 - 2nd Week of September 2026

This week, the global voice AI industry saw notable developments in three areas: on-device voice recognition, contact center integration, and enterprise voice agents.


1. Meta Unveils Real-Time Voice Recognition Model Muse Voice Transcribe

Meta Superintelligence Labs introduced Muse Voice Transcribe, its first real-time audio recognition model, on September 1st. By re-evaluating the audio buffer every 80 milliseconds to determine the end of speech, it handles multi-speaker separation for distinguishing over 20 speakers, streaming speech recognition, and endpointing with a single model. On-device deployment targeting Ray-Ban Meta and Oakley Meta glasses is key, and it has already been integrated into Meta AI dictation for Mac and the coding assistant Muse Code.

Analysis article on Meta Muse Voice Transcribe (shattered.io)


2. SoundHound Finalizes LivePerson Acquisition, Launching Omnichannel Strategy in Earnest

SoundHound AI announced on September 4th that it has officially completed the acquisition of LivePerson and appointed John Collins as its new CFO. Securing a customer base that includes 25 Fortune 100 companies and over 750 patents, the company plans to integrate LivePerson's digital messaging infrastructure into its agent platform OASYS to operate voice, web, mobile, SMS, and social channels through a single engine.

This move to unify voice and digital channels shows that the contact center market is shifting from channel-specific individual solutions to competition among integrated platforms. Deployment flexibility, encompassing both on-premises and on-device, is expected to become an increasingly crucial variable in this integrated competition.

SoundHound AI Completes Acquisition of LivePerson (SoundHound Newsroom)


3. Google Formally Launches Voice Conversation Features in Gmail, Docs, and Keep

Google officially launched Gmail Live, Docs Live, and Keep Live on September 3rd. First revealed at Google I/O last May, these Gemini audio-based interactive voice features allow users to search and summarize their inbox by voice in Gmail, draft documents hands-free in Docs, and convert rambling speech into organized notes in Keep.

Attempts to complete tasks using only voice within business applications, going beyond mere dictation, are spreading among big tech companies. In domestic corporate environments, demand for voice interfaces tailored to specific tasks, such as summarizing meeting minutes or drafting reports, is also expected to continue.

Google rolling out Gmail Live, Docs Live, and Keep Live (9to5Google)


4. ElevenLabs Officially Launches Agent Conversation Management Feature

ElevenLabs has officially launched Procedures and Conversation Triage Ticket features on ElevenAgents. Procedures are task-specific instructions called up only in specific situations, allowing a single agent to handle multiple tasks without having to pack all instructions into a system prompt. Triage Ticket is a feature for registering agent response quality issues as tickets and assigning them to a representative for follow-up actions.

ElevenLabs Changelog, August 24, 2026


5. Speechmatics Selected as Tresic's Conversation Intelligence Engine, Beating Google and Deepgram

Tresic, which provides conversation intelligence to communication and managed service providers, announced on September 9th that it has selected Speechmatics as the speech recognition engine for its Intelligence Cloud platform. This decision was made after a direct comparative evaluation of Google and Deepgram, with the key factors being speaker separation performance and recognition accuracy that precisely captures sensitive information like card numbers. The Speechmatics engine operates in two stages within Tresic's own infrastructure, ensuring that audio and transcripts do not leave the system.

Because missing even a single digit of a card number means masking fails, Tresic treated transcription accuracy as a compliance control item. This case demonstrates that domain-specific recognition accuracy and the ability to process data within own infrastructure are what actually drive vendor selection.

Tresic selects Speechmatics to power the speech layer of its conversation intelligence platform (GlobeNewswire)


Please feel free to contact us at contact@holamago.com for inquiries or collaboration proposals. We will return next week with more voice AI news curated from MAGO's perspective.