Menu

Menu

Menu

[Voice AI Weekly #7] Fourth Week of August 2026

[Voice AI Weekly #7] Fourth Week of August 2026

[Voice AI Weekly #7] Fourth Week of August 2026

Technology

Technology

Voice AI Weekly #7 - Fourth Week of August 2026

In the global voice AI industry this week, a prominent trend was voice input establishing itself as a default feature of operating systems and collaboration platforms rather than being confined to specific apps. The trend of extracting emotion and speaker information directly from the signal layer also continues.


1. Meta AI Launches Mac App Supporting System-Wide Voice Dictation

Meta newly introduced the Meta AI app for Mac on August 20th. It features a system-wide dictation function that allows users to write documents directly using voice in any app running on Mac, without being tied to a specific program. Utilizing its proprietary model, Muse Spark, it also includes a feature that understands what is currently on the screen and responds accordingly. Google also added a similar system-wide dictation feature to its Gemini Mac app last month.

It is noteworthy that voice input, which used to work only within specific apps, is becoming a basic feature at the operating system level. As the habit of working with voice across the desktop becomes more established, the utility of signal processing technologies that read emotion and intent along with it is expected to expand!

Source: Meta AI's new Mac app wants you to talk to your apps


2. AudioCodes Voca CIC Certified for Microsoft Teams Voice Agent Unify Integration

On August 17th, AudioCodes' contact center solution Voca CIC became the first to be certified under Microsoft Teams' new Voice Agent Unify integration program. It passed both third-party functional validation and Microsoft's security review, handling caller verification and intent determination before routing calls to Teams Phone auto-attendants or agents. Adopting companies use it as a managed service operated by AudioCodes without needing separate deployment.

Microsoft has officially introduced external voice agents to Teams through its own certification process. Enterprise voice AI is increasingly being incorporated as a standard component of major collaboration platforms. For vendors supplying features based on APIs, this certified integration structure could serve as a reference model.

Source: AudioCodes Voca CIC Now Certified for Microsoft Teams Unify Integration for Voice Agents


3. HeyBreez Secures $2.5 Million Seed Funding for Enterprise Voice AI Infrastructure

Delaware-based startup HeyBreez announced on August 20th that it raised $2.5 million in a seed round, bringing its cumulative funding to $3.8 million. Processing over 1 million multilingual calls per month, the company manages the end-to-end operational processes before and after calls, including carrier routing, retries, callbacks, SMS follow-ups, and CRM integration.

It is notable that investment is pouring into startups that bundle and sell operational solutions for post-call tasks, rather than just converting speech to text. The focus of competition in contact center automation is now on operational stability before and after calls, rather than model performance alone.

Source: HeyBreez Raises $2.5M in Seed funding


4. Smallest.ai Unveils 'Pulse STT Pro', an STT Bundling Diarization and Emotion Recognition

Indian startup Smallest.ai announced $21 million in cumulative funding, along with the release of its new voice platform 'Voice 4.0' and STT model 'Pulse STT Pro'. Supporting 38 languages, this model provides low-latency transcription bundled with speaker diarization, emotion recognition, code-switching recognition, and automatic PII masking as standard features.

Source: Smallest.ai gets $21 Million in Funding to Build Voice 4.0, the Next Generation of Enterprise Voice AI


5. NAVER Cloud Presents 'Sommelier', a FullDuplex Voice Preprocessing Technology, at an International Conference

NAVER Cloud presented its voice preprocessing technology 'Sommelier' at ACL 2026, an international conference in the field of natural language processing. By adopting a FullDuplex approach where a single model processes speech, video, and text simultaneously, the AI can listen to the user's voice in real-time even while generating a response, reflecting satisfaction or follow-up questions in the conversation.

The design of processing the voice signal itself prior to the text conversion stage shares the same direction as MAGO's Acoustic Intelligence. The fact that a domestic company brought research in this area to an international conference stage is also a useful reference for assessing the technological competitiveness of Korean voice AI.

Source: Voice AI Turns Human: Machines That Interrupt, Feel and Reply


As voice establishes itself as a major feature of platforms beyond specific apps, the competition in technology to read emotion and speaker information along with it is heating up.

For inquiries or collaboration proposals, please feel free to contact us at contact@holamago.com. We will return next week with more voice AI news curated from MAGO's perspective.