Microsoft unveils three AI models for speech, images
Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Microsoft unveiled public preview versions of three new AI models — MAI-Transcribe-1, MAI-Voice-1, and MAI-Image-2 — for speech recognition, speech synthesis, and text-to-image generation.
- MAI-Transcribe-1 offers enterprise‑grade accuracy across 25 languages while using roughly 50 % less GPU compute than leading alternatives.
- MAI-Voice-1 can synthesize 60 seconds of audio in under a second on a single GPU, enabling ultra‑fast speech generation.
- Azure AI Foundry now hosts these models, letting developers integrate them into applications and services.
- Microsoft already powers its Copilot, Bing, PowerPoint and Azure Speech products with the same models, and Copilot’s Audio Expressions and Voice Mode use MAI‑Voice‑1 and MAI‑Transcribe‑1 respectively.
- Microsoft holds an OpenAI stake valued at about $135 billion, highlighting the strategic tension between partnership and competition.
Why it matters: Enterprises and developers gain cheaper, faster AI capabilities for speech and image tasks, while Microsoft expands its AI portfolio and leverages its $135 billion OpenAI stake to balance partnership and competition. OpenAI faces heightened competition as Microsoft offers comparable services under its own brand in the enterprise market.
Ask SkimNews



