Microsoft's new speech recognition model undercuts rivals with 10-cent pricing
Microsoft AI released MAI-Transcribe-2, claiming it is faster and more accurate than competing offerings from OpenAI, Google, and ElevenLabs. The model is priced at 10 cents per hour of audio, a 72% reduction from the previous version's rate, significantly lowering costs for enterprises processing large volumes of audio.
Microsoft’s new MAI-Transcribe-2 model enters a crowded field of speech recognition tools, directly challenging offerings from OpenAI, Google, and ElevenLabs. The company positions the system as both faster and more accurate than these competitors, a claim that, if verified, would strengthen its enterprise appeal. The pricing shift is notable: at 10 cents per hour of audio, the cost drops 72% from the prior version. For businesses handling massive volumes of recordings—call centers, media archives, or meeting transcription services—this reduction could meaningfully alter operating budgets. The move reflects a broader industry trend toward aggressive price competition in AI services, where efficiency gains are passed to customers to capture market share.
This pricing strategy could accelerate adoption of AI transcription across small and mid-sized businesses that previously found such tools too costly. Lower barriers may also increase reliance on automated systems for sensitive audio, raising questions about data privacy and accuracy in critical contexts like legal or medical records. While efficiency gains are likely, organizations may need to weigh cost savings against potential errors or compliance risks.