Microsoft announced two new models today called MAI Voice 1 and MAI 1 preview. The company says these models are the first in-house foundation models from Microsoft AI. Microsoft is placing them into Copilot features and into experimental tools for users to try.
MAI Voice 1 is a speech generation model. Microsoft says it can create a minute of audio in under one second on a single GPU. The company already uses the model in Copilot Daily and in generated podcast-style segments that explain topics. Users can test the voice model in Copilot Labs, where they change voice and style.

MAI 1 preview is a text model designed for instruction following. Microsoft says it pre-trained and post-trained the model using about fifteen thousand Nvidia H100 GPUs. The company calls MAI 1 preview a glimpse of future Copilot offerings and says it will test the model on community benchmarks such as LMArena.
Performance Claims
Microsoft highlights speed and efficiency for MAI Voice 1. The company presented the single-GPU audio figure as evidence that the model can serve real-time and interactive uses. Engineers will test latency, heat, and cost across real product flows as the model moves from lab to large scale. Independent tests and third-party benchmarks will be important to confirm the claimed performance.
For the MAI 1 preview, Microsoft stressed that the model focuses on consumer use. The company said it trained the model on a large scale so it can follow instructions and answer everyday queries inside Copilot. Microsoft also described plans to orchestrate specialized models for different tasks rather than using a single giant model for everything. That approach mirrors trends in the wider field.
What To Watch
The new models Microsoft integrates with OpenAI models within Copilot. Microsoft now has OpenAI models in its Copilot stack. The new MAI models will be positioned next to the model so Microsoft can test what will best suit consumers. The balancing of the two sources of models will determine the customer experience and platform economics.
Also, watch model quality and safety. Microsoft will need to show that the MAI 1 preview can match or exceed other instruction following models on factuality and helpfulness. The company must also manage risks tied to audio generation, such as voice misuse and copyright questions. The initial preview and the LMArena tests will offer early signals on where the models sit on quality metrics.

A human view on impact
These releases show Microsoft moving to develop core AI capabilities internally. That does not end the company’s partnership with other labs, but it gives Microsoft more control over product timing and cost. For users, this could mean faster rollout of new voice-driven features inside Copilot and in Windows. For developers and businesses, it could mean new options for voice and assistant APIs from Microsoft Azure over time.