Multimodal AI in .NET: Vision, Audio and Speech
Build multimodal .NET apps that understand images, documents, speech and audio, and generate images and voice, with Microsoft.Extensions.AI and Azure AI.
12 articles about Multimodal AI: in-depth .NET and AI guides, senior interview questions and AI news on DotNet AI Hub.
Build multimodal .NET apps that understand images, documents, speech and audio, and generate images and voice, with Microsoft.Extensions.AI and Azure AI.
Moonshot AI released Kimi K3's weights in July 2026, a 2.8T-parameter multimodal MoE model with a 1M-token context and new terms for large commercial users.
At Build 2026 on June 2, Microsoft launched seven in-house MAI models, led by the MAI-Thinking-1 reasoning model and the MAI-Code-1-Flash coding model.
Google opened its Gemini 3.5 series at I/O on May 19, 2026, with Gemini 3.5 Flash, which it says beats Gemini 3.1 Pro on agent benchmarks, at $1.50/$9 pricing.
Google DeepMind released Gemma 4 in April 2026, multimodal open models from E2B to 31B, and moved Gemma to the Apache 2.0 license for the first time.
Mistral AI released Mistral Small 4 in March 2026, an Apache 2.0 open-weight 119B MoE model with 6.5B active parameters that unifies chat, reasoning and coding.
Google released Gemini 3 on November 18, 2025, starting with Gemini 3 Pro in preview, which topped LMArena at 1501 Elo and launched with Antigravity.
Google DeepMind announced Genie 3 in August 2025, a world model that turns text prompts into interactive 720p worlds at 24 fps that stay consistent for minutes.
Google made Gemini 2.5 Pro and 2.5 Flash generally available on June 17, 2025, previewed 2.5 Flash-Lite and gave developers stable thinking models.
OpenAI agreed to buy io, the AI device startup founded by Jony Ive, for about $6.5 billion in stock. What its hardware push could mean for developers.
OpenAI's o3 and o4-mini, released April 16, 2025, reason with tools and images, posted strong coding and math results and reached developers through the API.
Meta released Llama 4 Scout and Maverick in April 2025, its first natively multimodal mixture-of-experts open-weight models, with a 10M-token context claim.