More on the topic...
Generating detailed summary...
Failed to generate summary. Please try again.
Microsoft just rolled out MAI-Image-2.5, a next-gen image model that now ranks No. 2 on Arena’s Image Edit leaderboard—beating Nano Banana 2.1—and No. 3 for text-to-image tasks. There are two flavors: the full-fidelity MAI-Image-2.5 for top-quality outputs and MAI-Image-2.5-Flash for speedy, cost-effective runs. Both excel at detailed text rendering, coherent scene composition, and precise local edits—say removing motion blur or swapping out objects—while preserving face identity across changes in pose or expression.
Under the hood, the model grasps lighting, scale and perspective, so added elements look natural. Benchmarks show it outperforms GPT-Image-1.5 and Nano Banana Pro 2K on prompt adherence, visual quality and controlled editing. Microsoft has already embedded it in PowerPoint for auto-generated slide visuals and in OneDrive for fine-tuned photo cleanup. Developers can access it today via Foundry at $5 per 1 million text-input tokens, $8 per 1 million image-input tokens and $47 per 1 million image-output tokens. Flash drops those costs to $1.75, $1.75 and $19.50 respectively.
Safety measures include layered guardrails that filter harmful or policy-violating content. Still, the model can mirror biases in its training data and occasionally invent misleading visual details, so outputs need review in sensitive domains—identity, legal, medical, financial or news. You can test both versions in the MAI Playground or tap into them through OpenRouter’s API.
Questions about this article
No questions yet.