Live
AI models

Phi-4-Multimodal

Training compute
1.2×10²³ FLOP
Parameters
5.6B
Published
Mar 3, 2025

Phi-4-Multimodal is an AI model developed by Microsoft (United States), first published in March 2025. It works in the multimodal, language, vision and speech domain, on tasks such as language modeling/generation, question answering, visual question answering and 4 more.

Training it took an estimated 1.2×10²³ FLOP of compute (estimation method: operation counting,hardware). The model has 5,600,000,000 parameters. Training ran on 512 NVIDIA A100 SXM4 80 GB for about 672 hours.

Access: Open weights (unrestricted). Its weights are openly available. It is built on top of Phi-4 Mini. Epoch AI rates the confidence of this record as confident.

Full record
Organization
Microsoft
Country of organization
United States
Domain
Multimodal, Language, Vision, Speech
Task
Language modeling/generation, Question answering, Visual question answering, Speech recognition (ASR), Translation, Audio question answering, Character recognition (OCR)
Training compute
1.2×10²³ FLOP
Compute estimation method
Operation counting, Hardware
Parameters
5,600,000,000
Training hardware
NVIDIA A100 SXM4 80 GB
Chips used
512
Training time
672 h
Training power draw
402.0 kW
Model accessibility
Open weights (unrestricted)
Open weights
Yes
Base model
Phi-4 Mini
Epoch confidence
Confident
More from Microsoft
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models