लाइव
AI models

Masked Autoencoders ViT-H

Training compute
4.6×10²⁰ FLOP
Parameters
632M
Published
Nov 11, 2021

Masked Autoencoders ViT-H is an AI model developed by Facebook AI Research (United States and France), first published in November 2021. It works in the vision domain, on tasks such as semantic segmentation, image classification and image generation.

Training it took an estimated 4.6×10²⁰ FLOP of compute (estimation method: hardware,operation counting). The model has 632,000,000 parameters. It was trained on roughly 328M datapoints.

Access: Open weights (non-commercial). Its weights are openly available. It is built on top of ViT-Huge/14. The reference paper has 11,449 citations. Epoch AI rates the confidence of this record as speculative.

Full record
Organization
Facebook AI Research
Country of organization
United States, France
Domain
Vision
Task
Semantic segmentation, Image classification, Image generation
Training compute
4.6×10²⁰ FLOP
Compute estimation method
Hardware, Operation counting
Parameters
632,000,000
Dataset size
328M
Training time
69 h
Model accessibility
Open weights (non-commercial)
Open weights
Yes
Base model
ViT-Huge/14
Citations
11,449
Epoch confidence
Speculative
More from Facebook AI Research
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models