Masked Autoencoders ViT-H is an AI model developed by Facebook AI Research (United States and France), first published in November 2021. It works in the vision domain, on tasks such as semantic segmentation, image classification and image generation.
Training it took an estimated 4.6×10²⁰ FLOP of compute (estimation method: hardware,operation counting). The model has 632,000,000 parameters. It was trained on roughly 328M datapoints.
Access: Open weights (non-commercial). Its weights are openly available. It is built on top of ViT-Huge/14. The reference paper has 11,449 citations. Epoch AI rates the confidence of this record as speculative.