Pix2Struct-Large is an AI model developed by Google Research and University of Cambridge (United States and United Kingdom), first published in June 2023. It works in the vision domain, on tasks such as image captioning and visual question answering.
Training it took an estimated 1.7×10²⁰ FLOP of compute (estimation method: operation counting). The model has 1,300,000,000 parameters.
Access: Open weights (unrestricted). Its weights are openly available. Epoch AI rates the confidence of this record as likely.