Video-MME is an AI evaluation tracked by Epoch AI, with 50 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.
Scored models span August 2023 to June 2025. The best score recorded is 0.8, measured as Overall (no subtitles). Models come from ByteDance, Google DeepMind, Alibaba, Shanghai AI Lab,SenseTime,Tsinghua University,Nanjing University,Fudan University,Chinese University of Hong Kong (CUHK),Shanghai Jiao Tong University, OpenAI, NVIDIA,Massachusetts Institute of Technology (MIT),University of California (UC) Berkeley,University of California San Diego,University of Washington,Tsinghua University, Shanghai AI Lab, Anthropic.
| 01 | video-SALMONN-2plus | ByteDance | 0.8 |
| 02 | gemini-1.5-pro-001 | Google DeepMind | 0.75 |
| 03 | gemini-1.5-pro-001-feb24 | Google DeepMind | 0.75 |
| 04 | Qwen2.5-VL-72B-Instruct | Alibaba | 0.73 |
| 05 | InternVL2_5-78B | Shanghai AI Lab,SenseTime,Tsinghua University,Nanjing University,Fudan University,Chinese University of Hong Kong (CUHK),Shanghai Jiao Tong University | 0.72 |
| 06 | gpt-4o-2024-11-20 | OpenAI | 0.72 |
| 07 | gpt-4o-2024-08-06 | OpenAI | 0.72 |
| 08 | Qwen2-VL-72B-Instruct | 0.71 | |
| 09 | LLaVA-Video-72B-Qwen2 | 0.71 | |
| 10 | gemini-1.5-flash-001 | Google DeepMind | 0.7 |
| 11 | LinVT | 0.7 | |
| 12 | Aria | 0.68 | |
| 13 | ViLAMP-llava-qwen | 0.68 | |
| 14 | Oryx-1.5-32B | 0.67 | |
| 15 | LLaVA-OneVision 72B | 0.66 | |
| 16 | VideoLLAMA3-7B | 0.66 | |
| 17 | LLaVA-Video-7B-Qwen2 | 0.66 | |
| 18 | LLaVA-Video-7B-Qwen2-TPO | 0.66 | |
| 19 | VideoChat-Flash-Qwen2-7B_res448 | 0.65 | |
| 20 | gpt-4o-mini-2024-07-18 | OpenAI | 0.65 |
| 21 | ByteVideoLLM-14B | 0.65 | |
| 22 | NVILA-8B | NVIDIA,Massachusetts Institute of Technology (MIT),University of California (UC) Berkeley,University of California San Diego,University of Washington,Tsinghua University | 0.64 |
| 23 | LiveCC-7B-Instruct | 0.64 | |
| 24 | MiniCPM-o-2_6 | 0.64 | |
| 25 | Qwen2-VL-7B-Instruct | 0.64 | |
| 26 | VideoLLAMA2-7B | 0.62 | |
| 27 | InternVL2-40B | Shanghai AI Lab | 0.61 |
| 28 | MiniCPM-V-2_6 | 0.61 | |
| 29 | claude-3-5-sonnet-20241022 | Anthropic | 0.6 |
| 30 | claude-3-5-sonnet-20240620 | Anthropic | 0.6 |
| 31 | gpt-4-1106-vision-preview | OpenAI | 0.6 |
| 32 | mPLUG-Owl3-7B-241101 | 0.59 | |
| 33 | TimeMarker | 0.57 | |
| 34 | VITA-1.5 | 0.56 | |
| 35 | kangaroo | 0.56 | |
| 36 | VITA | 0.56 | |
| 37 | Video-XL-7B | 0.56 | |
| 38 | Video-CCAM-7B-v1.2 | 0.53 | |
| 39 | long-llava-qwen2-7b | 0.53 | |
| 40 | LongVA-7B | 0.53 | |
| 41 | Qwen-VL-Max | 0.51 | |
| 42 | InternVL-Chat-V1-5 | 0.51 | |
| 43 | SliME-Llama3-8B | 0.45 | |
| 44 | Chat-Uni-Vi-7B-v1.5 + 100k SG-WV | 0.43 | |
| 45 | Qwen-VL-Chat | 0.41 | |
| 46 | Chat-UniVi-7B-v1.5 | 0.41 | |
| 47 | sharegpt4video-8b | 0.4 | |
| 48 | Video-LLaVA-7B | 0.4 | |
| 49 | video_chat2_mistral | 0.4 | |
| 50 | ST-LLM | 0.38 |