Live
Benchmarks

Video-MME

Models scored
50
Best score
0.8
Metric
Overall (no subtitles)

Video-MME is an AI evaluation tracked by Epoch AI, with 50 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span August 2023 to June 2025. The best score recorded is 0.8, measured as Overall (no subtitles). Models come from ByteDance, Google DeepMind, Alibaba, Shanghai AI Lab,SenseTime,Tsinghua University,Nanjing University,Fudan University,Chinese University of Hong Kong (CUHK),Shanghai Jiao Tong University, OpenAI, NVIDIA,Massachusetts Institute of Technology (MIT),University of California (UC) Berkeley,University of California San Diego,University of Washington,Tsinghua University, Shanghai AI Lab, Anthropic.

Full record
Score metric
Overall (no subtitles)
Models scored
50
Best score recorded
0.8
Earliest model scored
Aug 20, 2023
Most recent model scored
Jun 18, 2025
Third-party reported
Yes
Leaderboard
01video-SALMONN-2plusByteDance0.8
02gemini-1.5-pro-001Google DeepMind0.75
03gemini-1.5-pro-001-feb24Google DeepMind0.75
04Qwen2.5-VL-72B-InstructAlibaba0.73
05InternVL2_5-78BShanghai AI Lab,SenseTime,Tsinghua University,Nanjing University,Fudan University,Chinese University of Hong Kong (CUHK),Shanghai Jiao Tong University0.72
06gpt-4o-2024-11-20OpenAI0.72
07gpt-4o-2024-08-06OpenAI0.72
08Qwen2-VL-72B-Instruct0.71
09LLaVA-Video-72B-Qwen20.71
10gemini-1.5-flash-001Google DeepMind0.7
11LinVT0.7
12Aria0.68
13ViLAMP-llava-qwen0.68
14Oryx-1.5-32B0.67
15LLaVA-OneVision 72B0.66
16VideoLLAMA3-7B0.66
17LLaVA-Video-7B-Qwen20.66
18LLaVA-Video-7B-Qwen2-TPO0.66
19VideoChat-Flash-Qwen2-7B_res4480.65
20gpt-4o-mini-2024-07-18OpenAI0.65
21ByteVideoLLM-14B0.65
22NVILA-8BNVIDIA,Massachusetts Institute of Technology (MIT),University of California (UC) Berkeley,University of California San Diego,University of Washington,Tsinghua University0.64
23LiveCC-7B-Instruct0.64
24MiniCPM-o-2_60.64
25Qwen2-VL-7B-Instruct0.64
26VideoLLAMA2-7B0.62
27InternVL2-40BShanghai AI Lab0.61
28MiniCPM-V-2_60.61
29claude-3-5-sonnet-20241022Anthropic0.6
30claude-3-5-sonnet-20240620Anthropic0.6
31gpt-4-1106-vision-previewOpenAI0.6
32mPLUG-Owl3-7B-2411010.59
33TimeMarker0.57
34VITA-1.50.56
35kangaroo0.56
36VITA0.56
37Video-XL-7B0.56
38Video-CCAM-7B-v1.20.53
39long-llava-qwen2-7b0.53
40LongVA-7B0.53
41Qwen-VL-Max0.51
42InternVL-Chat-V1-50.51
43SliME-Llama3-8B0.45
44Chat-Uni-Vi-7B-v1.5 + 100k SG-WV0.43
45Qwen-VL-Chat0.41
46Chat-UniVi-7B-v1.50.41
47sharegpt4video-8b0.4
48Video-LLaVA-7B0.4
49video_chat2_mistral0.4
50ST-LLM0.38
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks