Live
Benchmarks

Adversarial NLI

Models scored
15
Best score
0.58
Metric
Score

Adversarial NLI is an AI evaluation tracked by Epoch AI, with 15 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span January 2022 to April 2024. The best score recorded is 0.58, measured as Score. Models come from Microsoft, Mistral AI, Google DeepMind, Meta AI, OpenAI, Microsoft,NVIDIA.

Full record
Score metric
Score
Models scored
15
Best score recorded
0.58
Earliest model scored
Jan 27, 2022
Most recent model scored
Apr 23, 2024
Third-party reported
Yes
Leaderboard
01gpt-3.5-turbo-1106OpenAI0.58
02Phi-3-small-8k-instructMicrosoft0.58
03Meta-Llama-3-8B-InstructMeta AI0.57
04Phi-3-medium-128k-instructMicrosoft0.56
05Mixtral-8x7B-v0.1Mistral AI0.55
06Phi-3-mini-4k-instructMicrosoft0.53
07gemma-7bGoogle DeepMind0.49
08Mistral-7B-v0.1Mistral AI0.47
09phi-2Microsoft0.42
10Megatron-Turing NLG 530BMicrosoft,NVIDIA0.4
11text-davinci-001OpenAI0.35
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks