ModelBeat
Models/Grok 4
All models
X
grok-4

Grok 4 is xAI’s proprietary AI model, released Jul 9, 2025.

Context
Input
/ 1M
Output
/ 1M
Released
Jul 9, 2025
Modalities

Benchmarks

Scores on standardized evaluations. Higher is better — and rank shows where Grok 4 lands among all models tracked on Model Beat.

32.0
Intelligence Index
Epoch AI
32nd percentile of tracked models
43.0
Coding Index
Epoch AI
43rd percentile of tracked models
15.0
Agentic Index
Epoch AI
15th percentile of tracked models

Reasoning

8 evals
ARC-AGI66.7%

Abstract visual reasoning

ARC-AGI-216.0%

Harder abstract reasoning

GPQA Diamond87.0%

Graduate-level scientific reasoning

Humanity's Last Exam26.7%

Frontier of human expert knowledge

MMLU-Pro86.6%

Broad expert knowledge

SimpleBench60.5%

Common-sense trick questions

SimpleQA Verified47.9%

Factual accuracy & hallucination

WeirdML45.7%

Novel ML problem-solving

Coding

2 evals
LiveCodeBench81.9%

Contamination-free coding

SciCode45.7%

Scientific research coding

Math

3 evals
AIME 2024/202584.0%

Olympiad-qualifier math

FrontierMath19.7%

Research-level math problems

FrontierMath Tier 42.1%

Hardest research math

Agentic & Tools

5 evals
APEX15.2%

Multi-step agentic tasks

GDPval (win/tie rate)24.3%

Economically valuable work

METR task horizon1.8 h

Autonomous task length

Terminal-Bench27.2%

Command-line agentic tasks

τ²-bench74.9%

Tool-agent-user reliability

Changelog

5
  • Grok 4: Humanity's Last Exam improved from 23.9% to 26.7% (+12%)benchmark · Aug 6, 2026
  • Grok 4: ARC-AGI-2 dropped from 29.4% to 16.0% (-46%)benchmark · Aug 5, 2026
  • Grok 4: ARC-AGI dropped from 79.6% to 66.7% (-16%)benchmark · Aug 5, 2026
  • Grok 4: ARC-AGI-2 improved from 16.0% to 29.4% (+84%)benchmark · Aug 5, 2026
  • Grok 4: ARC-AGI improved from 66.7% to 79.6% (+19%)benchmark · Aug 5, 2026

Compare Grok 4 with

Frequently asked questions

What is Grok 4?

Grok 4 is a Grok-family AI model, from xAI, released Jul 9, 2025.

Who created Grok 4?

Grok 4 is an AI model developed by xAI, part of the Grok family, released Jul 9, 2025.

How does Grok 4 perform on benchmarks?

On standardized evaluations tracked by Epoch AI, Grok 4 scores 86.6% on MMLU-Pro, 81.9% on LiveCodeBench, 47.9% on SimpleQA Verified.

Is Grok 4 open source?

Grok 4 is a proprietary model (API access), available through its provider's API rather than as a downloadable open-weight model.

Benchmark data from Epoch AI (CC BY) and Artificial Analysis; pricing, specs & descriptions from OpenRouter. Data updated Jul 8, 2026.