DAX LLM Benchmark

Which LLM is best at DAX?
DAXBench tests how models understand, write, and reason about DAX.
Methodology designed by Maxim Anatsko.

Last updated: Jul 10, 2026

127 models · 30 tasks · Initial Release

Model Leaderboard

Ranked by score

ModelTasks
1
Gemini 3.1 Flash Lite PreviewHIGH
Google
97.38%
96.67%100%29/30
2
Claude Fable 5Auto (legacy)
Anthropic
96.92%
96.67%100%29/30
3
GPT-5.3 ChatAuto (legacy)
OpenAI
96.9%
96.67%100%29/30
4
Qwen3.5 Plus 2026-02-15MED
Qwen
96.84%
96.67%100%29/30
5
GLM 5Auto (legacy)
Z.AI
96.23%
96.67%100%29/30
6
Qwen3.7 MaxAuto (legacy)
Qwen
94.53%
93.33%100%28/30
7
Gemini 3.1 Pro PreviewHIGH
Google
94.49%
93.33%100%28/30
8
Gemma 4 31BAuto (legacy)
Google
94.46%
93.33%100%28/30
9
Qwen3.6 Plus Preview (free)Auto (legacy)
Qwen
93.92%
93.33%100%28/30
10
Qwen3.5 397B A17BAuto (legacy)
Qwen
93.89%
93.33%100%28/30
11
GPT-5.4 MiniAuto (legacy)
OpenAI
93.26%
93.33%100%28/30
12
Qwen3.6 Max PreviewAuto (legacy)
Qwen
91.55%
90%100%27/30
13
Qwen3.5-FlashMED
Qwen
90.82%
90%100%27/30
14
GLM 5.1Auto (legacy)
Z.AI
90.28%
90%100%27/30
15
Qwen3.6 Plus (free)Auto (legacy)
Qwen
89.66%
90%100%27/30
16
GLM 5V TurboAuto (legacy)
Z.AI
89.11%
86.67%100%26/30
17
GPT-5.3-CodexHIGH
OpenAI
88.65%
86.67%100%26/30
18
gpt-oss-120bAuto (legacy)
OpenAI
88%
86.67%100%26/30
19
Grok 4.5Auto (legacy)
xAI
87.87%
86.67%100%26/30
20
Claude Sonnet 4.6MED
Anthropic
87.39%
86.67%100%26/30
21
Claude Sonnet 4Auto (legacy)
Anthropic
87.35%
86.67%100%26/30
22
KAT-Coder-Pro V2Auto (legacy)
Kwaipilot
87.23%
86.67%100%26/30
23
GLM 5 TurboAuto (legacy)
Z.AI
86.65%
86.67%100%26/30
24
Gemini 2.5 Flash Preview 09-2025Auto (legacy)
Google
86.17%
83.33%100%25/30
25
GPT-5.1-Codex-MaxAuto (legacy)
OpenAI
85.63%
83.33%100%25/30
26
Claude Opus 4.8Auto (legacy)
Anthropic
85.41%
83.33%100%25/30
27
Gemini 3 Pro PreviewAuto (legacy)
Google
84.9%
83.33%100%25/30
28
Claude Sonnet 4.5Auto (legacy)
Anthropic
84.43%
83.33%100%25/30
29
o3Auto (legacy)
OpenAI
84.4%
83.33%100%25/30
30
Gemini 3.1 Flash LiteAuto (legacy)
Google
84.36%
80%100%24/30
31
Kimi K2 ThinkingAuto (legacy)
Moonshot AI
84.35%
83.33%100%25/30
32
GPT-5.4HIGH
OpenAI
83.82%
80%100%24/30
33
Grok 4.3Auto (legacy)
xAI
83.74%
83.33%100%25/30
34
Grok 4.20 BetaHIGH
xAI
83.04%
83.33%100%25/30
35
Claude Opus 4.5Auto (legacy)
Anthropic
82.71%
80%100%24/30
36
Grok 4Auto (legacy)
xAI
82.7%
80%100%24/30
37
Claude Opus 4.6Auto (legacy)
Anthropic
82.03%
80%100%24/30
38
Gemini 3 Flash PreviewAuto (legacy)
Google
81.38%
76.67%100%23/30
39
GPT-5.2Auto (legacy)
OpenAI
81.36%
80%100%24/30
40
R1Auto (legacy)
DeepSeek
81.33%
80%100%24/30
41
Grok Build 0.1Auto (legacy)
xAI
81.32%
80%100%24/30
42
Kimi K2.7 CodeAuto (legacy)
Moonshot AI
81.1%
80%96.67%24/30
43
DeepSeek V4 ProAuto (legacy)
DeepSeek
80.79%
80%93.33%24/30
44
Aurora AlphaAuto (legacy)
Openrouter
80.56%
80%100%24/30
45
GPT-5.2 ChatAuto (legacy)
OpenAI
80.1%
80%100%24/30
46
Qwen3 Max ThinkingAuto (legacy)
Qwen
80.09%
80%100%24/30
47
Gemini 2.5 FlashAuto (legacy)
Google
79.58%
76.67%100%23/30
48
Kimi K2.6Auto (legacy)
Moonshot AI
78.71%
76.67%96.67%23/30
49
DeepSeek V3.2 SpecialeAuto (legacy)
DeepSeek
78.31%
76.67%100%23/30
50
DeepSeek V3.1Auto (legacy)
DeepSeek
78.28%
76.67%100%23/30

About This Benchmark

Evaluation Method

Models are tested against DAX tasks of varying complexity using the Contoso sample dataset. Responses are evaluated for syntax correctness and output accuracy.

Scoring System

Harder tasks are worth more points. Correct solutions also earn bonus points for following DAX best practices, writing efficient code, and producing clear, readable output.

Task Categories

Tasks cover aggregation, time intelligence, filtering, calculation, table manipulation, iterator, context transition across basic, intermediate, and advanced levels.

Browse by Category

Browse All Tasks

Evaluating language model capabilities in DAX code generation

Created by Maxim Anatsko, BI developer specializing in Power BI automation and LLM workflows.