| Rank | Model configuration | |||
|---|---|---|---|---|
| 1 | Gemini 3.7 FlashGoogle | 2.783 | 0.6407 | 0.1768 |
| 2 | Grok 4.6xAI | 2.582 | 0.6095 | 0.1592 |
| 3 | Qwen3.8-MaxAlibaba | 2.565 | 0.6031 | 0.1504 |
| 4 | GPT-5.6 SolOpenAI | 2.533 | 0.6115 | 0.1487 |
| 5 | Claude Opus 5Anthropic | 2.527 | 0.6008 | 0.1402 |
| 6 | Gemma 4 31B ITGoogle | 2.516 | 0.6356 | 0.1810 |
| 7 | DeepSeek-V4-FlashDeepSeek | 2.413 | 0.6206 | 0.1653 |
| 8 | DeepSeek-V4-ProDeepSeek | 2.386 | 0.6203 | 0.1692 |
| 9 | Ministral 3 14B Instruct 2512Mistral AI | 2.201 | 0.6230 | 0.1636 |
| 10 | MedGemma 27B ITGoogle | 2.152 | 0.6266 | 0.1738 |
| 11 | Llama 4 Scout 17B-16E InstructMeta | 2.098 | 0.6211 | 0.1691 |
| 12 | Meditron3 70BEPFL LiGHT / OpenMeditron | 1.929 | 0.6103 | 0.1631 |
| 13 | HuatuoGPT-3-32BFreedomIntelligence | 1.826 | 0.5593 | 0.0906 |
| 14 | MedReason-8BUCSC-VLAA | 1.435 | 0.5921 | 0.1133 |
184Board casesBoard Simulation
4,844Specialist questionsSpecialist Turn
14Model configsAcross both tasks
2.783Top Conclusion alignmentGemini 3.7 FlashBoard Simulation
3.427Top Clinical equivalenceDeepSeek-V4-ProSpecialist Turn
Conclusion alignmentHigher is better
14 models
1
Gemini 3.7 Flash Best
2
Grok 4.6 Best
3
Qwen3.8-Max Best
4
GPT-5.6 Sol Best
5
Claude Opus 5 Best
6
Gemma 4 31B IT Best
7
DeepSeek-V4-Flash Best
8
DeepSeek-V4-Pro Best
9
Ministral 3 14B Instruct 2512 Best
10
MedGemma 27B IT Best
11
Llama 4 Scout 17B-16E Instruct Best
12
Meditron3 70B Best
13
HuatuoGPT-3-32B Best
14
MedReason-8B Best
No published models in this viewCompleted and audited configurations will appear here automatically.
EfficiencyQuality against generation cost
Mean output tokens per response, measured over the same 184 cases.
Complete metric dataAll values remain visible. Select a metric column to rank the chart.
Scroll horizontally to compare all metrics Clinical equivalenceHigher is better
9 models
1
DeepSeek-V4-Pro Best
2
DeepSeek-V4-Flash Best
3
Gemma 4 31B IT Best
4
Ministral 3 14B Instruct 2512 Best
5
Meditron3 70B Best
6
Llama 4 Scout 17B-16E Instruct Best
7
MedGemma 27B IT Best
8
HuatuoGPT-3-32B Best
9
MedReason-8B Best
No published models in this viewCompleted and audited configurations will appear here automatically.
EfficiencyQuality against generation cost
Mean output tokens per response, measured over the same 4,844 questions.
Across the boardProfile by target specialist
This configuration
Panel median
- surgeon1,766 q
- medical oncologist1,347 q
- radiation oncologist569 q
- radiologist407 q
- other307 q
- molecular pathologist207 q
- pathologist192 q
DeepSeek-V4-Pro
3.43 overall
molecular pathologist pathologist
DeepSeek-V4-Flash
3.33 overall
molecular pathologist radiation oncologist
Gemma 4 31B IT
3.08 overall
molecular pathologist other
Ministral 3 14B Instruct 2512
3.02 overall
other surgeon
Meditron3 70B
3.01 overall
pathologist radiation oncologist
Llama 4 Scout 17B-16E Instruct
2.97 overall
molecular pathologist medical oncologist
MedGemma 27B IT
2.88 overall
other pathologist
HuatuoGPT-3-32B
2.67 overall
molecular pathologist pathologist
MedReason-8B
2.09 overall
radiologist pathologist
Complete metric dataAll values remain visible. Select a metric column to rank the chart.
Scroll horizontally to compare all metrics | Rank | Model configuration | |||||
|---|---|---|---|---|---|---|
| 1 | DeepSeek-V4-ProDeepSeek | 3.427 | 5.6% | 11.8% | 0.6155 | 0.1829 |
| 2 | DeepSeek-V4-FlashDeepSeek | 3.328 | 6.7% | 13.7% | 0.6168 | 0.1811 |
| 3 | Gemma 4 31B ITGoogle | 3.080 | 6.5% | 10.0% | 0.6329 | 0.1983 |
| 4 | Ministral 3 14B Instruct 2512Mistral AI | 3.015 | 12.1% | 28.7% | 0.6077 | 0.1688 |
| 5 | Meditron3 70BEPFL LiGHT / OpenMeditron | 3.014 | 11.2% | 20.9% | 0.6430 | 0.2171 |
| 6 | Llama 4 Scout 17B-16E InstructMeta | 2.967 | 10.2% | 15.0% | 0.6423 | 0.2094 |
| 7 | MedGemma 27B ITGoogle | 2.881 | 13.7% | 23.1% | 0.6216 | 0.1841 |
| 8 | HuatuoGPT-3-32BFreedomIntelligence | 2.671 | 23.1% | 74.5% | 0.5587 | 0.1092 |
| 9 | MedReason-8BUCSC-VLAA | 2.093 | 43.5% | 52.0% | 0.5988 | 0.1463 |