LLM Evaluation Leaderboard

SpaceOmicsBench β€” 100 questions across 9 spaceflight omics modalities & 4 difficulty tiers

Judge: Claude Sonnet 4.6 100 Questions 1–5 Scale 9 Models 5 Dimensions
πŸ† Overall Leaderboard
Weighted score across five dimensions (Factual 25%, Reasoning 25%, Completeness 20%, Uncertainty 15%, Domain 15%). πŸ”’ proprietary API Β· πŸ”“ open-weight
🌟

Claude Sonnet 4.6 leads at 4.62

Top across all difficulty tiers, with Reasoning at 4.97 and Completeness at 4.77. The judge assigned 93 novel-insight flags.

πŸ”“

DeepSeek-V3: strongest open-weight model in this evaluation

4.34/5.00 β€” surpasses Claude Sonnet 4, GPT-4o, and Gemini 2.5 Flash. Especially strong on Hard & Expert questions.

⚠️

Gemini: thinking mode matters

With max_tokens=8192, Gemini 2.5 Flash scores 4.00 β€” uniform across difficulties (Easy 3.52 β†’ Expert 4.04). Previous 2.74 Expert score was a truncation artifact.

β‰ˆ

Bottom tier clusters at ~3.30

GPT-4o Mini (3.32) and GPT-4o (3.30) are nearly tied in this evaluation; the two Llama-70B backends score 3.31.

πŸ“ˆ Difficulty Profile
Compares how model scores change from Easy to Expert questions.
πŸ•ΈοΈ Five-Dimensional Breakdown
Select up to three models to compare across the five scoring dimensions. Uncertainty Calibration has the lowest mean score across the nine evaluated models.
🧬 Performance by Modality
Score breakdown across 9 omics modalities. Color: β–  low β†’ β–  mid β†’ β–  high (scale 1–5).
πŸ“‹ Full Data Table
Click column headers to sort. All scores are on a 1–5 scale and were assigned by Claude Sonnet 4.6. v2.1: Q27/Q28/Q64 ground truth corrected; Gemini re-evaluated with max_tokens=8192 (thinking mode).
# Model Score Easy Med Hard Expert Factual Reason Complete Uncert Domain Halluc Novel