SpaceOmicsBench

A multi-omics AI benchmark for spaceflight biomedical data — 21 ML tasks across 9 modalities plus a 100-question LLM evaluation, using data from Inspiration4, the NASA Twins Study, and the JAXA CFE study.

21
ML Tasks
9
Modalities
100
LLM Questions
7
Baselines
9
LLM Models
Baseline Results
Performance across all tasks. The best score in each row is highlighted. E2 and E3 are supplementary tasks with extreme class imbalance.
TaskNameCategoryTierMetric RandomMajority LogRegRFMLPXGBLGBM
Performance Analysis
Normalized composite scores, category radar, and difficulty distribution.
Normalized Composite Score — (score − random) / (1 − random)
RF Category Performance — radar view
Difficulty Tier Distribution — 21 tasks
B1 Feature Ablation — AUPRC by feature set
Insight: Without effect-size features (fold changes), RF and MLP reach AUPRC values of 0.863 and 0.847, respectively, suggesting that performance is not driven solely by effect-size thresholding.
LLM Evaluation
A 100-question benchmark with five-dimensional scoring across nine models.
Canonical v2.1 Results
The dedicated interactive LLM leaderboard contains the current nine-model ranking, dimension breakdown, modality heatmap, and full scored table.