A multi-omics AI benchmark for spaceflight biomedical data — 21 ML tasks across 9 modalities plus a 100-question LLM evaluation, using data from Inspiration4, the NASA Twins Study, and the JAXA CFE study.
21
ML Tasks
9
Modalities
100
LLM Questions
7
Baselines
9
LLM Models
Baseline Results
Performance across all tasks. The best score in each row is highlighted. E2 and E3 are supplementary tasks with extreme class imbalance.
Task
Name
Category
Tier
Metric
Random
Majority
LogReg
RF
MLP
XGB
LGBM
Performance Analysis
Normalized composite scores, category radar, and difficulty distribution.
Insight: Without effect-size features (fold changes), RF and MLP reach AUPRC values of 0.863 and 0.847, respectively,
suggesting that performance is not driven solely by effect-size thresholding.
LLM Evaluation
A 100-question benchmark with five-dimensional scoring across nine models.
Canonical v2.1 Results
The dedicated interactive LLM leaderboard
contains the current nine-model ranking, dimension breakdown, modality
heatmap, and full scored table.