QuadrillionQuadrillion

Results β€” MLE-Bench Low

We evaluated Qualia, our autonomous research agent, on MLE-Bench Low, a set of 22 Kaggle machine learning competitions that test an agent's ability to train models, engineer features, and produce competition-grade submissions. Our agent performed these with no human intervention other than the original prompt, spawning in many cases additional agents to help with testing out multiple modeling and feature engineering choices in parallel (see the Agents column), or to build on its prior attempts. Click on any row to see the chat history and all agents spawned, as well as the tasks into which the agent split the work. Qualia medalled on 77.2% of the tasks, tying for 4th place on the dataset while using far less time than the allotted maximum (24h).

πŸ₯‡ 14πŸ₯ˆ 2πŸ₯‰ 1β€” 5
16/22 medals
CompetitionMetricScoreMedal β–²RankPercentileAgentsTime
Aerial Cactus Identificationauc-roc1.0000β–²πŸ₯‡#11/1221
100%
142m
APTOS 2019 Blindness Detectionquadratic-weighted-kappa0.9380β–²πŸ₯‡#11/2929
100%
513h 20m
Detecting Insults in Social Commentaryauc-roc0.8478β–²πŸ₯‡#11/50
100%
112m
Dogs vs. Cats Redux: Kernels Editionlog-loss0.0304β–ΌπŸ₯‡#11/1315
100%
118m
Leaf Classificationmulti-class-log-loss0.0000β–ΌπŸ₯‡#11/1596
100%
123m
Plant Pathology 2020 - FGVC7mean-column-wise-roc-auc0.9920β–²πŸ₯‡#11/1318
100%
120m
Tabular Playground Series - Dec 2021multi-class-classification-accuracy0.9601β–²πŸ₯‡#11/1189
100%
314m
Jigsaw Toxic Comment Classification Challengecolumn-wise ROC AUC0.9876β–²πŸ₯‡Gold14/4539
99.7%
11h 58m
Histopathologic Cancer Detectionauc-roc0.9882β–²πŸ₯‡Gold10/1149
99.2%
21h 16m
NOMAD2018 Predict transparent conductorsmean-column-wise-rmsle0.0555β–ΌπŸ₯‡Gold8/879
99.2%
113m
Denoising Dirty Documentsroot_mean_squared_error0.0076β–ΌπŸ₯‡Gold3/162
98.8%
130m
Text Normalization Challenge - English Languageaccuracy0.9973β–²πŸ₯‡Gold10/261
96.6%
11h 45m
MLSP 2013 Birdsauc-roc0.9420β–²πŸ₯‡Gold4/81
96.3%
114m
The ICML 2013 Whale Challenge - Right Whale Reduxauc-roc0.9904β–²πŸ₯‡Gold9/129
93.8%
121m
Random Acts of Pizzaauc-roc0.7980β–²πŸ₯ˆSilver33/462
93.1%
142m
Text Normalization Challenge - Russian Languageaccuracy0.9823β–²πŸ₯ˆSilver32/163
81%
11h 24m
Spooky Author Identificationmulti-class-log-loss0.2891β–ΌπŸ₯‰Bronze100/1242
92%
52h 23m
Dog Breed Identificationmulti-class-log-loss0.2647β–Όβ€”None425/1281
66.9%
15h 0m
Tabular Playground Series - May 2022auc-roc0.9775β–²β€”None558/1152
51.6%
92h 20m
New York City Taxi Fare Predictionroot-mean-squared-error3.7864β–Όβ€”None864/1485
41.9%
123h 30m
RANZCR CLiP - Catheter and Line Position Challengeauc-roc0.9653β–²β€”None936/1547
39.6%
323h 30m
SIIM-ISIC Melanoma Classificationauc-roc0.8993β–²β€”None2008/3308
39.3%
922h 41m
MLE-Bench Low Results β€” Quadrillion