Results β MLE-Bench Low
We evaluated Qualia, our autonomous research agent, on MLE-Bench Low, a set of 22 Kaggle machine learning competitions that test an agent's ability to train models, engineer features, and produce competition-grade submissions. Our agent performed these with no human intervention other than the original prompt, spawning in many cases additional agents to help with testing out multiple modeling and feature engineering choices in parallel (see the Agents column), or to build on its prior attempts. Click on any row to see the chat history and all agents spawned, as well as the tasks into which the agent split the work. Qualia medalled on 77.2% of the tasks, tying for 4th place on the dataset while using far less time than the allotted maximum (24h).
| Competition | Metric | Score | Medal β² | Rank | Percentile | Agents | Time |
|---|---|---|---|---|---|---|---|
| Aerial Cactus Identification | auc-roc | 1.0000β² | π₯#1 | 1/1221 | 100% | 1 | 42m |
| APTOS 2019 Blindness Detection | quadratic-weighted-kappa | 0.9380β² | π₯#1 | 1/2929 | 100% | 5 | 13h 20m |
| Detecting Insults in Social Commentary | auc-roc | 0.8478β² | π₯#1 | 1/50 | 100% | 1 | 12m |
| Dogs vs. Cats Redux: Kernels Edition | log-loss | 0.0304βΌ | π₯#1 | 1/1315 | 100% | 1 | 18m |
| Leaf Classification | multi-class-log-loss | 0.0000βΌ | π₯#1 | 1/1596 | 100% | 1 | 23m |
| Plant Pathology 2020 - FGVC7 | mean-column-wise-roc-auc | 0.9920β² | π₯#1 | 1/1318 | 100% | 1 | 20m |
| Tabular Playground Series - Dec 2021 | multi-class-classification-accuracy | 0.9601β² | π₯#1 | 1/1189 | 100% | 3 | 14m |
| Jigsaw Toxic Comment Classification Challenge | column-wise ROC AUC | 0.9876β² | π₯Gold | 14/4539 | 99.7% | 1 | 1h 58m |
| Histopathologic Cancer Detection | auc-roc | 0.9882β² | π₯Gold | 10/1149 | 99.2% | 2 | 1h 16m |
| NOMAD2018 Predict transparent conductors | mean-column-wise-rmsle | 0.0555βΌ | π₯Gold | 8/879 | 99.2% | 1 | 13m |
| Denoising Dirty Documents | root_mean_squared_error | 0.0076βΌ | π₯Gold | 3/162 | 98.8% | 1 | 30m |
| Text Normalization Challenge - English Language | accuracy | 0.9973β² | π₯Gold | 10/261 | 96.6% | 1 | 1h 45m |
| MLSP 2013 Birds | auc-roc | 0.9420β² | π₯Gold | 4/81 | 96.3% | 1 | 14m |
| The ICML 2013 Whale Challenge - Right Whale Redux | auc-roc | 0.9904β² | π₯Gold | 9/129 | 93.8% | 1 | 21m |
| Random Acts of Pizza | auc-roc | 0.7980β² | π₯Silver | 33/462 | 93.1% | 1 | 42m |
| Text Normalization Challenge - Russian Language | accuracy | 0.9823β² | π₯Silver | 32/163 | 81% | 1 | 1h 24m |
| Spooky Author Identification | multi-class-log-loss | 0.2891βΌ | π₯Bronze | 100/1242 | 92% | 5 | 2h 23m |
| Dog Breed Identification | multi-class-log-loss | 0.2647βΌ | βNone | 425/1281 | 66.9% | 1 | 5h 0m |
| Tabular Playground Series - May 2022 | auc-roc | 0.9775β² | βNone | 558/1152 | 51.6% | 9 | 2h 20m |
| New York City Taxi Fare Prediction | root-mean-squared-error | 3.7864βΌ | βNone | 864/1485 | 41.9% | 1 | 23h 30m |
| RANZCR CLiP - Catheter and Line Position Challenge | auc-roc | 0.9653β² | βNone | 936/1547 | 39.6% | 3 | 23h 30m |
| SIIM-ISIC Melanoma Classification | auc-roc | 0.8993β² | βNone | 2008/3308 | 39.3% | 9 | 22h 41m |