Qualia is your autonomous AI researcher.

Use it to autonomously run experiments, interpret results and propose new directions, on your machine or in the cloud.

lm-training
(10)
Tasks
Knowledge

Baseline establishes the comparison point

The 124.4M-parameter transformer reaches a validation loss of 2.824 on WikiText-103. All three variants use the same validation set and training budget.

Sources
FigureBaseline training curve

Training loss decreases over the fixed training budget.

Rotary embeddings improve validation loss

RoPE with a 2,048-token context lowers validation loss to 2.741, a 3.0% improvement over the baseline. This experiment changes both position encoding and context length, so their individual effects remain untested.

Sources

Warmup and shared embeddings beat the baseline

Adding 1,000 warmup steps and tying input and output embeddings lowers validation loss to 2.709, a 4.1% improvement under the same training budget.

Sources

SwiGLU and RMSNorm deliver the largest gain

Gated feed-forward layers with RMS normalization achieve a validation loss of 2.641, down 6.5% from the baseline and the lowest of the three tested variants.

Sources

SwiGLU and RMSNorm lead this comparison

All three variants outperform the baseline. SwiGLU with RMSNorm is the strongest tested combination; combining it with warmup or rotary embeddings remains a follow-up experiment.

Sources
TableTransformer experiment leaderboard
VariantVal. lossGain
Baseline2.824
RoPE + longer context2.7413.0%
LR warmup + weight tying2.7094.1%
SwiGLU + RMSNorm2.6416.5%
FigureValidation loss comparison

SwiGLU + RMSNorm finishes with the lowest validation loss.

SwiGLU + RoPE

Validation loss: 2.583

Sources

SwiGLU + warmup

Validation loss: 2.562

Sources

Combine all improvements

Validation loss: 2.498

Sources

Learning-rate sweep

Validation loss: 2.607

Sources

Regularization sweep

Validation loss: 2.619

Sources

Repeat across three seeds

Validation loss: 2.646 ± 0.008 across three seeds

Sources

Compare follow-up experiments

Combined model: 2.498 · 11.5% below baseline

Sources
Train 125M transformer

Find three ways to improve this transformer and test them in parallel.

The baseline is trained. I’ll compare three improvements on the same validation set:

  • RoPE + longer context
  • LR warmup + weight tying
  • SwiGLU + RMSNorm

Each agent will train its variant and record the results in the task map.

Started 3 agents

All three agents responded. Let's aggregate the data and report results.

Made 1 change

SwiGLU + RMSNorm has the lowest validation loss: 2.641, down 6.5% from the baseline.

RoPE reached 2.741 and warmup with weight tying reached 2.709. The comparison plot and all three experiments are saved in the notebooks.

Go one level deeper. Combine the improvements, tune the winner, and check that the results hold up.

I’ll launch six follow-up agents: combine SwiGLU with RoPE, add warmup, combine all improvements, sweep learning rates, tune regularization, and repeat the winner across three seeds.

Started 6 follow-up agents

All six follow-up agents finished. I’m comparing their results with the first-round winner and checking the seed runs.

Combining all improvements wins: validation loss 2.498, down 11.5% from the baseline.

  • SwiGLU + RoPE: 2.583
  • SwiGLU + warmup: 2.562
  • All improvements: 2.498
  • Learning-rate sweep: 2.607
  • Regularization sweep: 2.619
  • Original SwiGLU winner, three seeds: 2.646 ± 0.008

The seed runs support the first-round gain. The combined model is the next candidate to validate across seeds; all six experiment notebooks are saved.

Qualia supports every step of the research process.

Process and work with data

  • Unify diverse datasources across files, datasets, and the web
  • Autonomously clean and preprocess data

Illustrative workflow: load the TAM dataset, coordinate eight agents, prepare data in Python notebooks, and compare a 62-second pipeline against a 192-second baseline. Dataset and benchmark values are simulated.

Build recurring pipelines

  • Run workflows on a recurring schedule to generate live reports
  • Autonomously refine pipelines based on feedback to improve future runs

Build and train ML models

  • Autonomously train models and improve existing ones, from linear regression to deep neural networks
  • Easily post-train 10B-parameter foundation models on your data for specialized tasks

Explore and test new ideas

  • Fan out agents to test features, architectures, and training strategies in parallel
  • Diagnose model failures and autonomously propose and run follow-up experiments

Illustrative workflow: twelve agents receive instructions through the app’s To Agent chat rows, then their experiments appear in the task map. Content and timing are simulated; the layout follows the cloud app.

Write up and share your results

  • Create publication-ready papers and writeups that clearly summarize research
  • Share your results and collaborate with others in chats

Illustrative workflow: synthesize experiments into a writeup with figures and limitations, inspect a linked claim, and trace it to the notebook output behind the conclusion. Data and timings are simulated.

Work autonomously

  • Give Qualia a research goal and let it work independently for hours
  • Receive progress updates in Slack to steer research as it runs

Why Qualia works

  1. Coordinates dozens of research agents.

    Qualia breaks research goals into focused experiments that agents work on in parallel. Agents share findings and can investigate each other’s results, using what they learn to guide the next experiments.

  2. Research accumulates across sessions.

    Qualia remembers your experiments, findings, and assumptions in a shared knowledge graph that persists across sessions. Agents build on what worked and what failed, connecting insights across your research to propose creative new ideas.

  3. Traceable.

    Everything the agent does or says has provenance and is traceable to a query or line of code. You can inspect the evidence behind each finding and follow the steps that led to it.

  4. A beautiful UI that runs wherever you are.

    Qualia is built on primitives that run on any platform — desktop, Linux cluster, or in the cloud. It runs as an agent that manipulates Jupyter notebooks, or as a headless terminal agent.

Infinite possibilities

Posttraining

Qualia can posttrain large LMs, intelligently understand where they are failing, and improve. Run it on Qualia Cloud to take advantage of our pool of GPUs, or use standard libraries like Tinker or Fireworks.

Biotech

Qualia can preprocess large amounts of multiomics data, build repeatable, auto-improving data pipelines, and autonomously fit interpretable models over it. Use Qualia's writeup feature to build clear, publishable reports. Qualia is HIPAA-compliant and can be run in ZDR mode.

Quantitative finance

Qualia helps you explore alternative datasets, test trading hypotheses in parallel, and understand where your models fail. Work with your existing data and code to run backtests, investigate results, and propose follow-up experiments, with every finding traceable to the analysis behind it.

Data science

Qualia works with your existing data and models to understand where predictions fail, discover useful features, and run experiments to improve them. You can delegate data preparation, feature engineering, and model iteration to coordinated agents that record their findings in a reproducible knowledge graph.

And many more...

Research together in Qualia Cloud

Run experiments with cloud compute in your browser, and bring your team into the same research workspace.

0:00

Illustrative collaboration preview: share a cloud workspace with Richard as a Commenter, then Robbie, Richard, and Patrick discuss a shared notebook and ask Qualia about the results. The multi-person chat previews UI in development; messages and research data are fictional.

Share notebooks, agent conversations, and evidence. Teammates can comment on findings and ask Qualia questions about the research.

What could you discover with Qualia?

Get started

Need Qualia for enterprise?

Qualia Enterprise provides your organization with additional capabilities, security and control.

Learn about Qualia for Enterprise
Qualia