Partner with us

Thesis

Models that think from first principles.

Our vision is low-cost models that think from first principles: they reason from measured facts and mechanisms instead of recalling what has already been written, propose ideas no one has tested, and use them to advance science and technology.

01 The gap

Discovery starts where the reading stops.

Models have become very good at recalling and recombining what people have already written. Discovery needs three things recall can’t supply: reasoning from measured facts, proposing what nobody has tried, and running the experiment that shows whether it holds.

People learn that by running experiments: many attempts, most of them wrong, each leaving a record of why. We think models can learn it the same way.

02 Our bet

Originality is learned from experiments, not from text.

Every problem runs through the same five steps. Steps 4 and 5 are what make the next problem take fewer experiments.

  1. Propose

    An idea, and the reason it should work.

  2. Run

    The experiment that could prove it wrong.

  3. Check

    A re-run with new random seeds, scored by code the agent can’t edit.

  4. Remember

    The result and why it failed or held, stored either way.

  5. Reinforce

    Train on results that passed the check, then start the next problem.

Memory

Grounded in experiments it ran.

A record of every experiment: what was tried, what happened, and why it failed or held. The model reasons from results it produced and verified, not only from text it was trained on.

Reinforcement learning

Rewarded for results that hold.

Models train on outcomes that passed the re-run. The reasoning steps behind those outcomes are reinforced, so the next problem needs fewer experiments to reach a verified result.

03 First principles

What we mean by a first-principles thinker.

Six behaviours we intend to measure. Fluent writing is not one of them.

  1. Starts from measured facts and mechanisms, not precedent.

  2. Proposes ideas that aren’t in what it has read, and states why they should work.

  3. Designs the experiment that could prove it wrong.

  4. Tells a reproducible result from a lucky run.

  5. Changes its conclusion when the evidence does, and records why.

  6. Needs fewer experiments on each new problem.

04 How we work

Six rules we build by.

  • First principles

    Research is a science we measure

    We treat autonomous research, and systems that improve their own ability to do research, as a new scientific discipline. We define its quantities mathematically, starting with experiments per verified result, and measure them on problems with known answers. Then we publish predictions beyond that data, each with the result that would disprove it.

  • Verification

    Checked before it’s kept

    A result enters memory only after a re-run with new random seeds reproduces it. The code that scores it can’t be read or edited by the agent that produced it.

  • Reinforcement

    Reward only what held up

    Reinforcement learning trains only on results that passed that re-run. An agent’s own report of success earns no reward, because an agent that can influence its reward learns to game it.

  • Human in the loop

    People own the measure

    We don’t expect AI to exceed people at every kind of judgment, so three decisions stay with people: which problems to solve, what counts as better, and whether a result ships. When we find another decision that needs a person, we add it to the system as a named, required step.

  • Product-led

    Proven on real workloads

    Research counts once it runs in a product on real workloads, starting with AIM for AI inference. A benchmark can be overfit. A live workload keeps changing and records the cost of every wrong decision, which keeps our claims honest over years rather than for one paper.

  • The long game

    Low-cost models that reason

    Our autonomous research systems are built toward one output: low-cost models that think from first principles and advance research themselves. We judge those models by what it costs them to reach a verified result, not by parameter count or benchmark scores.

05 The path

Where we start, and where it leads.

An intended sequence, not a schedule.

  1. 01In development

    Measured problems

    Start where better is a number: the quality, latency and cost of AI inference, with AIM.

    Explore AIM
  2. 02

    Agents that remember

    Research agents that keep a verified memory of every experiment and need fewer experiments on each new problem.

  3. 03

    Models that learn from it

    Reinforcement learning on that verified memory, so a model reaches results agents already found in fewer experiments.

  4. 04

    Low-cost models that advance research

    First-principles models, judged by the cost of each verified result, applied to open problems in science and software.

06 Work with us

Build it with us.

Researchers, engineers and partners who want models that discover: we’d like to hear from you.

How we’ll test this, and how we’ll report what we find.

Read the research