Thesis
Models that think from first principles.
Our vision is low-cost models that think from first principles: they reason from measured facts and mechanisms instead of recalling what has already been written, propose ideas no one has tested, and use them to advance science and technology.
01 The gap
Discovery starts where the reading stops.
Models have become very good at recalling and recombining what people have already written. Discovery needs three things recall can’t supply: reasoning from measured facts, proposing what nobody has tried, and running the experiment that shows whether it holds.
People learn that by running experiments: many attempts, most of them wrong, each leaving a record of why. We think models can learn it the same way.
02 Our bet
Originality is learned from experiments, not from text.
Every problem runs through the same five steps. Steps 4 and 5 are what make the next problem take fewer experiments.
-
Propose
An idea, and the reason it should work.
-
Run
The experiment that could prove it wrong.
-
Check
A re-run with new random seeds, scored by code the agent can’t edit.
-
Remember
The result and why it failed or held, stored either way.
-
Reinforce
Train on results that passed the check, then start the next problem.
Memory
Grounded in experiments it ran.
A record of every experiment: what was tried, what happened, and why it failed or held. The model reasons from results it produced and verified, not only from text it was trained on.
Reinforcement learning
Rewarded for results that hold.
Models train on outcomes that passed the re-run. The reasoning steps behind those outcomes are reinforced, so the next problem needs fewer experiments to reach a verified result.
03 First principles
What we mean by a first-principles thinker.
Six behaviours we intend to measure. Fluent writing is not one of them.
-
Starts from measured facts and mechanisms, not precedent.
-
Proposes ideas that aren’t in what it has read, and states why they should work.
-
Designs the experiment that could prove it wrong.
-
Tells a reproducible result from a lucky run.
-
Changes its conclusion when the evidence does, and records why.
-
Needs fewer experiments on each new problem.
04 How we work
Six rules we build by.
-
First principles
Research is a science we measure
We treat autonomous research, and systems that improve their own ability to do research, as a new scientific discipline. We define its quantities mathematically, starting with experiments per verified result, and measure them on problems with known answers. Then we publish predictions beyond that data, each with the result that would disprove it.
-
Verification
Checked before it’s kept
A result enters memory only after a re-run with new random seeds reproduces it. The code that scores it can’t be read or edited by the agent that produced it.
-
Reinforcement
Reward only what held up
Reinforcement learning trains only on results that passed that re-run. An agent’s own report of success earns no reward, because an agent that can influence its reward learns to game it.
-
Human in the loop
People own the measure
We don’t expect AI to exceed people at every kind of judgment, so three decisions stay with people: which problems to solve, what counts as better, and whether a result ships. When we find another decision that needs a person, we add it to the system as a named, required step.
-
Product-led
Proven on real workloads
Research counts once it runs in a product on real workloads, starting with AIM for AI inference. A benchmark can be overfit. A live workload keeps changing and records the cost of every wrong decision, which keeps our claims honest over years rather than for one paper.
-
The long game
Low-cost models that reason
Our autonomous research systems are built toward one output: low-cost models that think from first principles and advance research themselves. We judge those models by what it costs them to reach a verified result, not by parameter count or benchmark scores.
05 The path
Where we start, and where it leads.
An intended sequence, not a schedule.
-
Measured problems
Start where better is a number: the quality, latency and cost of AI inference, with AIM.
Explore AIM -
Agents that remember
Research agents that keep a verified memory of every experiment and need fewer experiments on each new problem.
-
Models that learn from it
Reinforcement learning on that verified memory, so a model reaches results agents already found in fewer experiments.
-
Low-cost models that advance research
First-principles models, judged by the cost of each verified result, applied to open problems in science and software.
06 Work with us
Build it with us.
Researchers, engineers and partners who want models that discover: we’d like to hear from you.
How we’ll test this, and how we’ll report what we find.
Read the researchReview build — contact address not yet connected.