Method · version 1.1

Measuring whether AI accelerates science

Definitions, source rules, and analysis standards for the tracker.

The tracker now separates three questions that are easy to mix together:

  1. What happened? The event ledger records important claims, results, benchmarks, failures, and institutional changes.
  2. Did the rate of progress change? Acceleration panels follow the same outcome over time with a stable definition.
  3. Did AI cause the change? Causal studies compare differently exposed targets, researchers, institutions, or fields.

The three layers support each other, but they are not interchangeable.

The event ledger is not a time series

The event ledger is selected editorially. Research effort, source coverage, search tools, and inclusion standards change over time. Older events can also be added retrospectively.

For that reason, the number of tracker events in a month or year must never be used as an estimate of scientific acceleration.

Events are still essential. They help us discover possible outcomes, record the evidence behind individual claims, and explain why a measured change may have happened. A strong event may later become one observation inside a stricter series.

Evidence lanes

Every series belongs to one lane:

  • Discovery outcomes — verified discoveries and scientific performance frontiers. These are the only series eligible for a future cross-domain synthesis.
  • Validation and translation — expert validation, clinical translation, and other downstream outcomes.
  • Research process — time, experiments, human effort, or cost per validated result.
  • Controls — reporting-dominated or lower-exposure series used to test alternative explanations.

Monitoring status and evidence lane answer different questions. A control can be fully monitored without being pooled with discovery outcomes. A discovery outcome can remain under audit until its unit, dates, and rebuild path are ready.

Each series also records whether its denominator is strong, partial, or absent. Process proxies may link to the outcome series they are intended to explain.

What qualifies as an acceleration panel

A series can enter the monitoring stage only when it has:

  • a clear unit of progress;
  • a stable counting or measurement rule;
  • a defensible dating rule;
  • a useful historical baseline or a denominator frozen before prospective observation;
  • a public and auditable source;
  • a documented update or rebuild path;
  • explicit limits on what the series measures.

Good units include:

  • a fixed open problem moving from open to verified as solved;
  • a protein target first crossing a predeclared ligand-potency threshold;
  • an independently confirmed physical-performance record;
  • the time required to complete a defined research loop;
  • a validated discovery per unit of search effort.

Paper counts, press releases, tool launches, and changing benchmark leaderboards can provide context. They are not treated as discoveries by default.

The construct must be named

Every series states what it actually measures. Examples include discovery, validation, disclosure, publication, benchmark performance, frontier magnitude, workflow productivity, and downstream use.

This matters because nearby measures can behave very differently. A rise in vulnerability disclosures is not the same as a rise in serious vulnerabilities found. A rise in papers is not the same as a rise in correct discoveries. A faster simulation is not automatically a faster end-to-end research programme.

Dates are part of the measurement

A scientific result can have several dates:

  • discovery;
  • experiment completion;
  • independent verification;
  • public announcement;
  • paper publication;
  • database entry.

A series must declare which date it uses and why. Dates are not silently substituted. When only an approximate date exists, the precision and basis must be recorded.

A database-maintenance date must not be treated as a discovery date. When historical solution dates cannot be reconstructed defensibly, the better design is to freeze the denominator and observe prospectively rather than manufacture a retrospective time series.

Quantity, quality, and denominators

Raw counts are rarely enough.

Where possible, each panel keeps three views:

  1. Quantity: how many outcomes occurred.
  2. Quality or magnitude: how important, severe, novel, or large the outcomes were.
  3. Denominator: how much opportunity or effort produced them.

Useful denominators include active targets, open problems, observing days, experiments attempted, submissions, researchers, compute, or search expenditure.

When no denominator exists, that absence is shown as a first-class limitation.

Descriptive and causal claims

The first estimand is descriptive:

How much did the rate or magnitude of publicly observable, externally checkable progress change relative to its earlier trajectory?

The stronger estimand is causal:

How much greater was progress among AI-exposed units than it would have been without that exposure?

A visible break in a chart does not establish causality. The same pattern can be produced by more funding, new instruments, reporting changes, backlogs, database expansion, incentives, or unrelated technology.

Each published estimate therefore receives a causal-status label:

  • descriptive_change;
  • ai_consistent_change;
  • quasi_experimental_evidence;
  • causal_experiment;
  • causality_unresolved.

Intervention dates are domain-specific

There is no single date when AI began affecting all of science.

Possible interventions include AlphaFold availability, broad access to language models, a scientific-agent release, an autonomous-lab deployment, or an institution's internal rollout. A series records the candidate dates and the evidence behind them.

Where possible, the primary intervention and analysis are specified before the post-period is examined. Exploratory breakpoints remain clearly labelled as exploratory.

A cohort freeze is not automatically an intervention. It can be only the point at which prospective observation begins.

Controls and falsification tests

Some outcomes should respond strongly to AI. Others should respond weakly because they depend mainly on new instruments, field observation, or physical infrastructure.

Low-exposure outcomes can serve as negative controls. If high- and low-exposure series move in the same way at the same time, the common cause may be funding, reporting, or another broad change rather than AI.

Controls do not make a study causal by themselves. They make alternative explanations easier to test.

Series lifecycle

A series moves through these states:

  1. candidate — a promising idea that has not passed the audit;
  2. auditing — the unit, dates, source, and rebuild path are being checked;
  3. monitoring — the definition is fixed enough for regular updates;
  4. paused — the source or definition currently prevents reliable updates;
  5. retired — the series is preserved for history but no longer used.

A candidate can be rejected without being hidden. The reason should remain visible.

Analysis and uncertainty

Published panels should report effect sizes and uncertainty, not only labels such as “accelerating.” Depending on the outcome, the primary model may use:

  • segmented count models;
  • survival analysis for problem-to-solution time;
  • frontier-improvement models;
  • difference-in-differences;
  • event studies;
  • hierarchical estimates across related series.

Robustness checks should vary plausible dates, baseline windows, partial-period treatment, quality thresholds, and attribution definitions.

A cross-domain summary must preserve heterogeneity. A mathematical proof, a potent ligand, a vulnerability disclosure, and a solar-cell record are not interchangeable units.

A new prospective panel with zero events should report that exact state. It should not annualize a short partial period into an acceleration or stagnation verdict.

AI attribution

AI attribution is recorded separately from the scientific outcome.

Useful evidence includes author reports, model transcripts, tool logs, public artifacts, independent reconstruction, and explicit method descriptions. Affiliation with an AI company is weaker evidence than a documented AI-assisted workflow.

Sensitivity analyses should distinguish narrow, medium, and broad attribution rules rather than force every case into a single binary label.

Tool availability, an AI-attempt link, or a formalized problem statement is not by itself evidence that AI contributed to a discovery.

Inspiration and differences

This measurement programme is inspired in part by METR's article “LLMs' Contribution to Discoveries” and the public tecunningham/ai-discovery-data repository. Their work shows the value of stable, rebuildable discovery series, local data provenance, and machine-checked numerical claims.

This tracker keeps its broader event ledger as a separate layer and adds explicit distinctions between descriptive change and causal attribution. No data or code from the reference repository is copied into this project by default; series should be rebuilt from their primary upstream sources or linked as external references with clear attribution.

Current implementation stage

The measurement contract, schemas, validation, public interfaces, and three-layer navigation are in place. Candidate series remain separate from monitored evidence, and study data builds remain separate from causal results.

Four monitored vertical slices now exercise different parts of the instrument:

  1. curl disclosures — a high-frequency fixed-codebase reporting control with committed source provenance, severity, and attribution sensitivities;
  2. OpenSSL disclosures — an independently maintained comparison panel with the same audit structure;
  3. NREL single-crystal silicon efficiency — a natural-science physical frontier with a frozen technology category, 50 eligible records, 26 record advances, and explicit area and metadata sensitivities;
  4. Erdős Problems 1–100 — a prospective fixed-risk-set cohort with 54 unresolved baseline problems, 46 resolved-before-freeze strata, and a verified-event protocol that refuses to treat catalogue edits as historical discoveries.

The first public causal study asks whether AlphaFold shortened the target-to-potent-ligand path. Its BindingDB feasibility audit now contains 849 core protein targets and passes five of eight evidence gates. It remains data_build / not_estimated because binding-site usability, covariate overlap, and pre-trend credibility are unresolved, and the strict 100 nM outcome contains only two post-treatment first-crossing events.

The next measurement priorities are:

  1. freeze a defensible AlphaFold binding-site or domain usability rule;
  2. audit BindingDB covariate overlap and pre-treatment event-study leads before any effect estimate;
  3. monitor the Erdős cohort prospectively and verify source changes before adding events;
  4. add research-effort denominators and exposure measures where public data allow them;
  5. estimate effects only after the relevant definitions and specifications are frozen.

The goal is not to publish an early headline. It is to build an instrument whose later headlines can be trusted.