33 tools on the map. 8 are wired today, each with a page saying exactly what it reads and what it writes.

Experiment throughput, and why growth teams get stuck

Experiment throughput is the number of well-run experiments a team finishes in a period. Most ideas lose and most teams are bounded by engineering time, so throughput is what a programme returns.

By Lucas Duys, reviewed by Lucas Duys. Published 2026-09-04.

What throughput is

Experiment throughput is the number of experiments a team takes from a hypothesis to a decision in a period, counted only when each one was run well enough to trust: sized before it started, run to its planned sample, read against a primary metric and its guardrails, and closed with a decision. A test that was stopped the day it looked good, or one that ran for a quarter and never reached a sample, does not count, because it did not produce a decision anyone should act on.

The reason the number matters is the arithmetic of ideas. At Microsoft, only about one third of the ideas tested improved the metric they were designed to improve [1]. At Bing, where thousands of experiments run a year, most fail, and the ones that succeed move key metrics by a fraction of a percent to about one percent [2]. Those are the figures of a mature programme with strong ideas and a large audience. A team that runs four experiments a quarter and expects two of them to win is planning against numbers nobody has reported.

If most ideas lose, the return on an experimentation programme is set by how many ideas it can test, not by how good any one of them looks in the planning meeting. That is what makes throughput the operating number. Every other property of a programme, the quality of hypotheses, the statistics, the tooling, shows up in the return through it.

Where teams get stuck

Ask a growth team why it ran four experiments last quarter rather than twelve and the answer is almost never a shortage of ideas. The backlog is long. The limits are elsewhere, and they stack.

Engineering time. An experiment is a change to the product, built to production quality, behind a flag, with exposure logged correctly. On most teams that work competes with the roadmap for the same engineers, and it loses, because a roadmap item ships something everyone agreed on and an experiment might ship nothing. The idea waits for a sprint, then for a review, then for a deploy window.

Traffic. A step deep in a funnel sees a fraction of the visitors the top sees, and a small effect on a small population takes a long time to detect. Detecting an effect ten times smaller needs a hundred times more users [1], so a low-traffic step can only test large changes, and a team that keeps queuing small ones there is queuing tests that cannot finish.

Analysis and trust. A result has to be read: the primary metric at its planned sample, the guardrails at their windows, the exposure checked for a sample ratio mismatch. On a team without an experimentation platform that is a person with a notebook, and the person is the bottleneck.

Coordination. Two experiments on the same screen interfere; two teams testing the same metric argue about attribution. Booking.com, which runs more than a thousand experiments at once, got there by decentralising experimentation and investing considerable time in training so that people outside the central team could run their own [3]. Without that, every experiment passes through the same few hands.

Which limit binds first

The order matters. Traffic is a property of the product and cannot be bought quickly. Analysis and coordination are organisational and take a year to change. Engineering time is the limit that binds first on most product-led teams, and it is also the one that can move.

How programmes grow

The published accounts of experimentation at scale agree on the shape of the path. A study of companies adopting continuous experimentation found that they seldom succeed in evolving the practice, and describes the ones that did in three phases: a technical phase, in which the platform and the instrumentation are built; an organisational phase, in which teams outside the central group learn to run and read their own tests; and a business phase, in which experiment results drive decisions about what to build at all [4]. Microsoft's own account of running over two hundred concurrent experiments a day names the same three kinds of work, cultural and organisational, engineering, and trustworthiness [1].

For a growth team that is not a search engine, the practical version is shorter. Throughput rises when the cost of getting one experiment live falls, when the tests that cannot finish are refused before they start, and when reading a result stops needing a specialist. Each of those is a bottleneck a team can name and measure this quarter: hours of engineering per experiment, the share of tests that reached their planned sample, and days from a finished test to a decision.

A worked example, on an example funnel

The front page draws a self-serve product with 10,000 visits a month, 4,400 signups, 3,900 reaching workspace setup, 1,900 creating a first project and 1,400 paying. The figures are the blueprint's example and not a customer's. On that funnel the onboarding experiment that delays setup until after the first project is written against the 3,900 who reach setup, and the pricing experiment against a smaller population still. At those volumes, a modest relative change at the setup step takes weeks to detect and the same change at the plan chooser takes longer than a reasonable test should run, which is the feasibility question the guide on whether a test is worth running works through. A team on this funnel does not get to twelve experiments a quarter by testing more things at the plan chooser. It gets there by testing where the traffic is, refusing the tests that cannot finish, and cutting the engineering cost of each one that can.

What Athren does about it

Athren's answer to the bottleneck is the loop on the front page. It reads the customer journey, finds where it is losing the most value, decides whether an experiment there is worth running at all, writes the hypothesis, builds the variant inside the customer's own components behind a flag they already have, and waits for a person to approve it before anything reaches traffic. The engineering work of getting an experiment live is what it takes off the team. The decision about what to test, and the approval of every change, stays with the team. The guardrail guide says what it watches while an experiment runs.

What to measure this quarter

Three numbers, kept honestly, say whether throughput is rising.

  • Experiments finished, where finished means run to the planned sample and closed with a decision. Not started, not shipped.
  • Engineering hours per experiment, from the moment the hypothesis is written to the moment the variant is live behind its flag. This is the cost the loop exists to cut.
  • Share of tests that reached their planned sample. A programme whose tests keep running out of time is one that starts tests it should have refused.

A team that moves the second and third numbers moves the first. A team that only counts the first learns nothing about why it is stuck.

Sources

  1. 1Online Controlled Experiments at Large Scale, Microsoft (KDD 2013). Read 2026-09-04.
  2. 2Seven Rules of Thumb for Web Site Experimenters, Microsoft and LinkedIn (KDD 2014). Read 2026-09-04.
  3. 3Democratizing online controlled experiments at Booking.com, arXiv, 2017-10-23. Read 2026-09-04.
  4. 4The Evolution of Continuous Experimentation in Software Product Development, Microsoft Research (ICSE 2017). Read 2026-09-04.

Give us one funnel.

Twenty five minutes, about how you run experiments today. No access, no commitment.