33 tools on the map. 8 are wired today, each with a page saying exactly what it reads and what it writes.

Experiment backlogs versus continuous experimentation

A backlog ranks experiment ideas before they run. Continuous experimentation makes running one cheap. Published success rates say ranking is the weaker lever, and which one a team needs depends on what its tests cost.

By Lucas Duys, reviewed by Lucas Duys. Published 2026-09-04.

The two models, stated plainly

A backlog is a list of experiment ideas, scored and ordered, worked from the top. Continuous experimentation is a system in which getting an idea in front of traffic is cheap enough that the order matters less than the rate. Both are real operating models with published practice behind them, and a team usually has one without having decided to.

The distinction is not tooling and it is not ambition. It is where the team spends its judgment. A backlog spends it on choosing before the test. A continuous programme spends it on the loop that runs after: instrumentation, reading, and deciding what the result changes. Neither is free, and picking one over the other is a decision about which of those two costs a team can actually pay.

What a ranked backlog is betting on

A backlog's ordering is a prediction. Put an idea at the top and you are claiming it will move the metric more, per unit of effort, than the ideas below it. That claim can be checked against how often ranked ideas turn out to be right at all.

The published numbers are not encouraging. A 2022 summary of historical success rates, meaning the share of experiments an organisation believes were true improvements to its overall evaluation criterion, gives 33 percent at Microsoft, 20 percent in Avinash Kaushik's account, 15 percent at Bing, 10 percent at Booking.com, Google Ads and Netflix, and 8 percent at Airbnb Search [1]. These are the rates of teams with platforms, specialists and years of practice.

Take the middle of that range. If nine in ten of the ideas a team ranks will not win, then the top of the list and the middle of the list are both mostly losers, and the ordering is sorting a set whose values it cannot see. Ranking harder does not fix that. It is not a failure of the scoring method; it is what a prior of ten percent does to any ordering built on estimates.

What the scoring method actually says about itself

RICE, the most widely copied of these frameworks, scores each item as reach times impact times confidence, divided by effort, giving total impact per unit of time worked. Impact is a five-point scale from 0.25 for minimal to 3 for massive; confidence is 100, 80 or 50 percent; effort is person-months [2].

Its author is candid about what those numbers are. "Choosing an impact number may seem unscientific," he writes, "but remember the alternative: a tangled mess of gut feeling," and says the scores "shouldn't be used as a hard and fast rule" [2]. That is a fair description of the tool, and it is also the point. RICE is a way of making a group's guesses explicit and comparable. It was never a claim that the guesses are accurate, and the success rates say they are not.

So a backlog is worth keeping for what it genuinely does: it makes disagreement visible, it forces effort to be named, and it stops the loudest idea winning by default. It is not a substitute for running the test.

What continuous experimentation asks for instead

The alternative is not to stop choosing. It is to make being wrong cheap, so that the choice carries less weight. The academic model of this is the RIGHT model, which sets out what a continuous experimentation system needs: the ability to release minimum viable products or features with suitable instrumentation, to design and manage experiment plans, to link experiment results with a product roadmap, and to manage a flexible business strategy [3].

That is a demanding list, and the same work names the obstacles: rapid design of experiments, instrumenting software to collect, analyse and store the data, and integrating experiment results into both the product development cycle and the software development process [3]. None of those is a scoring decision. All of them are engineering and organisational cost, paid once and then amortised over every test afterwards.

Booking.com's account of its own programme is the organisational half of the same bill. It describes designing for democratisation and decentralisation from the start: a central repository of successes and failures so knowledge is shared, a generic code library that keeps experimentation loosely coupled from business logic, close and transparent monitoring of the data pipelines so people trust the numbers, and safeguards that let anyone own an experiment end to end [4]. That is what it costs to move the bottleneck off a central team.

Which one a team needs

The honest answer is that it depends on the price of one test, and that price is knowable.

Where an experiment is cheap to build and the surface has traffic, the ranking is the wrong place to spend the argument. Run more of them, in the order they are ready, and let the results do the sorting. The throughput guide covers what bounds that rate and what to measure.

Where an experiment is expensive, or the surface is too thin to detect anything, the backlog is still the right tool, because the effort term in the score is real and the ordering saves genuine money. But the first question there is not what to rank; it is whether the test can finish at all, which the feasibility checker settles in four fields and the feasibility guide works through. A ranked list of tests that cannot reach a sample is a list of tests nobody should run in any order.

Most teams are a mix, and the useful move is to split the list rather than to pick a side: the cheap, high-traffic tests leave the backlog and go into flow, and the backlog keeps the expensive ones it was always for.

What Athren does about it

Athren works on the price of one test rather than on the ranking. It reads the customer journey, finds where it is losing the most value, writes the hypothesis, builds the variant inside the customer's own components behind a flag they already have, and waits for a person to approve it before anything reaches traffic. The scoring meeting is not what it replaces. The engineering cost that makes a backlog necessary is. What runs while a test is live is in the guardrail guide.

What to change this quarter

Three moves, in order, none of which requires a new tool.

  • Split the backlog by cost, not by score. Two lists: tests that could be live this week if someone had the time, and tests that need a build. The first list is not a backlog and should stop being managed as one.
  • Measure the price of one test. Engineering hours from written hypothesis to live variant. It is the number that decides which model a team can afford, and most teams have never counted it.
  • Refuse the tests that cannot finish before scoring anything. A feasibility check costs minutes and removes items that would otherwise sit near the top of a list for a quarter.

Sources

  1. 1A/B Testing Intuition Busters: Common Misunderstandings in Online Controlled Experiments, Kohavi, Deng and Vermeer (KDD 2022), 2022-08-14. Read 2026-09-04.
  2. 2RICE: Simple prioritization for product managers, Intercom, 2018-01-05. Read 2026-09-04.
  3. 3The RIGHT model for Continuous Experimentation, Journal of Systems and Software 123 (Aalto University research portal), 2017-01-01. Read 2026-09-04.
  4. 4Democratizing online controlled experiments at Booking.com, arXiv, 2017-10-23. Read 2026-09-04.

Give us one funnel.

Twenty five minutes, about how you run experiments today. No access, no commitment.