33 tools on the map. 8 are wired today, each with a page saying exactly what it reads and what it writes.

Retention experiments

What retention experimentation is, why its window makes it the hardest stage to test, which metrics measure it, and the guardrails a slow test needs.

What experimentation on retention is

Retention is the stage after the funnel ends. Every other stage turns on a decision made once: create the account, finish setup, choose the plan, pay. Retention turns on a decision made again, and again, after everyone has stopped watching. Experiments here change what a returning user meets: what the home opens on, what is shown as new since the last visit, what happens to a workspace that has gone quiet, and what is offered when someone decides to stop.

Why it is the hardest stage to test

Three problems arrive together. The number is a cohort measured over a window, so the last unit enrolled still owes the test that whole window, and a period still running is not a number yet [1]. The effect is usually small, because one surface rarely changes a habit. And it is the effect most likely to drift: treatment effects in online experiments are not always stable over time and can increase or decrease as users learn a change [2]. Kohavi and colleagues recommend two weeks and a look for an effect that quickly diminishes, while reporting that novelty and primacy effects are uncommon in practice [3].

The metrics

The primary metric is a retention rate over a cohort, and its definition has to be settled before the test, because the same data answers differently depending on the definition chosen. Amplitude names three: returning on a specific day, on that day or after it, and inside brackets the team defines, and its worked example reads far higher under the second than under the first [4]. PostHog groups users into cohorts by when they first perform a start event, and lets the period be hours, days, weeks or months [1]. Make the qualifying event mean the product was used rather than opened.

The window, and what it does to feasibility

Microsoft's experimentation platform reports that a typical A/B test there runs for seven days [5]. A retention question cannot. A four week read needs four weeks per unit on top of enrolment, so a month of enrolment is a two month test, and that arithmetic settles most retention experiments before anyone writes one down. Run to the planned duration, because repeated looks at a test in flight have to account for peeking [5]. Size it with the calculator and the sample size guide, price the calendar with the feasibility guide and the feasibility checker, and be ready to decide by judgment and ship under guardrails instead.

The guardrails

A returning surface fails in two ways. It buries what it does not show, so creation of new work belongs beside the retention number rather than under it. And a test running for weeks has weeks in which assignment can drift; a sample ratio mismatch makes the result untrustworthy whatever the primary metric says [5]. Guardrails are the parts of a product that must not degrade even though they will not necessarily improve [5]. Support volume and cancellation inside the window are the ordinary two, read daily while the primary metric is read once; the guardrail guide says why.

The patterns

Opening the workspace home on work in progress is the example written out in full below. The same instinct sits elsewhere on the returning path: a summary of what changed since the last visit, a prompt in the second week rather than the second session, a notification that lands on the thing it is about. Each changes what a return costs, each needs the same long window, and that is why a team runs one at a time rather than four, as the throughput guide argues.

Sources

  1. 1Retention, PostHog Docs. Read 2026-09-04.
  2. 2Novelty and Primacy: A Long-Term Estimator for Online Experiments, arXiv (Sadeghi, Gupta, Gramatovici, Lu, Ai, Zhang), 2021-02-18. Read 2026-09-04.
  3. 3Seven Rules of Thumb for Web Site Experimenters, Kohavi, Deng, Longbotham and Xu, KDD 2014, via exp-platform.com. Read 2026-09-04.
  4. 4Interpret your retention analysis, Amplitude Docs. Read 2026-09-04.
  5. 5Patterns of Trustworthy Experimentation: During-Experiment Stage, Microsoft Research, Experimentation Platform, 2021-01-25. Read 2026-09-04.

Retention experiments

Give us one funnel.

Twenty five minutes, about how you run experiments today. No access, no commitment.