1. Success, defined first

A target lift on the primary metric, sized honestly: frequency lift for casual and first-time learners, reactivation rate for lapsed. Small and compounding. Current-user retention sits at 84% and moves roughly a point a year; a large number on a slide would be an own-goal.

Pairing rate is measured first. Per-pair value is hypothesised high. Pairing rate is the thing nobody knows.

2. Measurement

The showcase page states the metric and the guardrails. The design behind them:

Per person, not per pair. Pairs are asynchronous by design, and averaging across a pair would hide one buddy showing up daily while the other shows up twice a month. Days active per buddy per month, against unpaired learners matched on segment and baseline activity.

Individual-level randomisation where the feature allows it. Paired learners self-select, and engaged users are naturally more social. Where randomisation isn't possible, the comparison uses matched cohorts, and the reactivation read in particular uses a holdout: lapsed learners who were invited and couldn't accept, against lapsed learners who were invited and could.

Isolating the pull. Activity in a week with no quiz exchanged can't be explained by the quiz. That is the read on whether knowing where your buddy is does anything on its own.

Guardrails and what each protects against. Solo lessons completed, against Course Buddies cannibalising individual learning. 30-day retention, against a short-term social bump masking later drop-off. Invite behaviour, against spam pressure eroding trust.

3. When to stop

Kill when iteration stops moving the primary metric: the meaningful design variations have been tried and the delta has flattened below the bar with no remaining variation that plausibly clears it. Not one failed test. The levers below are that list of variations, written before launch, so that "we tried the reasonable things" means something specific.

Scope and kill are different signals. If guardrails are clean and one segment converts but the overall number misses, that is a narrow-the-scope signal. Kill is when even the best-responding segment can't clear the bar, or a guardrail can't be protected without gutting the feature.

A clean kill produces a real learning: interpersonal commitment doesn't convert this population as hypothesised. That feeds the next bet.

4. Levers

The showcase page names three: no material reward, choosing from four cards, and review clearing at node level. Four more would be pulled before those, or alongside them, depending on what the first read shows.

A shared daily quest. The call is no shared target; the quiz exchange is the daily driver. The alternative is a daily goal both buddies' activity feeds, complete so many lessons or earn so much XP between you, with a reward for both. This is different from the rejected idea of putting Course Buddies on the Quest page; it's a shared target inside the feature's own surface. It is the most likely thing to try if DAU underperforms and the strongest counter to the intrinsic bet, because it changes what a day with a buddy is for. The signal is daily return among paired learners flat against solo.

The sender answering before sending. The call is that the sender answers all five before the quiz goes. This is the learn-by-teaching mechanism, and it is also friction on every send. Removing it halves the cost of sending and removes the learning claim. It stays unless send abandonment mid-compose is high and quizzes per pair would rise materially without it. If that trade has to be made, the feature becomes an engagement mechanic that happens to involve practice, and the doc should say so rather than keep the claim.

Public buddy status on profiles. The call is that a pair is visible only to the two people in it; a third party sees nothing. The alternative is showing active pairs on a profile the way Courses and Achievements are shown. Seeing that a friend has a buddy is one of the few ways anyone discovers the feature, and it's the trigger for the out-of-app conversation pairing depends on. V1 doesn't hide pairs; it chooses when they're visible, at a milestone post rather than continuously, because Duolingo announces relationships as events and never displays them as standing status. A permanent line is checkable in a way a post is not. If pairing rate is starved for discovery, this is the lever, and it costs a visibility change and no new mechanic. The signal is pairing rate against awareness, and whether pairs overwhelmingly form from a milestone post.

Default position during review. The call is that a rewound buddy's Path opens at their review point while a range is outstanding, not at their own solo progress. The alternative is opening at the solo position with the range reached by scrolling. This determines whether rewound learners actually work their range or drift back to solo progression and never converge. The signal is range completion rate, and the share of sessions that touch the range at all.

The remaining levers, one line each: