logo
|
Blog
  • DelightRoom
  • Alarmy
  • DARO
  • DelightHub
  • KOEN
Careers
Product

Time is Undefeated

The experiment group that won last year lost this time.
DelightRoom's avatar
DelightRoom
Sep 29, 2024
Time is Undefeated
Contents
The experiment group that won last year lost this time.Deciding to run the experiment againBackground: Same D1 retention, different meaningRe-experiment: Same variable, different situationCore metric and guardrail metricA different result from beforeWhat else changed?

The experiment group that won last year lost this time.

We've run countless experiments on when to show the purchase screen during the new user onboarding process. The hottest(?) variable was exactly when to show it. We couldn't just look at a single core metric; we had to monitor guardrail metrics that might take a negative hit. After fierce debate, the current version was established and maintained as our baseline. At the time, we had found the optimal timing.

(Past related post 1: The importance of driving the first subscription while considering retention — Optimizing the timing of the onboarding purchase screen)

(Past related post 2: The subscription and retention balancing game — 2x subscription conversion improvement vs. 10% D1 retention improvement)

Deciding to run the experiment again

About nine months passed, and we decided to revisit the 'optimal' timing for showing the onboarding purchase screen. Watching funnel data trends inside the product and reviewing user observation cameras made us increasingly feel that the current timing was no longer appropriate.

(Left) Last year's experiment (Right) This experiment

Background: Same D1 retention, different meaning

We had definitely chosen the experiment group that increased the subscription conversion rate without dropping D1 retention through an A/B test. However, as time went by, the fact that the control and experiment groups had the same D1 retention began to mean something entirely different.

First off, even if the same proportion of users dropped off, exactly when they dropped off was different. Unlike the previous purchase screen timing, the current timing was disadvantageous for product growth in multiple ways. They might have been users who would have dropped off later anyway, but not even getting the chance to introduce and pitch our product became a huge bottleneck. This constraint was particularly tough this year, as Alarmy was accelerating its value expansion from just waking up to full sleep management.

Secondly, even if the same percentage of users churned, the impression they left with was completely different. This was something hard to catch with quantitative data alone; we only realized it after running several user observation camera sessions. Users who finished the entire new user onboarding and then churned left with an impression like, 'This service doesn't seem to solve my current problem.' On the other hand, users who dropped off because they hit the purchase screen early in the onboarding left thinking, 'Oh, this is a completely paid app.' We concluded that the negative impact of the latter was distinctly worse.

Even if the final number of remaining users is the same, what they think as they drop off can be completely different.

We felt that the former users might eventually come back to us someday, whereas the latter users would never look our way again. I also worried they might spread negative word-of-mouth to their friends. 'I didn't use it because it wasn't a good fit for me, but it might work for you, so give it a try' and 'That app is totally paid, I think you have to pay to use it' would clearly drive viral effects in completely opposite directions.

Re-experiment: Same variable, different situation

So we decided to re-run the exact same experiment from before. The variable was identical: 'exposure timing.' However, our product had changed in the meantime. Due to our expansion into sleep services, sleep-related content had been added to the onboarding flow. For the experiment group that saw the purchase screen later, they would now be exposed to more 'sleep-related content' compared to the last experiment. On top of that, our App Store listing was also quite different from before. It was our first time re-running the exact same experiment, so I was incredibly curious about the results.

Core metric and guardrail metric

This time, we flipped it: we set our core metric as D1 retention and our guardrail metric as the subscription conversion rate. Since we now clearly understood the trade-off relationship between the two metrics, we kicked off the work with a specific, acceptable drop threshold set for the guardrail.

A different result from before

Surprisingly, the result (the winner) was different.

Over nine months, the market, the users, and our product had all evolved, and the results seemed to reflect those very changes.

Perhaps the onboarding value props presented before the delayed purchase screen were more persuasive than before. Last time, showing the purchase screen later resulted in lower subscription conversion with no significant difference in D1 retention. This time, however, D1 retention came out higher with no significant difference in subscription conversion.

To our absolute relief, we were finally able to change the onboarding purchase screen timing that had been acting as a major bottleneck. Now, each cross-functional squad can freely use onboarding as a blank canvas to pitch our product's value to their hearts' content. For context, we did see a slight drop in subscription conversion in certain countries. But just as this experiment lost in the past but won this time, I figure it's a drop we can easily recover by solidifying our value proposition during onboarding. We plan to patch it up with various follow-up growth initiatives.

What else changed?

Running the same experiment but getting a different result gave us some multi-dimensional lessons. But aside from the product outcome, there were other things that had changed. Nine months ago, we were a bit scattered about whose side to take when D1 retention and subscription conversion clashed. The process of selecting a winner was messy. But this time, from planning to determining the winner, we were able to make clean decisions without any real confusion. We rarely have disagreements now because everyone shares a similar understanding of what an acceptable trade-off threshold looks like. Even if opinions differ slightly, we all know we can reach a decision through a quick discussion. Weighing these two core metrics against each other has simply become our daily routine. Alongside our product, our team had grown too.

We've been steadily stacking up our past growth lessons and smoothly rolling out follow-up plans, but we hadn't really looked back at variables we had already validated once. We also hadn't thought much about reviewing qualitative data collected through things like user observation cameras alongside the quantitative data. As the market and users change, we shouldn't blindly trust our previously confirmed winners. We need to build a habit of periodically questioning our product too.

Please look forward to what other revisit cases we'll share on the blog next!

(Ah, I should also write a post sharing some lessons about the observation cameras sometime.)

Share article
Contents
The experiment group that won last year lost this time.Deciding to run the experiment againBackground: Same D1 retention, different meaningRe-experiment: Same variable, different situationCore metric and guardrail metricA different result from beforeWhat else changed?

Delightroom

RSS·Powered by Inblog