Alarmy's A/B Testing Journal #1
1
While running A/B tests, I often wondered what hypotheses other teams were testing and what results they were getting. So, I wanted to take this opportunity to share some of the A/B testing experiences we've had at Alarmy.
App Store Listing A/B Testing
What should we write in our app store listing to get more downloads? What images should we use on the store page to encourage more users to download it?
On the Play Store, you can easily answer these questions through store listing experiments. I've picked out a few tests we ran at Alarmy related to this and summarized the hypotheses and data.
Feature Graphic
Target: Japan
Hypothesis: Emphasizing our slogan (SLEEP IF U CAN) will increase the download rate.
The result was a +6.2% to +26.4% increase in downloads compared to the original, and surprisingly, localizing the slogan actually yielded worse results.
Short Description
Target: Korea
Hypothesis: Specifying our target audience will make people think, "This is exactly me!" and increase the download rate.
The result showed a +0.6% to +20.1% improvement over the original. After testing various target audience definitions, we found the right audience for Alarmy was "people who can't wake up well with regular alarms" (which seems obvious in hindsight).
Full Description
Target: Germany
Hypothesis: Introducing the photo dismiss mode in the very first line will intrigue users and increase the download rate.
The original first line of the description stated that we were featured in major foreign media like CNET and Gizmodo, but we changed this to a sentence introducing the photo dismiss feature. Since the 'full description' section is only visible after pressing the read more button, we honestly didn't expect much before the experiment. However, the result was a +7.0% to +26.6% boost, which was a much bigger improvement than we anticipated.
Screenshots
Target: US
Hypothesis: Changing the toilet image to a bathroom sink will increase downloads.
I think this is one of the most interesting hypotheses among the A/B tests we ran. A friend saw the screenshot with the toilet picture and said it felt somewhat gross, noting especially that women don't usually leave the toilet seat up like that (something I hadn't even thought of!). They recommended changing it to a picture of a sink.
The screenshot above doesn't appear on the first store screen; you have to swipe to the second one to see it. For that reason, I didn't think simply swapping the registered photo would lead to any improvement, but the result was a +0.9% to +6.1% lift.
So What?
Thanks to running A/B tests on various hypotheses across many countries like this, Alarmy's store acquisition rate remains around 50%. This means that out of every two people who view our store listing, one definitely downloads Alarmy.
In the image above, you can see benchmarking results comparing the acquisition rates of similar apps. The 75th percentile is 34%, which shows Alarmy is maintaining a pretty solid acquisition rate (51.2%). Of course, we could increase this rate by running CPI (Cost Per Install) marketing campaigns, but Alarmy is currently not doing any marketing at all.
Ultimately, if you set up various hypotheses and consistently run store listing A/B tests, you can improve your store acquisition rate, so I highly recommend giving it a try.
Conversion Copy A/B Testing
I want to introduce one instance among our in-app copy A/B tests that yielded an unexpected result. Although it's gone now, we used to run a rewarded ad system within the Alarmy app like the one below. It was a feature where downloading one of the suggested apps (ad apps) would upgrade you to the Pro version for 30 days.
The reward users felt here was '30 days of Pro version usage'. What if the reward increased? Wouldn't they naturally download more suggested apps? Based on this hypothesis, we ran an A/B test changing the number of reward days. We tested 30-day, 60-day, and 100-day rewards, aiming to see how much the conversion rate would increase as the reward got bigger. Here were the results.
30-day reward: 6% | 60-day reward: 14% | 100-day reward: 10%
Our expectation that the largest reward of 100 days would have the highest conversion rate was spectacularly wrong. Thinking, "Why on earth?" we conducted user interviews and found that the 100-day reward felt too generous, making it seem less scarce compared to the 60-day offer (simply put, the "too good to be true" effect).
In the end, we improved the conversion rate by 1.5 times just by changing the reward from 30 days to 60 days (of course, we also factored in the scenario where users receiving the 30-day reward might install an additional ad app within 60 days). Being able to boost the conversion rate by over 1.5x by changing a single number like this is exactly the charm of A/B testing.
To add a bit more, having a list of 3 suggested apps was more efficient than 5 (the paradox of choice where too many options lower purchasing power), and placing the one with the highest unit price first with a 'Like' badge also helped increase efficiency.
Things to Watch Out For
Limited Resources
In truth, doing A/B testing frequently is not easy for a small startup. It would be great if we had overflowing resources to run A/B tests indefinitely, but ultimately, resources are limited, so you have to consider the potential impact an A/B test will bring before proceeding.
If you run A/B tests with the mindset of simply "Let's try changing the button color to every color of the rainbow!" or because the team can't agree and you want to "See who is right!", it's hard to get good results.
Ultimately, I believe you can achieve much better results if you first formulate a highly plausible hypothesis and think about whether that A/B test is worth running within your limited resources (if you understand your users well, your hypotheses will hit the mark more often than you'd expect).
Comparing A/B Tests by Time Period
What if you ran an A/B test by serving Option A in the first week and Option B in the second week? You might make the wrong choice due to external factors. Because visitor behaviors can change from moment to moment based on outside influences, it's a basic rule to run A/B tests during the exact same time period.
Let's take an example. The graph below shows Alarmy's weekly store acquisition in India. The period on the right has a 16% higher acquisition rate, but in reality, we didn't make any changes between the left and right periods (the shift occurred purely due to external factors).
Like this, significant differences can arise due to external factors even if you do absolutely nothing. If you run a time-based A/B test in a situation like this, what happens? If a worse option was introduced during the right-side period, you might end up making a decision that actually makes things worse than before.
Adequate Sample Size
It goes without saying, but you need an adequate sample size to run an A/B test. If the sample size is small, it's hard to make proper decisions, so it's helpful to look at results using a statistical significance calculator like the one below.
When looking at A/B testing case studies, it sometimes looks like you can achieve massive improvements with very little effort. But in experience, you mostly only salvage a few wins out of many tests, and dramatic improvements like the ones you hear about from the outside rarely happen.
Ultimately, in a resource-strapped startup, it can be more effective to focus on the pain points or improvements users are already experiencing rather than obsessing over a few percent improvement. Therefore, as I mentioned above, rather than blindly running A/B tests, I recommend formulating good hypotheses and proceeding while considering the potential impact of the A/B test.