How do you run Google Play store listing experiments for a work app?
Google Play lets you test listing changes on live traffic and give different audiences their own listing. For a work app this answers practical questions with evidence: does a scanned-invoice screenshot beat a folder view, does a French listing help in Quebec, should lapsed users see what changed? This guide covers the settings, a worked test for a fictional scanning app and how Play compares with Apple.
How do Google Play store listing experiments work?
You create up to two variants of your listing and Play shows them, alongside the current listing, to a share of visitors you choose. Per Play Console Help on store listing experiments, you can test graphics or localized text in up to five languages, measure first-time installers or retained installers, set the audience percentage, a minimum detectable effect and a confidence level, and Google stops an experiment after six months.
When the result is clear, you apply the winning variant to the main listing. Nothing about the app changes; only what visitors see before installing.
Scan documents with your phone and save them as PDFs.
Scan invoices, receipts and signed forms into clean, searchable PDFs.
When is a work app ready to run Play experiments?
You need enough listing traffic for a result within weeks, not months. A niche app with a few dozen store visitors a day can still test, but only large differences will show up, so pick bold variants and accept a higher minimum detectable effect.

Experiments fit best after the basics are sound: a clear title and short description that follow the Google Play listing rules, a complete screenshot set and a working first-run experience. Testing a listing that sends people into a broken onboarding only measures the onboarding.
| Stage | What to test | What not to test yet |
|---|---|---|
| Early, low traffic | Two very different first screenshots | Icon colour, caption wording |
| Growing, steady search traffic | Short description promise, feature graphic | Five changes at once |
| Several markets | Localized text in French or Spanish | Untranslated screenshots |
| Mature listing | Icon, order of frames, preview video | Changes without a hypothesis |
How do you set up a store listing experiment?
| Setting | Choice | Why |
|---|---|---|
| Type | Default graphics experiment | Testing screenshots and feature graphic |
| Variants | 1 (Variant A) | One variant reaches a result faster than two |
| Audience | 50% of visitors | Half see current, half see the variant |
| Metric | Retained first-time installers | Counts installs that stay |
| Minimum detectable effect | Set larger for low traffic | Small effects need very large samples |
| Confidence level | Keep Google's default unless you know why to change it | Higher confidence takes longer |
| End | When a result is reported, or stop by month six | Play ends experiments after six months |
Write down the start date, traffic share and hypothesis in a log. If you run a promotion or ad push during the test, note it, because a sudden change in traffic mix can bias the result.
What does a worked experiment look like?
ScanForm's team believed office users were the paying group, but the listing showed a generic folder view. The hypothesis: leading with a scanned invoice would attract more of the users who pay.
| Current listing | Variant A | |
|---|---|---|
| First screenshot | Your scans, in folders | Invoice scanned, edges fixed |
| Short description | Scan documents with your phone and save them as PDFs. | Scan invoices, receipts and signed forms into clean, searchable PDFs. |
| Feature graphic | Logo on a gradient | From paper to a clean PDF |
| Who it speaks to | Anyone with paper | Office staff and bookkeepers |
Illustrative outcome: the variant wins on retained installers, and the team checks one more thing before applying it, whether new users reach a first saved PDF at the same or a higher rate. Only then does it become the default. Illustrative example written by ShoutEx for this guide, not a benchmark.
Test big ideas, not pixels.
Small listing tweaks need huge traffic to show a result. Test a different first screenshot, a different promise or a different audience, and keep everything else constant.
When should you use custom store listings instead?
Experiments find one better listing for everyone. Custom store listings give specific audiences their own listing permanently. Play Console Help allows up to 50 per app, targeted by country, pre-registration, search terms, Google Ads traffic, user behaviour such as churned or lapsed users, or custom audiences, each with its own URL.
Welcome back. Batch scanning, auto-naming and shared folders are new.
| Audience | Targeting | Listing focus |
|---|---|---|
| Quebec | Country plus French localization | French captions, local document types |
| Searches for "receipt scanner" | Search terms | Receipts first, expense export |
| Google Ads traffic | Google Ads | Matches the ad's promise; see Google App campaigns |
| Churned users | User behaviour | What is new since they left |
| Partner members | Unique URL | Names the partner's workflow |
How do Play experiments compare with Apple product page optimization?
| Google Play experiments | Apple product page optimization | |
|---|---|---|
| Variants | Up to 2 plus current | Up to 3 treatments plus original |
| What can change | Graphics or localized text | Icon, screenshots, app previews |
| Length | Stopped after 6 months | Up to 90 days |
| Metric | Installs or retained installers | Conversion rate reported by Apple |
| Audience versions | Up to 50 custom store listings | Up to 70 custom product pages |
Apple's details are on its product page optimization page, and the iOS side of audience pages is covered in product page tests. ShoutEx view: run the same hypothesis on both stores when you can. If a business-document first frame wins on Android and loses on iOS, that tells you something useful about who uses each platform.
What mistakes and metrics matter for Play experiments?
- Stopping early because one variant leads after three days.
- Measuring installs only when retained installers are available.
- Testing during a campaign spike that changes who visits.
- Applying a winner without checking activation in your own data.
- Leaving custom listings stale after a redesign.
| Metric | Question it answers | Where to find it |
|---|---|---|
| Retained first-time installers by variant | Which listing brings users who stay? | Play Console, experiment results |
| Conversion by custom listing | Does each audience listing persuade? | Play Console, store listing performance |
| First saved PDF by variant | Did the winner bring the right users? | Your analytics |
| Paid conversion by listing | Did it bring paying users? | Your billing, joined by listing URL |
Frequently asked questions
How many variants can a Google Play experiment have?
Up to two variants tested against the current listing.
What can I test in a Play store listing experiment?
Graphics such as the icon, screenshots, feature graphic and video, or localized text in up to five languages.
How long can a Google Play experiment run?
Google stops experiments after six months. Most tests should reach a result much sooner if traffic is adequate.
What metric should I use for Play experiments?
Retained first-time installers when available, because it counts installs from people who keep the app.
How many custom store listings can I create?
Up to 50, each with its own URL and targeting such as country, search terms, Google Ads traffic or lapsed users.
Is there an Apple equivalent to Play experiments?
Yes. Apple product page optimization tests up to three treatments for up to 90 days, and Apple custom product pages give audiences their own pages.
Can a small app run Play experiments?
Yes, but only large differences will show up with little traffic. Test bold changes and expect results to take longer.
Sources & further reading
Regulator rules, platform policies and local data change. These sources let you check the facts on this page, last checked October 6, 2026.