One Button Label, Tested on a Six-Figure Sample
The main call to action on a live payment form did not describe what actually happened after the click. I designed the alternative and ran it as an A/B test on real payment traffic. Every monitored metric grew and the change shipped.
Context
The product is a subscription-based B2C/B2B SaaS service with a four-team billing department. This is work from the customer-facing payment form – the last screen before money moves.
Problem
The primary call to action on the payment form was labelled “Place the order”.
That label did not describe what actually happened when a user pressed it. At the final step of a payment funnel, the gap between what a button promises and what it does is not a copy problem. It is a conversion problem, and it sits at the point where hesitation is most expensive.
My Role
Product Designer. I identified the mismatch, designed the alternative labelling, and put it through a controlled A/B test on live payment traffic.
Decision
I treated a wording change as something to prove rather than something to argue.
Button copy is normally settled in a review meeting, by whoever argues best. That works until the change sits on a live payment funnel, where being wrong costs revenue directly and nobody finds out for a quarter.
The fix itself was small – one label. Which means the entire value of the work was in measuring it properly, at a scale where the result could not be dismissed as noise.
The hypothesis came from user expectation, not from taste. The existing label did not reflect the meaning of what would happen after the click. Framing it that way is what made the result interpretable when the metrics moved: if the numbers went up, they went up for a stated reason.
Results
The experiment ran for roughly five weeks in early 2023, across a six-figure user sample – over 164,000 users. The experimental version won, all metrics showed significant growth, and it was deployed to production.
| Metric | Delta |
|---|---|
| First Payments | +5.4% |
| Paid Users | +4.78% |
| Trials | +4.46% |
| Payment form successful submit | +3.67% |
Every monitored metric moved in the same direction. That is what makes the effect credible – a single lucky metric is an anecdote, four metrics moving together is a result.
Why This Is the Number I Lead With
I have a larger percentage in my portfolio. A later experiment on a different payment surface produced +52% on first payments, and I deliberately do not lead with it, because it rests on 13 conversions against 21 and the experiment’s own summary describes the outcome as similar conversion rates. That case study is about closing a legacy payment form, and its real achievement is elsewhere.
This one is the opposite trade: a smaller effect on roughly a hundred times the sample, with an unambiguous conclusion and nothing to walk back if someone asks how it was measured.
A +5.4% improvement on a six-figure sample is a stronger claim in front of anyone competent than +52% on 13 conversions. Knowing which of your own numbers to put forward is part of the job.
Key Takeaways
- Interface wording is a revenue variable on a payment surface, not a style choice. It deserves the same evidence as a structural change.
- Scale is what converts an opinion into evidence. The same edit shipped on taste would have been unfalsifiable.
- State the hypothesis before the test, in terms of user expectation. Otherwise a moved metric tells you nothing about why it moved.
- Report the number that survives the follow-up question.