How to Test Loyalty Program Offers: Holdouts, Variables and Metrics
|
How this guide was prepared. Last updated October 2026. It draws on Brandmovers' experience running loyalty programs and offers, including the distributor program cited below, described as its case page reports it. The rule on "free" offers was checked against the FTC's guide. Examples are illustrations, not benchmarks. This is general information, not legal advice. |
A loyalty offer test compares how members respond to different versions of an offer, and to no offer at all, by randomly assigning eligible members to each group. The no-offer holdout shows whether an offer changed behavior or only rewarded purchases members would have made anyway.
Bonus points, discounts and other offers are among the most used tools in a loyalty program, and among the most expensive. An offer that looks successful because many members redeemed it may have cost margin on purchases that would have happened without it. Testing offers properly answers a narrower and more useful question: which offer changes what members do, by how much, and at what cost. This guide covers why to test, what to test, how to design a valid test, what to measure, how to read results and roll out, and the rules that apply to the offers themselves.
Key Takeaways
|
Why test loyalty program offers?
Offers cost margin, and an untested offer can reward purchases members would have made anyway; testing shows which offers change behavior at a cost the program can sustain.
Without a test, the usual evidence is the redemption count: how many members used the offer. That number cannot separate members who bought because of the offer from members who would have bought regardless. This matters most for lapsing members, some of whom return on their own, so a rise after an offer can look like the offer's effect when nothing is there to compare it against. Frequent offers, tested or not, can also teach members to wait for a deal; a long-running holdout and the full-price purchasing measure below can show whether that is happening. A testing habit turns each offer into evidence for the next one.
What can you test in a loyalty offer?
Test one variable at a time, such as urgency, frequency, value and type, spend threshold, cross-category incentive, targeting or channel, against a holdout.
|
Variable |
Example A vs B |
What it answers |
Main metric |
Watch out for |
|---|---|---|---|---|
|
Urgency |
24-hour vs 72-hour redemption window |
Whether a shorter window lifts response |
Incremental purchases during and after the window |
Pulling purchases forward; overuse |
|
Frequency |
Monthly vs quarterly offers |
How often members can receive offers before response falls |
Incremental purchases per member over the period |
Fatigue and unsubscribes |
|
Value and type |
20% off vs $20 cashback; bonus points vs discount |
Which form motivates at the lowest cost |
Incremental margin per member |
Set both versions at a similar cost per member, or the test compares value as well as form; compare net of cost |
|
Spend threshold |
Bonus for spending $50 vs $75 |
Where a threshold moves basket size |
Basket size against the holdout |
Members already above the threshold |
|
Cross-category |
Bonus points on a related category |
Whether members try a new category |
New-category purchase rate |
Shifting spend from one category to another |
|
Targeting |
Offer to lapsing members vs all members |
Who needs the offer at all |
Incremental response by segment |
Paying members who were already active |
|
Channel and message |
Email vs app; different framing |
How best to deliver the offer |
Response by channel |
Members differ by channel |
The guides to redemption rates, average order value and personalized offers for churn cover the goals these tests usually serve.
How do you design a valid loyalty offer test?
Write a hypothesis, randomly assign eligible members, include a no-offer holdout, change one variable, run the test across a buying cycle and decide the success measure before launch.
- Hypothesis: state what the offer should change and for whom, such as "a 72-hour bonus raises second purchases among members who have not bought in 60 days."
- Random assignment: assign eligible members to groups at random, so that apart from chance the groups differ only in the offer.
- No-offer holdout: keep a group that receives no offer, so you can see what would have happened anyway.
- One change at a time: if two things change between groups, you cannot tell which one caused the result, unless the test has a group for each combination.
- Length and size: set the length in advance and run the test across at least one normal buying cycle. Before launch, use a sample size calculator: enter the current purchase rate, the smallest lift worth acting on and the confidence level you want, and it returns the members needed in each group. A smaller holdout costs less in withheld offers but needs a larger lift to show clearly. Stopping early because a result looks good makes chance results look real.
- No overlap or spillover: keep members in one test at a time, or a second offer will blur the first. Offers shared on deal sites, within a household or across buyers at one business account can reach the holdout and shrink the measured difference.
- Success measure in advance: decide the metric and the smallest result worth acting on before the test starts.
Smaller programs. If the member base cannot fill several groups, test fewer versions at once, usually one offer against a holdout, make the difference between versions larger or run the test longer. A result from groups too small to separate signal from noise is a direction, not a decision.
B2B programs. In a channel program, assign whole accounts rather than individual buyers, since several people at one distributor or dealer may buy through the same account. Account counts are often small and buying cycles long, so the guidance for smaller programs usually applies.
BLOYL™, Brandmovers' enterprise loyalty platform, supports A/B testing against a control group, and its rules engine sets earning and redemption rules by customer segment, purchase channel, time window or behavioral action, so tests and holdouts can be set up inside the program.
Case study (disclosed by Brandmovers). A leading Canadian regional distributor's B2B loyalty program, run on BENGAGED™, Brandmovers' B2B loyalty platform, shows why random assignment matters. "Sales among enrolled customers grew by an average of 25%, while non-enrolled customers saw only a 5% average increase" (distributor case study). Customers were not randomly assigned to enroll, so the gap reflects both the program and differences between the customers who joined and those who did not. Random assignment avoids that problem in an offer test, because members who get the offer and members who do not are chosen by chance rather than by choice.
What should you measure in a loyalty offer test?
Measure incremental purchases, spend and margin against the holdout, net of the offer's cost, and check what happens after the offer ends; redemptions only show use.
- Incremental response: purchases per member in each offer group minus purchases per member in the holdout over the same period, counting every member assigned to a group, not only those who opened or redeemed the offer.
- Incremental margin: the extra margin from those purchases minus the cost of the offer, including the cost of rewards paid on purchases the holdout shows would have happened anyway.
- After the offer: purchases in the weeks after the offer ends, to see whether the offer created new purchases or only moved them earlier.
- Side effects: unsubscribes, complaints and changes in full-price purchasing.
- Longer-term behavior: retention and visit frequency for members in each group over the following months.
Illustrative example. A program tests a 72-hour bonus against a 24-hour bonus for lapsing members, with a no-offer holdout. The 24-hour group redeems more often, but over the following eight weeks its purchases fall back to the holdout's level, while the 72-hour group stays ahead of the holdout by more than normal variation. On redemptions alone the 24-hour offer wins; on incremental margin net of each offer's cost, including after the offer ends, the 72-hour offer does. This is an illustration, not a benchmark.
How do you read results and roll out a winning offer?
Act on a result only when the difference exceeds normal variation, holds after the offer ends and pays after its cost; then roll out with a holdout in place.
- Check the size: small differences between groups can come from chance, and the more versions a test compares, the more likely one looks like a winner by chance alone. Repeat the test or extend it if the result is close.
- Check the economics: a higher response rate is not a win if the offer's cost exceeds the extra margin.
- Roll out gradually: extend the offer to similar segments first, and keep a small holdout so the effect can still be measured. A holdout gives up the offer's gains for those members, so size it to what the measurement needs.
- Retest: member behavior and competing offers change, and a new type of offer can draw a response partly because it is new, so results go stale.
The guide to A/B testing a loyalty program covers testing at the program level, such as pilots before launch.
What rules apply to loyalty program offers?
State each offer's terms, conditions and expiry clearly, follow the FTC's guide for offers described as "free," and keep every test group's terms clear and consistent.
"Free" offers. The FTC's guide says the conditions of a "free" offer "should be set forth clearly and conspicuously at the outset of the offer," and that a "free" offer is understood to be "based upon a regular price for the merchandise or service which must be purchased" (16 CFR 251.1). A bonus item should not be paid for by raising the price of the item that must be bought. The guide also says a single size of a product "should not be advertised with a 'Free' offer in a trade area for more than 6 months in any 12-month period," which matters when testing how often to run "free" offers.
Terms and expiry. Each offer should state what members must do, what they receive and when it ends, and honor those terms for every member who qualifies.
Fairness between groups. Members in different test groups may compare offers. Keep differences within what the program terms allow, and avoid tests that would leave one group worse off than the program's published benefits. Holdout members should still receive every standard program benefit; they miss only the offer being tested.
Competing B2B buyers. Testing different offers on resellers that compete with each other needs care. The FTC explains that a seller "charging competing buyers different prices for the same 'commodity' or discriminating in the provision of 'allowances'" may be violating the Robinson-Patman Act, which requires a seller to "treat all competing customers in a proportionately equal manner" (FTC).
This is general information, not legal advice.
Frequently Asked Questions
-
Randomly assign eligible members to groups, give each group a different version of the offer, keep one group with no offer and change only one variable. Run the test across a normal buying cycle and compare incremental purchases and margin against the no-offer group.
-
Without a group that receives no offer, you cannot tell whether members bought because of the offer or would have bought anyway. The holdout shows the baseline, so the difference between each offer group and the holdout is the offer's effect.
-
Incremental purchases and margin against the holdout, net of the offer's cost, plus what happens after the offer ends. Redemption counts show how many members used the offer, not whether it changed their behavior.
-
At least one normal buying cycle for the product, plus enough time after the offer ends to see whether purchases were created or only moved earlier. Groups also need to be large enough that a difference between them is not just noise.
-
Yes, if every offer is described clearly, stays within the program's published terms and is honored for everyone who qualifies. Avoid tests that would leave one group with less than the program promises. This is general information, not legal advice.
Conclusion
Testing loyalty offers is worth the effort when the test answers the right question: did the offer change what members did, and did it pay? Compare offers against a no-offer holdout, assign members at random, change one variable at a time, measure incremental margin net of cost and what happens after the offer ends, and roll out winners with a holdout still in place.
|
Want to test which loyalty offers actually change behavior? Brandmovers runs loyalty programs on BLOYL, with A/B testing against a control group and a rules engine that sets earning and redemption rules by customer segment, purchase channel, time window or behavioral action. Request a demo to talk through your offer testing with the Brandmovers team. |
Sources


