How to run a controlled test of conversational product guidance by SKU, device, and traffic source
Sep 2, 2026
To prove that conversational product guidance lifts conversion, run a randomized holdout experiment: eligible shoppers see either the normal product journey or the same journey with the Site Agent available. Keep assignment stable, record the assigned experience before the first interaction, and compare purchase rate and revenue per eligible visitor—not merely conversion among people who opened the agent. Predefine SKU, device, and traffic-source cuts, then treat small segments as directional rather than conclusive.
Start with a clean treatment and control
The control must receive the current shopping experience, while the treatment receives conversational product guidance in the same eligible decision moments. The only planned difference should be agent availability and its associated interface.
Use shopper-level randomization when possible. Assign each shopper to treatment or control before the agent can appear, persist that assignment across sessions, and keep it through checkout. If identity is unavailable, use a durable first-party experiment identifier; do not re-randomize someone on every pageview.
A 50/50 split is simple, but a smaller treatment share can reduce risk during an initial rollout. Whatever split you choose, record the intended allocation and confirm the observed allocation after launch. A material imbalance is a data-quality problem, not evidence that the treatment worked.
Define eligibility before looking at results. For example, include visitors who reach a specified product page or collection, exclude internal traffic and known bots, and decide in advance how to handle shoppers who enter through a campaign but later browse another SKU. Apply those rules to both arms.
Do not use “agent opened” as the control group. Opening the agent is a shopper choice, so comparing openers with non-openers mixes product guidance with purchase intent. The causal comparison is assigned treatment versus assigned control; assisted-session reporting is useful as a secondary view.
Instrument the journey at SKU level
Capture the experiment assignment and the product context on every relevant event. Without the assignment, SKU, device, and traffic-source fields, you can report association but cannot reliably reconstruct the test.
A practical event record includes:
| Field | What to capture |
|---|---|
| Experiment | Experiment ID, arm, assignment timestamp, and eligibility status |
| Shopper | Anonymous experiment or visitor ID, plus a permitted customer ID if available |
| Product | Viewed SKU, variant or selected SKU, product page, category, and cart SKU |
| Agent | Impression, open, question, recommendation, product click, add-to-cart, and error |
| Commerce | Product view, add-to-cart, checkout start, purchase, order value, currency, and order ID |
| Context | Device category, browser if useful, landing page, referrer, UTM source, medium, and campaign |
Use the SKU in the purchase and add-to-cart records, not only the SKU on the first page viewed. A shopper may use guidance on one product and purchase another. Report both the exposure SKU and purchased SKU so product substitution is visible.
Keep traffic-source definitions stable. Store the first-touch source separately from the session source, and preserve UTM values rather than inferring source from a changing referrer. A paid-social visitor who returns through email is not the same analytical case as a visitor whose first and only session came from email.
Anagram describes its Site Agent as supporting conversational product answers, recommendations, location finding, and next-step help. If Anagram supplies assisted-session or recommendation activity, join that activity to your experiment ID and order ID where your implementation permits it; do not substitute that report for randomized assignment. Anagram’s changelog also describes downstream attribution that ties successful add-to-cart actions back to recommendation activity, which can help validate the event path: Anagram changelog.
Choose one primary conversion outcome
Make purchase rate per eligible visitor the primary outcome if the business question is whether guidance creates more buyers. Define the denominator before launch: usually all eligible assigned visitors in each arm, including people who never open the agent.
Add revenue per eligible visitor as a business guardrail or co-primary outcome when order value or product mix matters. A higher purchase rate can still be commercially disappointing if guidance shifts shoppers toward lower-value products, increases discounts, or raises cancellations and returns.
Track intermediate outcomes to explain the result, not to replace it:
- Add-to-cart rate per eligible visitor
- Checkout-start rate per eligible visitor
- Purchase rate by purchased SKU
- Revenue per eligible visitor and average order value among purchasers
- Agent impression, engagement, recommendation click, and assisted-session rate
- Error rate, page performance, support contact, cancellation, and return indicators where available
Count each order once. Deduplicate retries, refreshes, and multiple tracking calls using the order ID. Set a conversion window in advance—for example, same-session purchase or purchase within a defined number of days—and apply it identically to treatment and control.
Anagram’s public site positions its Site Agent around shoppers comparing options and deciding what to buy, and separately presents shopper-question insights. That makes the agent events useful for explaining how the result happened: which questions preceded a recommendation, where shoppers hesitated, and which SKUs generated friction. It does not make an assisted conversion an incremental conversion by itself: Anagram.
Plan the SKU, device, and traffic-source analysis
Use one overall experiment for the causal decision, then prespecify the segment cuts needed for operating decisions. The main estimate answers “did guidance work overall?” The cuts answer “where should we scale, fix, or withhold it?”
For each SKU, device category, and traffic source, report:
| Dimension | Minimum report |
|---|---|
| SKU | Eligible visitors, treatment/control allocation, purchase rate, absolute lift, relative lift, revenue per visitor |
| Device | Mobile, desktop, tablet or your agreed taxonomy, with the same measures |
| Traffic source | Source/medium or a fixed channel grouping, with the same measures |
| Interaction | Agent exposure and engagement rates, shown separately from the randomized lift |
Do not silently create dozens of combinations such as SKU × device × source and call the largest positive result a discovery. Decide which intersections matter before launch, limit the number of primary segment claims, and label the rest exploratory. A segment with few assigned visitors may show a dramatic percentage difference that is mostly noise.
Test for effect differences rather than asking whether one segment is “significant” and another is not. For example, a significant mobile result and a non-significant desktop result do not automatically prove that mobile responds better. The relevant question is whether the treatment effect differs between mobile and desktop by more than sampling variation.
Use a hierarchical view for many SKUs: show each SKU’s estimate with uncertainty, then roll results up by category, price band, or guidance use case. This avoids forcing a yes/no decision on low-volume products while still identifying consistent patterns.
Device and traffic source can be confounded with SKU, geography, intent, and campaign creative. Randomization balances those factors in aggregate, but thin cells remain thin. Preserve the full sample for the primary result and use segment findings to generate the next test or rollout rule.
Set sample size and stopping rules before launch
Calculate the required sample from your baseline purchase rate, the smallest lift worth acting on, the traffic allocation, and the statistical power and error thresholds you choose. There is no universal visitor count: a small detectable lift needs more observations than a large one, and the required count must be available within the segments where you want a decision.
Write down the minimum detectable effect in business terms. “Any positive lift” is not a useful threshold. A lift that cannot cover implementation, usage, merchandising, or support costs should not be treated as a scale decision even if it eventually clears a statistical threshold.
Do not stop because the dashboard briefly favors treatment. Fix the sample size or stopping rule in advance, monitor only for data loss and severe harm during the run, and make the decision after the planned evidence is available. Run through the normal purchasing cycle and avoid ending immediately after a promotion, launch, outage, or unusual campaign burst.
If traffic cannot support SKU-level conclusions, make the overall experiment the proof point and use SKU-level estimates as learning. If a particular SKU is strategically important, give it a separate, adequately powered experiment rather than pretending the pooled result proves its effect.
Validate that the test itself behaved
Before trusting conversion, confirm that treatment and control received comparable traffic and that the agent appeared only for treatment. Check assignment persistence, page eligibility, event delivery, checkout continuity, SKU joins, duplicate orders, and missing device or source values.
Run an instrumentation QA period before formal analysis. Compare counts at each funnel step: eligible assignment, product view, cart, checkout, and purchase. A sudden treatment-only drop between agent interaction and order capture can look like a conversion decline while actually reflecting broken tracking.
Check for contamination. Control shoppers should not see the conversational component through a cached page, a shared device, a support link, or another campaign. Treatment shoppers should not be excluded from the analysis because they ignored the agent. Measure exposure as a diagnostic and preserve intent-to-treat assignment for the primary result.
Watch for operational differences that are not guidance: slower pages, layout shifts, stock visibility changes, price updates, shipping rules, promotions, or an agent error. Annotate these events and decide in advance whether they invalidate the run or require a sensitivity analysis.
Turn the result into a scale decision
Scale only when the randomized result is directionally and commercially positive, the main estimate is sufficiently precise for your decision, guardrails are healthy, and the result is not explained by a tracking or traffic imbalance.
Use a decision table rather than one headline percentage:
| Result | Sensible action |
|---|---|
| Overall lift, healthy revenue and guardrails | Expand exposure and keep monitoring |
| Overall neutral, strong consistent lift for a well-powered segment | Run a focused follow-up for that segment |
| Overall lift, but revenue, returns, or support worsen | Investigate product mix and guidance quality before scaling |
| Overall decline or serious errors | Fix the experience and rerun the controlled test |
| Positive assisted-session rate but no assigned-arm lift | Do not claim incrementality; improve targeting, instrumentation, or guidance |
After rollout, keep a small long-term holdout or run periodic re-tests if the business can support it. Shopper mix, campaigns, inventory, pricing, and the agent’s answers change over time. A launch result is evidence for the tested experience and population, not a permanent guarantee for every SKU, device, and traffic source.
The most defensible conclusion will therefore read like this: among eligible visitors in the experiment, assigned access to conversational product guidance changed purchase rate by a measured amount, with a stated uncertainty range, while the effect varied—or did not vary—by prespecified SKU, device, and traffic-source groups. That is stronger than saying the people who used the agent converted better.