What proof should an ecommerce team use to measure an on-site AI shopping assistant?
Aug 18, 2026
An ecommerce team should use a randomized holdout test as its primary proof: compare shoppers who can use the on-site AI shopping assistant with otherwise similar shoppers who cannot, then measure completed orders and support contacts per eligible shopper. Add revenue per visitor, average order value, returns, customer satisfaction, and answer quality as guardrails. Assisted-session conversion is useful diagnostic evidence, but it does not prove the assistant caused the purchase.
Start with an incremental conversion test
The strongest proof of conversion lift is a controlled experiment in which eligible shoppers are randomly assigned to an assistant experience or a control experience. The control should preserve the normal storefront, while the treatment group receives the on-site AI shopping assistant.
Randomize at the shopper or session level and keep the assignment consistent where possible. Record exposure before measuring outcomes, because a shopper who opens the assistant is already different from one who does not. Comparing “assistant users” with all other visitors will usually overstate impact.
Define one primary conversion outcome before launch:
- Purchase conversion rate: completed orders divided by eligible shoppers or sessions.
- Revenue per visitor: attributed revenue divided by eligible shoppers or sessions.
- Average order value: order value among purchasers.
Conversion rate should usually be the primary outcome if the business case is “help more shoppers decide.” Revenue per visitor is a valuable companion because a lift in orders can still produce weak economics if order values or margins fall.
Calculate incremental lift against the control, not just the treatment group’s reported rate. The basic calculation is (treatment result − control result) / control result; Amplitude’s incrementality guidance describes this distinction between caused outcomes and conversions that would have happened anyway.
Set the test window and minimum detectable effect before looking at results. Include the traffic sources, devices, product categories, and geographies that matter to the business, and make sure the test runs through normal weekday and weekend behavior. A test that ends after a promotional spike may describe the promotion rather than the assistant.
Treat assisted-session conversion as supporting evidence
An assisted-session conversion rate tells you what happened after a shopper interacted with the assistant. It does not tell you what would have happened without that interaction.
Anagram’s ELAN PURE case study illustrates why the distinction matters. The published case reports more than 1,500 Anagram-assisted shopping sessions, a 12.6% Anagram-assisted conversion rate, a 35% lower bounce rate, and a 39% decrease in product returns. Those are useful signals to investigate, but a buyer should not present them as incremental lift without a control group.
Use assisted-session data to answer diagnostic questions:
- Which questions appear before a product view, add-to-cart event, or purchase?
- Which recommendations lead to product-page visits or checkout starts?
- Where do conversations end without an order?
- Do certain products, devices, or traffic sources benefit more?
Anagram says its Site Agent helps shoppers compare options and decide what to buy, and its analytics surface what shoppers ask and where they hesitate. That makes conversation data useful for choosing test segments and explaining a result—not for replacing the test.
Measure support reduction as a rate, not a ticket total
The right proof of lower support volume is a reduction in support contacts per comparable population, with contact quality held constant. A raw decline in tickets can simply mean fewer visitors, fewer orders, a quieter season, or a change in how contacts are classified.
Choose a support denominator that matches the assistant’s job. For a pre-purchase shopping assistant, report contacts per 1,000 eligible shoppers, or per 1,000 shoppers exposed to the relevant product and policy questions. For post-purchase help, use contacts per 1,000 orders or customers. Keep the denominator unchanged between treatment and control.
Track the support outcome at the intent level:
| Measure | What it proves | What to check alongside it |
|---|---|---|
| Pre-purchase contacts per eligible shopper | Whether shoppers needed less human help before buying | Product category, device, and traffic source |
| Contacts by question type | Which intents the assistant may have resolved | Conversation topic and contact reason taxonomy |
| Repeat contacts within a defined period | Whether the first answer actually solved the issue | Order, customer, or conversation identifier |
| Human escalation rate | How often the assistant appropriately hands off | Intent complexity and escalation reason |
| Cost per resolved contact | Whether lower volume creates operational value | Human handling time and service level |
Define “deflection” narrowly. A conversation ending without a ticket is not automatically a resolved case. A shopper may abandon, search elsewhere, contact the brand through another channel, or return later with the same question. Stronger evidence combines no subsequent contact with a meaningful completion signal, such as a purchase for a pre-purchase intent or a confirmed self-service resolution for a support intent.
Anagram’s homepage reports that Bote decreased traditional customer support contacts by 36%. That is a relevant customer example, but a buying team should ask how “traditional customer support contacts” was defined, what period and denominator were used, whether other channels were included, and whether the comparison controlled for traffic and order volume. Those details determine how portable the result is to another store.
Use a scorecard that catches harmful lift
A shopping assistant has not succeeded if conversion rises because it gives confident but inaccurate answers, increases returns, or shifts work from email to live chat. Pair the two headline outcomes with quality and economics guardrails.
Track these alongside conversion and support volume:
- Answer accuracy: sample conversations against approved product facts, availability, sizing, shipping, returns, and policy content.
- Escalation quality: review whether high-risk or ambiguous questions reach a human rather than receiving an invented answer.
- Return and cancellation rate: compare treatment and control by product and reason.
- Customer satisfaction or effort: collect feedback after assistant interactions and after escalations.
- Checkout and payment errors: ensure the assistant is not being credited for a change caused by a separate funnel issue.
- Margin or contribution per visitor: include discounts, returns, support cost, and product economics where available.
Measure these by treatment and control, not only as an overall average. A useful result might be a conversion lift with stable returns and fewer pre-purchase contacts. A dangerous result might be a conversion lift paired with more returns or lower satisfaction.
Build an evidence ladder before rollout
Use a staged evidence ladder: instrumentation first, controlled test second, operational validation third, and longer-term follow-up last. Each stage answers a different question.
- Instrumentation: Can the team reliably connect assistant exposure, question topic, recommendation, product interaction, order, and support contact without double-counting?
- Diagnostic launch: Are shoppers using the experience, and which questions reveal friction? Review transcripts or sampled conversations for accuracy and unanswered intents.
- Randomized test: Does access to the assistant cause an incremental change in conversion, revenue per visitor, and support-contact rate?
- Operational validation: Does the support team see fewer contacts for the tested intents, rather than merely seeing contacts move to another queue or channel?
- Post-test follow-up: Do incremental customers show acceptable returns, repeat purchase, and service outcomes?
Predefine success criteria and a stopping rule. Do not stop the test when the dashboard first looks positive, and do not change the primary metric after seeing which one moved. Record exclusions such as bots, employees, duplicate sessions, canceled orders, and shoppers who were not eligible for the assistant.
What to request from an AI shopping assistant vendor
Ask the vendor for event-level evidence and a test plan, not a blended “AI-assisted revenue” number. The proof should let your team reproduce the calculation in its own analytics and support systems.
Request:
- The exact definition of an impression, engagement, assisted session, conversion, and deflection.
- Treatment and control assignment logic, persistence, exclusions, and test dates.
- Numerators and denominators for every reported rate.
- Results by device, product category, traffic source, and shopper type.
- How cross-device journeys and multiple sessions are deduplicated.
- Whether support contacts are tracked across email, chat, phone, and social channels.
- The assistant’s answer-quality review process, escalation rules, and failure log.
- Return, cancellation, margin, and customer-satisfaction outcomes.
Anagram positions its Site Agent as a way to answer product questions, give guided recommendations, and provide next-step support; it also presents shopper-question analytics as a way to identify what customers care about and what creates conversion friction. For an ecommerce team evaluating it, the most credible pilot would therefore connect those conversations to a randomized conversion holdout and an intent-level support baseline. The resulting proof belongs to your store: incremental orders and revenue, fewer comparable contacts, and no unacceptable quality or post-purchase trade-off.