How to measure whether a seasonal campaign changes ChatGPT product recommendations
Sep 2, 2026
A seasonal campaign can change ChatGPT product recommendations only if you can compare the same shopper intents before and after the campaign, record the products and their positions, and separate that movement from model variability and normal seasonal demand. Build a fixed panel of real shopper prompts, repeat each prompt consistently, measure product-level inclusion and rank, and use an untreated comparison set or matched market where possible.
Start with a fixed panel of real shopper prompts
Use real shopper questions as the measurement panel, then freeze the wording and intent mix for the campaign test. This shows whether ChatGPT’s recommendations changed for the buying situations your customers actually care about—not merely for a handful of convenient prompts.
Pull prompts from sources such as:
- On-site search and conversational questions
- Product-page questions and support conversations
- Pre-purchase questions about fit, compatibility, ingredients, use cases, price, or delivery
- Seasonal questions that include the campaign’s occasion, audience, or constraints
Group prompts by intent rather than treating every question as equivalent. For example, “best hiking jacket for wet weather” and “best hiking jacket for a summer trip” may concern the same category but represent different recommendation jobs.
Keep a prompt registry with the exact text, intent, market, language, date, and eligible products. Do not silently rewrite prompts after the campaign launches. If the panel changes, you are measuring a different population of questions.
Anagram’s AI Visibility product is designed to show how a brand appears in ChatGPT across important prompts and topics, alongside competitors and the sources shaping answers. Its shopper-question insights can help an ecommerce team turn on-site questions into candidate prompts for this panel. See Anagram’s AI Visibility.
Measure products, not just brand mentions
A higher brand mention rate does not prove that the campaign changed which products ChatGPT recommends. Capture the products named in every answer and calculate product-level measures for each prompt group.
| Measure | What it tells you |
|---|---|
| Product inclusion rate | How often a product appears in answers to eligible prompts |
| Top recommendation rate | How often it is the first or lead recommendation |
| Mean and median position | Where the product appears when it is included |
| Share of recommended products | The product’s portion of all product recommendations in the panel |
| Competitor inclusion rate | Whether competing products gained or lost presence in the same answers |
| Citation coverage | Whether the answer cites a page that supports the product or relevant attribute |
| Recommendation framing | Whether the product is presented as best overall, best for a use case, alternative, or poor fit |
Define product identity before collecting results. A model may name a product by title, line, variant, or retailer listing. Map those names to a canonical SKU or product family, and keep the mapping rules unchanged across pre- and post-campaign readings.
Also record eligibility. If a product is out of stock, discontinued, unavailable in the shopper’s region, or outside the stated budget, its absence should not be counted as a visibility failure. A recommendation must first satisfy the prompt’s constraints to be a meaningful win.
Create a baseline before the campaign
Run the frozen prompt panel before the seasonal campaign changes the pages, merchandising, or product availability. The baseline should include the full answer—not just whether the brand appeared—so you can compare product inclusion, rank, framing, competitors, and cited sources later.
Repeat prompts rather than relying on one observation. ChatGPT can produce different recommendations for the same wording, and OpenAI describes shopping research as personalized, conversational, and based on a user’s requirements and preferences. That makes one-off spot checks weak evidence of a campaign effect. Read OpenAI’s explanation of shopping research.
Use the same test conditions for every reading:
- Exact prompt text and prompt order, where practical
- Same country, language, device or account state, and product availability assumptions
- Same model or ChatGPT shopping mode
- Same time window and collection procedure
- Same rules for counting products, variants, citations, and rank
Store the timestamp and response for every run. Record model or mode changes as breaks in the series rather than blending them into a single trend.
Compare pre- and post-campaign movement
Calculate the change for each product and prompt group, not only the aggregate campaign result. A useful first comparison is the post-campaign inclusion rate minus the baseline inclusion rate, with the same calculation for top recommendation rate, position, share of recommendations, and citation coverage.
Then inspect the distribution behind the average. A campaign may improve a product’s visibility for “gifts under a budget” prompts while reducing its presence for everyday use-case prompts. The aggregate number can hide that trade-off.
A practical reporting table looks like this:
| Prompt group | Product | Baseline inclusion | Post-campaign inclusion | Change in lead position | Main source change |
|---|---|---|---|---|---|
| Seasonal gift | Product A | Record the observed rate | Record the observed rate | Record the observed movement | Note new, lost, or unchanged citations |
| Use case | Product B | Record the observed rate | Record the observed rate | Record the observed movement | Note the sources ChatGPT used |
Use observed rates rather than claiming certainty from a small number of runs. Include the number of prompt runs, the number of answers containing each product, and the number of eligible products in each group.
Separate campaign impact from seasonality
A before-and-after change is evidence of association, not proof that the campaign caused ChatGPT to change its recommendations. Product pages, reviews, retailer listings, inventory, competitor campaigns, model updates, and seasonal demand can all move at the same time.
Strengthen the counterfactual in one of three ways:
- Hold out prompt groups. Keep a comparable set of non-seasonal or unrelated intent groups untouched. Compare the campaign-targeted change with the change in the holdout.
- Use matched products or categories. Pair targeted products with similar products that did not receive the campaign treatment, then compare their changes in the same prompt groups.
- Use matched markets. If the campaign runs in selected regions, compare those regions with similar regions where the intervention did not run, while holding prompt language and availability rules consistent.
The strongest practical result is not “Product A appeared more often after launch.” It is closer to: “Product A’s presence and lead-recommendation rate rose for the targeted seasonal prompt groups relative to a comparable untreated set, while eligibility and test conditions stayed stable.”
If you cannot create a control or holdout, label the result as a monitored change. Do not turn a time-series movement into incremental revenue or causal lift.
Track the sources behind recommendation changes
A product can move up because ChatGPT found stronger evidence, not because it interpreted the campaign creative directly. Compare cited URLs and the claims those pages support before and after the launch.
Look for changes in:
- Product-page coverage of the seasonal use case
- Clear specifications, compatibility, sizing, ingredients, or materials
- Retailer or review pages that ChatGPT cites instead of the brand site
- Conflicting price, availability, or product details
- Competitor pages that answer the same seasonal question more directly
Anagram says its AI Visibility dashboard shows citation sources and where competitors are winning. That gives the team a way to investigate a visibility change rather than treating the recommendation output as a black-box score. The next action might be a product-page correction, a missing comparison, or better coverage of a shopper question—not another broad campaign.
Connect ChatGPT visibility to real shopper outcomes
Visibility is an upstream measure. To prove business value, connect the prompt and product findings to what happens after shoppers arrive on the site.
Keep three layers separate:
- AI visibility: product inclusion, position, competitor share, framing, and citations in the fixed prompt panel
- On-site behavior: visits from ChatGPT where identifiable, Site Agent engagement, product views, comparison actions, and add-to-cart events
- Commercial outcome: conversion, revenue, margin, support deflection, or other agreed business measures
Do not claim that a ChatGPT recommendation caused an order simply because both occurred. Use tagged links where the journey supports them, preserve referral and landing-page data, and compare exposed traffic with a suitable on-site control when possible.
Anagram’s branded Site Agent is a separate on-site layer: the company describes it as helping shoppers with conversational product answers and guided recommendations while they compare options and decide what to buy. Its shopper-question analytics can connect the questions customers ask on the site with the questions the team monitors externally. That connection helps explain whether a seasonal visibility change addresses an actual buying need.
Build the campaign readout buyers can trust
A credible seasonal readout should let another team member reproduce the comparison and see its limits. Include:
- The frozen prompt panel and intent definitions
- Product-to-SKU or product-family mapping rules
- Collection dates, model or mode, market, and availability assumptions
- Number of repeated runs per prompt
- Product inclusion, lead position, rank, share, framing, and citation results
- Competitor and holdout comparisons
- Any model, catalog, pricing, inventory, or site-content changes
- On-site and commercial outcomes, reported separately from visibility
- A list of findings that are directional rather than causal
For the next seasonal campaign, start the baseline before launch, repeat the same panel during the campaign, and run a final post-campaign read after the campaign conditions end. Keep an always-on set of general shopping prompts beside the seasonal set; it helps reveal whether the campaign changed targeted recommendations specifically or shifted broader product visibility as well.