Model training and data
Was Amplify AI trained separately for each category?
No. Amplify AI uses a single global model, with country and category as model features. The model learns how ads score in each country (for example, the US vs. Mexico) and in each category.
During development, we found the model is most predictive when it sees a broad, consistent range of advertising. Gaps in category coverage or inconsistent audiences add noise.
To ensure breadth and consistency, Zappi collects its own human training data. This data covers every country and platform where the model is available, across 20+ parent categories (such as food, beverages, personal care, and financial services) and 100+ child categories.
Any single vendor's data carries the biases of its customer base: skewed toward larger brands and certain categories, and not representative of the full range of ads running on TikTok, Instagram, Facebook, and YouTube. Instead, Zappi uses a purpose-built data asset based on two principles:
- Coverage: We sample across categories, brands, and creative types so the model sees the full variation it needs to judge any ad. Training on narrower category sets measurably reduces performance, because similar creative gives the model too little signal to distinguish effective from ineffective work.
- Audience consistency: Every ad uses the same recruitment logic: a category relevant sample, meaning anyone for whom the category is at all relevant. This keeps audience variation from adding noise the model would otherwise try to learn from.
Is Amplify AI based on global or country-specific data?
Both. Human data is collected by country and feeds a single global model, where country is a feature. The model learns market-specific patterns (how US ads score vs. Mexican ads) and patterns that hold across markets.
Amplify AI is released in a market only when:
- enough human data exists for that country, and
- the model is validated as predictive of human responses in that market.
The model is global, but results in each market are locally grounded and locally validated. For example, results for a US ad reflect US consumers. The model is trained and validated on US human data, and results are not extrapolated from other markets.
How do categories in Amplify AI differ from other Amplify solutions?
Amplify AI uses standard categories only, while other Amplify solutions also support customer-specific categories. Because the model is trained on category relevant samples, results always reflect the standard category. The human training data does not represent the specific, niche audiences some customers use, so custom categories are not available.
What do "category relevant" and "platform relevant" mean?
These criteria qualify respondents in the training data:
- Category relevant: Respondents rate how relevant the category is to them (something they buy or use often, keep up to date with, or consider personally important). Respondents who answer "not at all relevant" are screened out of the survey.
- Platform relevant: Respondents must regularly use the platform the ad is designed for (such as TikTok or Instagram) to be included in that platform's results.
Category availability and customer fit
Which parent categories are in the training data?
These parent categories are well represented in each country's training data, so ads in them receive high-quality predictions out of the box:
- Food
- Beverages
- Restaurants
- Alcoholic Beverages
- Financial Services
- Telecommunications & Internet (Service Providers)
- Consumer Healthcare
- Consumer Electronics
- Travel & Accommodation
- Home Care
- Baby, Parenting & Childcare
- Personal Care & Beauty
- Home Goods, Furnishings, Furniture & Garden
- Social Media, Apps & Entertainment (Software, Services)
- Automotive
- Retail
- Pet Care
- Large & Small Home Appliances
- Apparel, Footwear & Accessories
To see the current list of approved parent and child categories, draft a configuration in Acme like this one.
What if the parent category is available but the child category I need isn't?
Submit a Sample Operations New Category request ticket in Salesforce (SFDC). Include the parent category and your proposed child category name, then follow the standard process. Sample Operations may suggest a different child category name based on active categories in the market and overall customer demand.
Which parent categories aren't in the training data, and can customers in them still use Amplify AI?
These parent categories, and their child categories, are not currently used in the training data:
- Tobacco & Smoking Alternatives
- Public Services, Government, Education & Charity
- Pharmaceuticals (big pharma; OTC healthcare is included)
- Personal and Home Services
- Toys, Games & Sports Equipment
- Office & Stationery
- Non-profit organizations & charities
Other categories not listed here may also be missing. For a missing parent category, the model still predicts, but with lower confidence. Customers in these categories may still be able to use Amplify AI, with caveats.
Before contacting Sample Operations or setting customer expectations on availability, contact Kim Malcolm to discuss.
Broad audiences and targeting
Key points at a glance
- Amplify AI is a different tool for a different job than tailored, human-audience research.
- The category sets a broad scope, not a narrow target. The model is trained on everyone for whom the category is at all relevant, and predicts for that group based on the ad, the country, and the category. Changing the audience name in the platform doesn't change the training data, so it doesn't change results. This setup gives a reliable, predictive read on assets the model has never seen, at the speed and scale high-volume workflows need.
- It's built on a validated human methodology and continuously checked against new human responses. Results land within 1 point of the human score 84% of the time. It provides a strong signal at scale, not in-depth precision. Use human responses for in-depth precision.
- It answers questions like "Is this derivative campaign asset good to run?" or "Which of these social video assets should get more or less investment?" To learn early whether and why a core idea or core execution works, use human research. Once a core video asset is validated and optimized with humans, use Amplify AI for the assets that follow. For always-on social content, Amplify AI gives a directional signal at scale to guide boosting and investment decisions and to learn what works in aggregate.
Why does Amplify AI use broad category audiences instead of tailored ones?
The human training data is collected from broad category audiences. Mixing audiences in the training data adds noise and weakens the signal.
Amplify AI is built for fast, directional reads across a high volume of assets, so you can check, choose, and invest in stronger assets and learn at scale. Results reflect how well an ad lands with people for whom the category is relevant.
Narrower audiences would make the model precise for one use and unreliable for others as priorities change. A strong, consistent foundation keeps the model accurate in a changing media landscape.
Are broad audiences less accurate?
No. A broad audience reflects Ehrenberg-Bass Institute principles on where brand growth comes from and helps the model stay accurate as the market evolves. It includes everyone for whom the category is at all relevant: anyone who could become more predisposed to the brand in the short or long term.
Each prediction draws on thousands of respondents and ads in and around the category. This is a different kind of precision from a bespoke audience read, suited to a different type of asset, not a lesser one.
The data is Zappi-owned (no external bias), built on a human methodology proven to predict in-market outcomes, and continuously checked against new human responses to close gaps and catch drift.
Can Amplify AI target a specific audience, such as adults 25+?
No. Amplify AI predicts results at the total level only and can't restrict or skew toward a demographic group. It uses a machine learning model rather than purely synthetic responses (which could be aggregated by demographic), because the machine learning approach proved far more accurate. Predicting broad subgroups is under exploration for the future.
Can I report results by subgroup, such as age or gender?
Not yet. Subgroup reporting (such as male/female or younger/older) needs significantly more training data to reach reliable confidence. It's planned for future releases as the data asset grows.
What does Amplify AI provide if not a tailored audience read?
A directional, quantified read on how effectively an asset makes the brand more likely to come to mind powerfully and positively, driving brand growth among a category relevant sample. It's built for bulk review: flagging which assets are strong, which need work, and which to pull. It doesn't provide audience-specific diagnostic depth on why. That tradeoff is what enables speed and scale.
Does Amplify AI replace tailored audience research?
No. Amplify with human responses is for high-risk, high-value assets that need rigorous, in-depth validation: pre-testing before launch, understanding what works and why, and confirming flagship creative is ready for spend. Amplify AI picks up after that, covering the cutdowns, versions, and edits a campaign produces that wouldn't otherwise get detailed human research. It confirms a wave of content is good to run without diagnostic-level time or cost per asset.
Validation
What does the 84% validation figure mean?
The 84% figure is an adjacency agreement rate, not a statistical confidence interval. In 84% of cases, Amplify AI's predicted Creative Power score is within one point of the human respondent score on the five-point scale.
The agreement is strongest at the ends of the scale:
- When Amplify AI predicts an ad as weak (1), the human result is weak or limited (1 or 2) in 91% of cases.
- When Amplify AI predicts an ad as excellent (5), the human result is strong or excellent (4 or 5) in 100% of cases.
The underlying human methodology is linked to in-market outcomes: ads with high Creative Power are 1.7x more likely to drive Brand Lift. See the attached deck for detailed validation results.
How is the model validated?
Accuracy testing uses a holdout validation approach. We hold back 15% of human-tested ads from the training data, train the model on the remaining 85%, then test its accuracy on the held-out ads. Validation figures therefore reflect how well the model predicts ads it has never seen, not how well it reproduces scores from ads it learned from.
How do you check for model drift after launch?
Validation is ongoing, not a one-off exercise.
- Retraining: As we add categories, markets, and data, the model is retrained about every 2 to 4 weeks.
- Benchmarking: We continue collecting new human advertising data. Even after retraining slows, we plan to test around 100 digital ads per month in the US and UK with human respondents and run Amplify AI on the same ads in parallel. This provides an ongoing comparison of predicted vs. human results, so we can monitor whether performance changes over time.
- Filling gaps: We use human data collection to fill gaps in the dataset, such as adding breadth across categories or ad types, rather than letting whichever customer tests happen to run determine the training data.
Use Case
How does Amplify AI handle novel or unconventional creative?
Amplify AI is trained on a broad range of advertising rather than a series of narrow category models. During development, we found that single-category models were less predictive because the creative range within one category was often too narrow. Adding breadth across categories, brands, creator-led and brand-led content, ad lengths, and platforms improved predictive performance. This is one reason Zappi collects its own human advertising dataset rather than relying on customer projects.
Like any predictive model, Amplify AI is strongest when a new ad relates to patterns in its training data. For creative that is unlike anything the model has seen across any category, such as a truly breakthrough or unconventional execution, prediction confidence is lower.
Creative that is completely unlike anything in the dataset is relatively rare, but this is a useful guardrail. For a genuinely new idea, we recommend human research early in development to confirm it resonates, stands out, and cues the brand as intended.
When should I use Amplify AI, and when should I use human research?
Match the research method to the risk and precision the decision requires.
Amplify AI works well for:
- getting directional signals on assets that would otherwise receive no consumer feedback
- identifying stronger and weaker creative across a set to guide investment
- evaluating creator-led content at scale
- finding patterns across multiple assets or campaigns
- adding a brand and creative signal to media or engagement data
Amplify AI is not recommended for:
- Major hero investments or other high-risk, foundational creative decisions as the only research input. Use human research for these.
- Fine distinctions between subtle variations of the same execution. Its strength is directional comparison and learning across a broader body of work, not precision between marginal variants.
- Unprecedented creative. As described above, human validation is the better fit.
Amplify AI does not replace human research. Use human research when the decision calls for depth and confidence, and Amplify AI when asset volume means you would otherwise get little or no consumer signal.