
Kroger has tested whether artificial intelligence can evaluate campaign creative before media spend begins and still predict what shoppers will actually buy. Working with Vidmob and MMA Global, the retailer connected visual and messaging attributes in its ads to e-commerce conversion data across Meta and Google’s DV360.
The result moves predictive creative scoring closer to a commercial decision tool. It also raises a more interesting question for marketing teams: if generative AI makes it easy to produce more variants, can a brand build a feedback system that decides which variants deserve budget without flattening the work into a formula?
Creative volume is becoming cheap. Decision quality is becoming scarce.
Table of contents
Jump to each section:
- What Kroger measured before media spend
- Why creative scoring changes the allocation question
- The risk of turning evidence into a template
- What marketers should know about predictive creative scoring
What Kroger measured before media spend
The MMA Global study combined Kroger’s live campaign results with Vidmob’s creative data platform and MMA Global’s research standards. The model analyzed elements such as messaging, narrative structure, branding and the presence of people on screen, then tested whether those attributes were associated with completed e-commerce purchases.
1,934 creative assets from Kroger campaigns across Meta and DV360 formed the study’s evidence base.
The team trained the model on a full year of campaign data and validated it against a later quarter the model had not seen. That forward validation matters. A model that explains past results can still be little more than a tidy description of history, while a model tested on unseen campaigns has a stronger claim to guide future decisions.
81% predictive accuracy linked creative scores to e-commerce conversion performance in the validation data.
The study also identified a set of validated creative guidelines. Assets aligned with those guidelines performed materially better than those that did not.
Up to 4 times higher conversion was recorded for creative that followed the data-backed guidelines.
This is not simply another claim that AI can recognize objects in an ad. Kroger’s test connected creative attributes to sales outcomes rather than using attention or view-through rates as a proxy. That distinction makes the score more relevant to marketers who must defend both production choices and media allocation.
Why creative scoring changes the allocation question
Most creative optimization starts after launch. Teams put several assets into market, watch platform performance and shift spend once winners emerge. Kroger’s model brings part of that decision forward, before weak work consumes meaningful budget.
The common assumption is that predictive scoring mainly helps teams edit an ad. The contrasting reality is that it may create just as much value by changing which finished ads receive distribution. The strategic implication is that creative intelligence belongs in media planning, not only in the production workflow.
More than twice as many conversions could be generated by reallocating the same budget toward higher-scoring creative, according to the study model.
That reframes the old split between “working” media and “non-working” creative. Media teams have long been expected to justify bids, audiences and placements with performance data, while creative investment is often treated as a prerequisite whose commercial effect is harder to isolate. A pre-launch score tied to conversion gives both sides a shared signal.
A score is useful only when it changes a decision.
For Kroger, that decision can be an edit, a pause or a shift in media weight. The operational opportunity is a tighter loop among creative, analytics and media teams, with each new campaign adding evidence that can improve the next brief. When every marketer can access similar generation tools, proprietary performance learning becomes more defensible than access to the tools themselves.
The risk of turning evidence into a template
Predictive scoring introduces discipline, but discipline can become conformity if teams treat the model as a universal recipe. The Kroger findings came from one retailer’s campaigns, objectives, platforms and conversion data. They are evidence for Kroger’s context, not a law of advertising.
That boundary matters as creative teams generate more variants. A guideline can help reject weak options, but it can also encourage teams to repeat familiar patterns because the model recognizes them. The safest work may score well precisely because it resembles what has already succeeded, while a distinctive idea may carry features the model has not learned to value.
The more useful role for predictive scoring is to narrow uncertainty, not eliminate judgment. ContentGrip’s recent interview on Picsart’s Vera ad-testing workflow makes the same boundary practical: a model can recommend a change, another agent can produce it, and the marketer still decides what ships.
Predictive scoring should make creative debate more informed, not make creative debate disappear.
Marketers should also watch for feedback-loop bias. If media dollars continually flow toward the patterns a model already favors, those patterns collect more conversion data and become even more influential in future scoring. Teams need room for controlled exploration, otherwise optimization can quietly reduce the range of ideas that reach the market.
What marketers should know about predictive creative scoring
Kroger’s test offers a useful operating model for teams dealing with more assets, tighter budgets and rising pressure to connect creative decisions to revenue.
Start with the business outcome. A score becomes more valuable when it predicts the result the business actually cares about. Kroger linked creative attributes to e-commerce conversion rather than stopping at an attention proxy.
Validate on unseen work. Historical fit is not enough. Teams should ask whether a model can hold up on later campaigns, different placements and changing audience conditions before they make it part of routine approvals.
Connect creative and media decisions. The score should inform both what gets revised and what receives spend. That turns creative intelligence from a production aid into an allocation signal.
Preserve a budget for learning. High-scoring assets can carry the core campaign, while controlled tests give unfamiliar ideas a chance to generate new evidence. Without exploration, the model can only deepen what the brand already knows.
The larger shift is not from human creativity to machine judgment. It is from loosely connected creative and media processes to a shared evidence system where each can challenge the other.
For marketers, that system may become a more durable advantage than faster asset generation. Competitors can buy the same models and connect to the same platforms. They cannot instantly reproduce a brand’s accumulated understanding of which creative choices drive its customers to act.
The brands that learn fastest will not be the ones that obey every score. They will be the ones that know when a score is strong enough to guide spend, when it should trigger a better question and when an unproven idea deserves a fair test.