Kroger, Vidmob and MMA Global on September 30, 2026 published research covering 1,934 of the grocer's video and image ads, concluding that a model trained on a year of Kroger's campaign data could sort ads it had never seen into stronger and weaker converters with 81% accuracy.

In Short

Kroger handed almost 2,000 of its online ads to a company called Vidmob, which built a computer model that guesses whether an ad will lead to purchases, and the model got it right about eight times in ten on ads it had not been trained on. This matters for anyone paying for ads on Meta or Google, because the study says roughly one dollar in every two Kroger spent there was backing the weaker ads, and that ads showing real people enjoying food beat glossy product shots. Kroger now plans to write these rules into its creative briefs and run a live head-to-head test, which is the part that will show whether the numbers hold outside a spreadsheet.

What was measured

The research, released as a whitepaper titled "Predictive Creative Scoring Leads to Real Ecommerce Outcomes" and announced in New York at 9:00 a.m. Eastern Daylight Time, is unusual less for its conclusion than for its source material. Vendor studies of creative performance usually pool anonymised assets; this one names a single retailer and sets out a methodology with enough numbers attached to check it against itself. Not all of those numbers agree.

According to the whitepaper, the analysis covered a 12-month test period running from January 2025 to January 2026, followed by a validation period in the first quarter of 2026. The dataset comprised 1,934 pieces of live video and image creative: 1,062 ads on Meta and 872 on Google's Display & Video 360 (DV360). The key performance indicator is given simply as "Conversion by platform" - ecommerce purchase conversions, read separately for each buying platform.

The work followed three steps, which the document labels Collect, Analyze and Validate. Vidmob used statistical testing and predictive modelling to find which creative attributes were associated with stronger performance; then, according to the whitepaper, "Models trained on 2025 data were tested against a separate period in Q1 2026." Vidmob's underlying system, according to the company, draws on 40 models, more than 20 million analysed assets, more than three trillion creative elements tied to performance and 13 platform integrations.

Each asset was coded along two dimensions. Message covers what an ad communicates: brand identity and emotional tone, the value proposition and the proof points offered for it, call-to-action phrasing, and Kroger-specific elements such as loyalty offers. Craft covers how an ad is built: live action against animation, the use of talent, emotional intensity, and what the whitepaper calls attention flow - the opening hook and where a viewer's focus lands within the first three seconds. Every asset in the study was scored on both.

From 231 attributes to 22 guidelines

The filtering was steep. According to the whitepaper, the study tested 231 attribute levels, of which 105 passed standard statistical testing. A candidate guideline then had to clear three further conditions: it had to have moved real conversions, it had to hold up in live-action creative, and the predictive model had to agree. Anything lacking agreement across all three was dropped. Twenty-two survived, about 9.5% of what was tested, split between 17 guidelines for DV360 and five for Meta. A footnote states that the guidelines were applied regardless of creative concept.

The document does not say which significance threshold counted as "standard", nor whether any correction was applied for running 231 tests at once. At a conventional 5% threshold, roughly a dozen of 231 tests would be expected to pass by chance even if no attribute had any effect at all. With 105 passing, chance cannot account for most of the result, but the second and third filters carry much of the weight in separating signal from noise - and they are described only in outline.

The live-action condition matters too. Kroger's creative includes animated work - two images reproduced in the whitepaper show 3D-animated characters - so the final list appears limited to rules that survive within live-action ads. How many of the 1,934 assets were animated is not disclosed.

The headline numbers

Three figures carry the announcement: 81% accuracy, a fourfold conversion-rate gain and up to 2.2 times the conversions from the same budget. According to the whitepaper, the first is the model's "accuracy in predicting the asset to be a low or high performer based on ecommerce conversion metric across Meta and DV360," measured on campaign data the model had not seen in training.

That is a binary call - high or low - and its meaning depends on where the line sits. If the split falls at the median, a coin toss would score 50%, so 81% is a substantial improvement on chance. The whitepaper does not specify the cut-off. Nor does it report accuracy separately for Meta and DV360, the number of assets in the validation set, or the balance between false positives and false negatives. For an advertiser deciding whether to kill an asset before it runs, wrongly rejecting a good ad costs something different from wrongly passing a weak one.

Where the documents diverge

Several of these figures are described inconsistently across the source material, and the differences bear directly on how the results can be read.

  • The 4x figure appears as "up to a 4X higher conversion rate" on page four of the whitepaper, as "roughly a 4X improvement" on page five and as a "4x average improvement" on page 13. The press release says creative aligned with the model's recommendations "averaged 4x higher conversion rates." An average and a ceiling are different statistics. Vidmob has also described it, in material sent to publications, as a comparison of higher-scoring with lower-scoring ads, whereas the whitepaper ties it to whether the guidelines were followed.
  • The cost figure carries the same ambiguity. Page four and the press release cite "up to 70% lower cost per conversion." Page 12 states that, on average, conversions driven by high-scoring creative cost approximately 70% less than those driven by low-scoring creative.
  • The validation window is given as "a one month validation period (Q1 2026)" on the methodology page and as "a full quarter of unseen data" in the executive summary. The press release refers to an independent Q1 2026 dataset without specifying its length.
  • The guidelines are described as "22 validated guidelines across both platforms" and, in a chart, as "22 double-proven guidelines." Yet the same paragraph labels the five Meta guidelines "directional", in contrast to the 17 "validated" ones for DV360.
  • MMA Global's research head is named Vas Bakopoulos in the press release and Vas Bakolopous in the whitepaper's acknowledgements.

None of these differences overturns the findings. Together, though, they leave open whether 4x and 70% describe typical outcomes or the best cases in the dataset. Neither Kroger nor Vidmob published confidence intervals for any of the three headline figures.

Faces over packshots

The most concrete creative finding concerns people. According to the whitepaper, Meta and DV360 each responded to their own levers, but one theme held on both platforms: "Ads built around relatable, human narratives consistently outperformed glossy product shots and transaction-focused creative."

On Meta, 87% of assets featuring a relatable human scenario ranked above the median, against 38% of assets without one. Ads showing people reacting or enjoying the moment ranked above the median 76% of the time, compared with 45% without. Simple arithmetic on those figures, assuming the median splits the Meta set evenly, implies that only about a quarter of Kroger's Meta assets contained a relatable human scenario and roughly one in six showed people reacting or enjoying. Most of the grocer's Meta creative, in other words, lacked the attribute most strongly associated with above-median performance.

Kroger ties the result to its brand purpose, "Feed the Human Spirit." According to Kroger, Vidmob and MMA Global, campaigns featuring relatable human moments, such as people sharing and enjoying food, produced stronger ecommerce results than transactional shopping imagery or product-first creative.

Two qualifications apply. Both published attribute comparisons come from Meta, where the guidelines are labelled directional; the 17 DV360 guidelines, which the study describes as validated, are not itemised. And the comparisons are correlational. Assets with people in them may differ from those without in format, placement, audience or flight timing, and the published figures control for none of those.

The reallocation arithmetic

According to the whitepaper, a review of Kroger's historical media spend found that approximately 50% of all investment had gone behind creative in the weaker tier. On DV360 specifically, lower-scoring creative carried both lower conversion rates and higher CPMs, so the weaker ads were also dearer to deliver per thousand impressions.

The whitepaper separates two effects. The guideline-following lift of 4x reflects what becomes achievable when new creative is built to specification from the outset. The reallocation-only lift of 2.2x reflects what could be gained immediately by moving existing budget towards existing higher-scoring assets, "without producing new assets."

The whitepaper calls the 2.2x figure "potential", and the word carries much of the weight. It is an estimate from historical data, not the result of a campaign in which money was actually moved, and it assumes that high-scoring assets would keep converting at their observed rates as more spend flowed through them - an assumption that creative fatigue, frequency and auction dynamics routinely test. The DV360 CPM gap raises a further question the document leaves unanswered: did lower-scoring creative cause higher prices, or did it simply run in more expensive formats and placements, such as video rather than display? That is a textbook case of endogeneity, in which whatever decides where an ad runs also shapes how it performs.

Scale is unknown as well. Kroger reported $1.18 billion in advertising spend for 2025 in its March 31, 2026 annual report, but the whitepaper does not disclose how much of that ran on Meta and DV360. The dollar value of the weaker tier therefore cannot be calculated from the published material.

Platform automation is heading the same way regardless. Google deprecated manual creative rotation in DV360, with display and video line items migrating automatically from August 12, 2026 to creative selection that follows the line item's bidding goal; under maximise-conversions bidding, the platform now favours creatives expected to drive the most conversions. Kroger's test period predates that change. Part of the reallocation the study models is now something Google's system attempts by itself, at least within DV360 line items that bid for conversions.

What the companies said

No Kroger executive is quoted in the announcement. Kroger's position is set out only in the whitepaper, whose acknowledgements list five Kroger staff, three from Vidmob and one from MMA Global.

Alex Collmer, founder and executive chairman of Vidmob, framed the result as a question of selection rather than production. "Every brand can produce more creative than ever. The advantage now belongs to the ones who know which creative will drive the outcome they need," he said, according to Vidmob. "Kroger's results show what happens when brands treat creative performance as ownable IP: every choice becomes measurable, and every campaign makes the next one smarter."

Bakopoulos, head of insights and research at MMA Global, linked the work to the speed of production. "As AI makes creative production ten times faster, judgment has to scale with it," he said. According to MMA Global, the Kroger results show that predictive scoring can bring that judgment in before a campaign goes live. He described the next phase in terms of measurement infrastructure: "The next chapter goes deeper, to the segment level, and wider, into mix models and buying tools, so creative earns a seat at the table where investment decisions are made."

What comes next for Kroger

The whitepaper sets out how Kroger intends to put the findings to work. The guidelines are to be written directly into the creative briefing process - "not a report referenced once and shelved," in the document's words. Kroger frames the scores as complementary to media mix models, explaining why creative inside a channel performed as it did, with coverage meant to extend beyond Meta and DV360 to the rest of its media mix.

The step that matters most has not happened yet. According to the whitepaper, Kroger plans an in-market forward test "pitting high-scoring creative directly against low-scoring creative to turn this predictive model into measured, real-world proof." Until it reports, the evidence is observational: it shows which ads converted better in the past and that a model can recognise them, not that the creative caused the difference. A randomised holdout or split test is the conventional way to settle that.

Among areas for further research, the document lists YouTube, Snapchat, TikTok and Instagram - the last a curious inclusion, since Instagram is sold through Meta and the Meta results are not split between Facebook and Instagram placements.

There is also a gap between the outcomes Kroger cares about and the one the study measured. The executive summary names ROAS and sales lift as the outcomes creative can move, yet the KPI was platform conversion. What counted as a conversion, which attribution window applied, and whether the figures came from Meta's and Google's own reporting are not stated.

Why this matters for advertisers

The study arrives as creative becomes the main input advertisers still control by hand, and research keeps finding a gap between belief and practice: a Perion and EMARKETER survey of 111 US marketers published on April 15, 2026 found that 89.2% considered creative important to campaign performance, while only 3.6% said its performance was well understood and actively optimised. The same survey found that 72.1% waited for performance to decline before changing creative at all.

Generative tools have widened that gap rather than closed it. Research WARC published with TikTok on July 14, 2026 found that 88% of 400 marketers reported higher creative volume since adopting generative AI, but only 45% reported better quality. On Meta, more than 9 million small businesses were using at least one AI ad creative tool by the second quarter of 2026, while the Advantage+ suite hands decisions about audience, placement, budget and creative to Meta's machine learning systems. If production is cheap, as Bakopoulos's "ten times faster" remark implies, what is scarce is the decision about what to run.

Pre-flight scoring has become a crowded category. Smartly added Creative Predictive Potential on November 18, 2025, scoring assets before campaigns go live. System1 said on April 9, 2026 that it was beta testing a predictive AI creative measurement tool inside its ad testing platform. Mediaocean's NIVO layer, which went live on June 11, 2026, included a Predictive Scoring Agent among its twelve agents, and Brunner, a Pittsburgh independent agency, acquired the creative prediction platform AdSkate on July 16, 2026. Vidmob itself was among the vendors whose agents Horizon Media said in June its own buying agents could interact with. Most describe their predictive power in general terms; the Kroger study is a rare case of an accuracy figure, a named advertiser and a methodology appearing in one document.

For Kroger, the research sits within a wider build-out of its advertising business. Kroger Precision Marketing joined Google's Commerce Media Suite on March 24, 2026, bringing SKU-level sales reporting for Kroger shoppers reached through DV360 and YouTube. The grocer's retail media network has kept growing since: Kroger Precision Marketing increased profit 24% in the second quarter of 2026, while e-commerce sales rose 20%. The creative study, however, concerns Kroger in a different role - as an advertiser buying media on other companies' platforms to sell groceries online, rather than as a seller of shopper audiences to consumer goods brands.

MMA Global's involvement connects the work to measurement research the alliance has done with Kroger before. A whitepaper TransUnion and MMA Global released on October 2, 2025 found that at Kroger, traditional tools captured just 17% of the complete effect of marketing on online sales, with long-term effects reaching six times the short-term conversion rate lift. Set against that, the new study measures only the portion of impact that conversion tracking can see. Does creative tuned for short-window conversions also perform over the longer horizon that MMA's earlier work emphasised? Neither document addresses it. MMA is a nonprofit alliance of chief marketing officers operating in 16 countries with more than 825 corporate members.

That observational character matters. Karan Dhir, a product manager at Genentech, argued in September 2026 that Facebook's own experiments show standard ad measurement often off by a factor of three, because ad delivery concentrates on people already likely to convert. The same selection problem can apply to creative: if Kroger's human-moment ads were served more often to warmer audiences, part of their advantage would belong to the audience rather than the ad. The forward test is the point at which that can be checked.

Timeline

Summary

Who: Kroger, the US grocery retailer; Vidmob, a creative data company whose models scored the ads; and MMA Global, the marketing alliance that collaborated on the research. Alex Collmer, Vidmob's founder and executive chairman, and Vas Bakopoulos, MMA Global's head of insights and research, are quoted.

What: A whitepaper analysing 1,934 Kroger video and image ads - 1,062 on Meta and 872 on DV360 - which found that a predictive model classified unseen ads as high or low converters with 81% accuracy, that guideline-following creative showed up to or on average 4x higher conversion rates depending on the document, and that reallocating spend towards higher-scoring existing assets could yield up to 2.2x conversions. Of 231 attribute levels tested, 22 became guidelines, 17 for DV360 and five directional ones for Meta. Ads built around relatable human moments outperformed product and transactional creative.

When: Published on September 30, 2026. The test period ran from January 2025 to January 2026, with validation in the first quarter of 2026. An in-market forward test is planned but undated.

Where: Kroger's ecommerce advertising on Meta and Google's Display & Video 360, with results announced in New York.

Why: Kroger wants a repeatable way to connect creative decisions to business outcomes and to embed scoring in its briefing process. For the wider market, the study offers rare named-advertiser evidence for pre-launch creative scoring as AI multiplies creative volume, though its findings are observational, several figures are described inconsistently across documents, and the experimental test that would establish causation has not yet run.