A scorecard is a fixed set of criteria, each carrying a weight, applied repeatedly to the same class of subject so that results can be set side by side. An advertiser rating six supply-side platforms on the same eight questions is using one, as is a retailer grading its delivery times and returns against competitors in its category. What separates a scorecard from a report is structure and repetition: criteria are agreed before the assessment begins, stay constant between rounds, and produce a grade rather than a narrative.

The instrument exists because judgement does not scale and single numbers mislead. A team evaluating twelve vendors cannot hold twelve impressions in mind and weigh them fairly, so each becomes one composite figure a procurement committee can act on. That compression is the whole value of a scorecard, and the source of every complaint about it.

How a scorecard is built

Construction involves four contestable decisions.

The first is the unit being rated. Supply-side platforms, agencies, publishers, creatives and measurement vendors each need different criteria, and a scorecard rating two unlike things produces a comparison that only looks rigorous.

The second is the criteria, usually drawn from whatever the organisation already measures, so scorecards inherit the blind spots of existing reporting. Local television measurement guidelines published by CIMM and TVB in May 2026 show a more deliberate approach: the buyer's guide was structured for use as a request-for-information checklist or vendor scorecard, each question asking providers to demonstrate methodology, current capability, known limitations and roadmap, and to supply documentation rather than verbal assurance.

The third is weighting. A scorecard treating fee transparency and account service as equally important encodes a claim about what the buyer values, and that claim is rarely written down. Take an illustrative supply-path scorecard: working media share at 30%, invalid traffic at 20%, fee disclosure at 20%, path duplication at 15% and technical support at 15%, each scored from one to five. Moving fee disclosure to 5% reorders the ranking without changing a single measurement.

The fourth is the scale, and with it normalisation. Converting a measured percentage and a one-to-five opinion to a common range silently equates a rate derived from billions of impressions with the view of one account director.

Two further settings decide whether any of it matters. Thresholds fix what happens at each band, such as a partner below a cut-off losing budget at the next review, and cadence fixes how often the assessment repeats. Run once, a scorecard is an audit; run quarterly, it governs.

Origin and evolution

The idea of a balanced panel of indicators predates advertising, with French firms using a tableau de bord of operational indicators from the 1930s. The modern form dates to a Harvard Business Review article published in the January to February 1992 issue, pages 71 to 79. Robert S. Kaplan, then professor of accounting at Harvard Business School, and David P. Norton, president of the consulting firm Nolan, Norton and Company, opened with the proposition that an organisation gets whatever it chooses to measure. Financial measures, they argued, report past actions and need complementing with operational ones. The article drew on a 1990 Nolan Norton Institute project covering twelve companies whose value rested largely on intangible assets.

Kaplan and Norton set out four perspectives: financial, customer, internal processes, and innovation and learning. Two contributions mattered more than the categories: they codified a loose collection of metrics under one name with a consistent taxonomy, and required each measure to derive from strategy rather than from availability. Follow-up articles appeared in 1993 and 1996.

Analyst firms then industrialised vendor comparison. IHS Markit published an Internet of Things platform vendor scorecard ranking nine suppliers in 2018 on market presence and momentum, scoring device volumes, growth rate, vertical breadth, ecosystem development and technical innovation. That template, a fixed criteria set applied to a named competitive field, moved into technology procurement and then into media.

Programmatic advertising acquired its own criteria late. The Association of National Advertisers published its programmatic media supply chain transparency study in December 2023, using log-level data supplied through TAG TrustNet across 21 member companies, $123 million of spend and 35.5 billion impressions recorded between September 2022 and January 2023. Its findings, among them an average campaign running across 44,000 websites and 36 cents of each dollar reaching a valid, viewable impression, converted directly into procurement criteria. Inclusion lists, direct seller relationships, log-level data access and measurability rates became scorecard rows.

Where scorecards appear in marketing

Four families use the word, and they are not interchangeable. Strategic scorecards track organisational performance in the Kaplan and Norton lineage, usually owned by finance. Vendor and partner scorecards rate suppliers during selection and through the contract term, covering agency reviews and quarterly assessments. Media quality scorecardsgrade inventory, paths and placements against verification signals, and are the version programmatic practitioners meet daily. In-product scorecards are features named after the concept. Google added a Store Quality scorecard to Merchant Center on August 27, 2024, letting retailers compare delivery, shipping costs, returns and seller ratings against others in their category, and Kochava shipped a Creative Scorecard ranking creatives by impressions, clicks and new users in its second-quarter 2026 release. Looker Studio uses the word differently again, for a chart type displaying one headline figure, with a January 2025 update adding control over whether dimensions or metrics appear as the primary field.

Limitations and disputes

The most serious objection is structural. Once a scorecard governs budget, its criteria become targets, and inventory optimised against them scores well without being worth buying. Research published on July 28, 2026 by TAG, the ANA and Fiducia found AI-generated inventory grading as premium more than 70% of the time, with an invalid traffic rate of 0.05% against 0.32% for clean supply, viewability of 77.2% against 74.9%, and a TrueCPM of $7.08 against $6.15. The junk outscored real inventory on every row a media quality scorecard contains, and cost 15% more. Fraud operations exploit the same arithmetic: IAS blocked 800 domains whose traffic beat legitimate activity on every metric a buyer would check.

Weighting opacity is the second complaint: vendor rankings appear as conclusions rather than workings, leaving them impossible to re-run under different priorities.

Self-scoring is the third. Google's Optimization Score rates an account from 0% to 100% against recommendations Google itself generates, and the company reported that accounts raising the score by ten points saw conversions rise 10%. Reddit shipped its own Optimization Score in July 2025. In both cases the party selling the inventory sets the criteria, scores the account and quantifies the benefit of compliance. The tension runs the other way too: BidSwitch opened independent bidstream audits to its supply partners at no cost while retaining an interest in the paths reviewed.

Misalignment is the fourth. A report published on April 7, 2026 by In-Store Marketplace and Catalyst Media Consulting identified four incompatible scorecards applied to the same in-store activation: agencies asking how the channel compares with others, merchants asking whether product moved, retail media networks asking whether spend returned a margin, shopper marketing teams asking whether joint business plan commitments were met. Every party measures diligently; none agrees on the verdict.

Scorecards also work as instruments of persuasion. Nielsen's Gauge sets no linear rates directly, yet a statement from VAB chief executive Sean Cunningham described it as a "scorecard of relative scale among video advertising options" shaping channel allocation, with Nielsen's March 12, 2026 upfront planning guide putting streaming at 66.7% of ad-supported television time among adults aged 18 to 49. The row over a delayed February 2026 report turned on control of television money rather than methodology.

Who completes the form matters too. Asked in 2026 to grade artificial intelligence on a one-to-ten scale, PHD executives in Singapore reported productivity gains of 10% to 20%, below the figures in vendor-commissioned surveys.

Not the same as

score is one value describing one subject. IAS Quality Attention rates a single impression on a scale to 100, treating 65 and above as above-average attention. A scorecard combines several such values across several criteria.

dashboard displays whatever metrics are available, arranged to suit the viewer. A scorecard fixes them, weights them and produces a verdict. The Looker Studio component named scorecard is a dashboard element, not an evaluation framework.

benchmark supplies the reference values a scorecard scores against. The IAS Media Quality Report, whose 21st edition drew on more than 300 billion daily digital interactions and found made-for-advertising rates four times higher on mobile web display than desktop, and DoubleVerify's first-quarter 2026 report placing global authentic viewable rates at 74%, rank no suppliers at all.

certification is pass or fail against process standards, not a graded comparison. TAG certifies companies against programme requirements and the Media Rating Council accredits measurement methodology; a firm either holds the seal or does not.

Recent developments

Two 2026 findings pushed against the grading apparatus at once. The July work by TAG, the ANA and Fiducia put a matched dataset behind the argument that quality criteria now reward the inventory they were built to exclude. The IAB Spain supply-side platform guide, citing the ANA Programmatic Transparency Benchmark 2025, found 41% of programmatic investment reaching genuine, measurable, viewable impressions, with 26.1% consumed by platform fees and data costs, putting a documented price on the fee disclosure row.

Agentic buying, meanwhile, has arrived without an evaluation framework. Vendor-reported efficiency claims from agentic pilots accumulate faster than any standardised method for reading them across platforms and campaign types, leaving the new category scored against criteria built for the old one.

Timeline

  • 1930s: French firms adopt the tableau de bord, a fixed panel of operational indicators
  • 1990: The Nolan Norton Institute runs a twelve-company study of performance measurement in firms dependent on intangible assets
  • January to February 1992: Kaplan and Norton publish "The Balanced Scorecard - Measures That Drive Performance" in Harvard Business Review, pages 71 to 79
  • September to October 1993: "Putting the Balanced Scorecard to Work" documents the process for building one
  • January to February 1996: "Using the Balanced Scorecard as a Strategic Management System" reframes the tool as a management system, followed by the book of the same year
  • 2018: IHS Markit publishes an Internet of Things platform vendor scorecard ranking nine suppliers on market presence and momentum
  • December 2023: The ANA publishes its programmatic media supply chain transparency study, converting log-level findings into procurement criteria
  • August 27, 2024: Google announces the Store Quality scorecard in Merchant Center at its Think Retail event
  • January 2025: Looker Studio adds primary-field and sorting controls to its scorecard chart type
  • July 2025: Reddit introduces an Optimization Score and recommendations in Ads Manager
  • March 12, 2026: Nielsen's upfront planning guide puts streaming at 66.7% of ad-supported television time among adults 18 to 49
  • Late March 2026: The VAB disputes Nielsen's delayed February Gauge report and its methodology change
  • April 7, 2026: In-Store Marketplace and Catalyst Media Consulting identify four conflicting scorecards applied to in-store media
  • April 15, 2026: IAB Spain measures working media at 41% of programmatic investment, citing the ANA Programmatic Transparency Benchmark 2025
  • May 2026: CIMM and TVB publish local television currency guidelines structured as a vendor scorecard
  • July 2026: The 21st edition of the IAS Media Quality Report is released, and Kochava ships a Creative Scorecard widget
  • July 28, 2026: TAG, the ANA and Fiducia find AI-generated inventory grading as premium more than 70% of the time

Summary

Who. Advertisers, agencies and procurement teams build and run scorecards; vendors, publishers, exchanges and agencies are rated by them. Industry bodies including the ANA, CIMM, TVB, TAG and the IAB supply the criteria sets, and verification firms supply much of the underlying data. The format traces to Robert Kaplan and David Norton.

What. A fixed set of weighted criteria applied repeatedly to the same class of subject, producing a comparable number or grade. Variants include strategic scorecards, vendor and partner scorecards, media quality scorecards, and product features named after the concept.

When. The balanced scorecard was published in January 1992, analyst vendor scorecards were standard technology procurement by the 2010s, and programmatic criteria sets took shape after the ANA transparency study of December 2023.

Where. In procurement and agency reviews, supply path decisions, retail media and in-store measurement, television currency debates, and inside advertising platforms that ship scorecard features of their own.

Why. Scorecards decide which vendors keep budget, so their criteria become targets. Work published in July 2026 found synthetic inventory outscoring genuine supply on the standard media quality criteria, which makes how a scorecard is weighted, and who supplies its data, a commercial question rather than an administrative one.