Microsoft Clarity described on September 30, 2026 how its engineers turned the vetting of new tracking-script versions into a reusable Claude Code skill, which according to the company returns a ship, investigate or block verdict in about half an hour.

In Short

Microsoft, which runs the free Clarity website analytics tool, wrote a detailed instruction file that lets an AI coding assistant test whether a new version of Clarity's tracking script is safe to release. Before that, an engineer wrote database queries by hand, and slow queries or misread numbers could hide real problems or invent fake ones, according to Microsoft. Now a single typed command produces a written report, and a person still makes the final decision on whether to ship.

What Microsoft published

The Clarity blog post is dated September 30, 2026, credited to Manu Nair, and titled "How We Use an AI Skill to Ship Clarity Releases with Confidence". According to Microsoft, the team built a skill that "queries billions of rows of telemetry, catches statistical traps" and writes a structured evaluation report for engineers deciding whether a build goes out.

It is an engineering write-up rather than a product release. Nothing in the post changes what site owners see in the Clarity dashboard. What it describes is the quality gate behind the script that those site owners install, and the way Microsoft has tried to make that gate repeatable.

The post appeared one day after the previous entry on the Clarity blog, which covered adjustable brand terms in AI Visibility, a feature PPC Land reported on September 29.

The problem of judging a script release

Every new version of the Clarity script raises one question, according to Microsoft: "is this version safe to ship?" The post follows it with a short verdict of its own. "That sounds simple. It isn't."

Clarity captures a wide set of signals, according to the company. These include the three Core Web Vitals - Interaction to Next Paint (INP), Cumulative Layout Shift (CLS) and Largest Contentful Paint (LCP) - along with session recording fidelity and script overhead. Before a release, Microsoft compares dozens of metrics across device type, recording mode, browser and operating system. The comparison spans hundreds of thousands of websites that run different script versions at the same time. A regression in any critical metric, the post says, could degrade the experience for sites that rely on Clarity.

Measurements of this kind, taken from live visits rather than lab runs, fall under what the industry calls real user monitoring. The approach produces large, noisy distributions, and that is where the old process struggled.

Historically the work was manual, according to Microsoft. An engineer wrote ad-hoc queries against the analytics data warehouse, looked over the results and made a judgment call. The queries were fragile because the warehouse enforces server-side timeouts, and certain aggregation patterns over billions of rows failed regularly. No systematic protocol existed for what to check, for how to read differences between populations sampled at unequal rates, or for deciding when an apparent regression was a statistical artifact.

How the skill is put together

A Claude Code skill, as the post defines it, is a reusable workflow described by a SKILL.md file in a repository under .claude/skills/, optionally accompanied by reference documents, templates, examples or scripts. Microsoft's skill is invoked with a single command, for example /clarity-evaluation 0.8.59 0.8.60-beta, which compares a production version with a beta.

The company stresses that the file is not code to be run as written. In its words, the skill "isn't just a script" but "a methodology document". It holds four kinds of content:

  • Domain knowledge - what each metric means, how to interpret it and which thresholds matter.
  • Query optimization rules - patterns that perform well in the warehouse and patterns that exceed its execution limits.
  • A phased evaluation protocol - data collection, anomaly investigation and final assessment.
  • Interpretation guardrails - rules on sampling artifacts, population bias and when an aggregate number can be trusted or not.

On each run, per the post, the agent follows seven steps. It parses the arguments and recognizes a beta evaluation; fetches the latest metric definitions from an internal repository, since those definitions change over time; generates a complete TypeScript script with the query optimizations built in; and runs that script against the warehouse, a job that takes roughly 20 to 30 minutes. It then interprets the output through a tiered framework, investigates surprising findings, and writes a report with a ship, investigate or block verdict.

The tiers run in a fixed order. Script overhead comes first (is the new version sending more data or blocking the main thread for longer?), followed by Web Vitals, then data completeness, such as whether end-of-session payloads arrive and whether session durations stay consistent. Errors and health form the fourth tier, covering crash patterns, fatal errors and resource limits. Engagement signals - clicks, rage clicks and intent - come last.

Size of the artifacts

According to Microsoft, the skill file runs to about 380 lines of Markdown. The TypeScript script it generates is 200 to 300 lines per run, and it is disposable: regenerated from scratch each time for the specific versions and date range, so that new metrics or query patterns can be absorbed without editing generated code. The output is written to a file named in the pattern version-evaluation-[version].md.

One inconsistency is visible in the post. The text describes a SKILL.md file, while the architecture diagram labels the methodology file .claude/skills/clarity-evaluation.md. Microsoft does not explain the difference, and it may simply reflect the diagram's shorthand.

The post also does not name the data warehouse. The function quantilesTDigest and the PREWHERE clause that it cites match the syntax of ClickHouse, but Microsoft does not state that, and the identification here is PPC Land's own observation.

The timeouts and the quantile fix

The first attempts failed on query timeouts. The underlying table holds billions of rows, according to the post, and the warehouse limits each query's run time (the architecture diagram gives the limit as roughly 60 seconds). Naive aggregations simply did not finish.

The decisive change involved percentile calculation. Microsoft says the warehouse's default quantile() function relies on reservoir sampling, which it characterizes as using memory proportional to the number of rows per aggregation state. Its example query requested seven separate percentile calculations: P50, P90 and P99 of data size, P50 and P90 of thread-blocked time, and P75 of INP and of CLS. Run across many GROUP BY groups, the company says, memory use "explodes" and the cluster's budget was exceeded even with sampling.

The replacement was a family of T-Digest-based functions, which Microsoft describes as using constant memory per state. They also accept a batched form. A call such as quantilesTDigest(0.5, 0.9, 0.99)(DataSize) returns an array holding P50, P90 and P99 from a single digest - "three percentiles for the memory cost of one", in the post's phrasing. According to Microsoft, that single change took the evaluation from a 100% timeout rate to a 100% success rate. The rule now sits in the skill in plain terms: "Always use T-Digest quantiles, never reservoir sampling."

Three further optimizations made it into the skill, according to the post:

  1. Partition pruning. Putting the partition key next to the version filters in the PREWHERE clause lets the engine skip irrelevant partitions entirely and reduces disk reads.
  2. Client-side domain extraction. Grouping by parsed URL domains inside the database timed out even at small sample rates, and the post calls GROUP BY domain(Url) a "cluster-killing operation". Fetching raw URLs and parsing domains in application code proved, in Microsoft's words, orders of magnitude faster.
  3. Asymmetric sampling. The larger population, production, is sampled at a lower rate than the smaller beta population. Query times stay manageable while statistical power is kept where the comparison needs it.

The third item is the setup for the most instructive episode in the post.

When a tenfold increase was not real

During one evaluation, according to Microsoft, queries showed a client-side diagnostic flag running 10 times higher on the beta than on production. It looked like a serious regression: something in the new code appeared to trigger a warning condition ten times as often.

The sampling rates differed by roughly the same factor. Production was sampled at about one-tenth of the beta rate, and raw counts drawn from two differently sampled populations cannot be compared directly. A metric that seems ten times worse may only reflect ten times as much data.

To test that, the team identified the websites contributing most to the flag and queried them at full resolution, with no sampling, on both versions. The rates matched to within hundredths of a percentage point. The apparent regression was, in the post's words, "a pure sampling artifact".

The lesson became a dedicated phase of the skill, which the post labels Phase 2.5. Any metric showing a surprising result - more than a 20% change in either direction - triggers a drill-down to the individual site level. The rule reads: "If rates differ at the site level, it's real." If they match, the aggregate gap is an artifact. Microsoft made the check bidirectional, so that artifacts producing false improvements or false stability get the same scrutiny as false regressions.

Comparing like with like

A second guardrail concerns operating mode. Clarity runs either in analytics-only mode or with session recording, and recording adds DOM-mutation capture that naturally generates more data. Some metrics therefore differ by design between the two. Microsoft's skill stratifies every comparison by mode and forbids comparing across strata - the post cites "never compare Playback=0 vs Playback=1" as an explicit rule.

When the two modes disagree, the skill trusts the recording-enabled cohort, because it covers a larger and more representative population, according to the post.

The 0.8.60-beta numbers

For its recent evaluation of 0.8.60-beta against 0.8.59, Microsoft gave the following results.

SignalResult per Microsoft
Fatal errorsZero
Data transmitted (recording-enabled)4-30% reduction
Main-thread long tasks (mobile)20-40% reduction
Cumulative Layout Shift12-30% improvement
Largest Contentful Paint5-7% improvement
One flag with an apparent 10x increaseSampling artifact confirmed by deep-dive

The verdict printed in the post is one word: "Ship."

Microsoft also offers a counterfactual. Without the skill, the same analysis would have taken an engineer several days of writing queries and debugging timeouts, and the sampling artifact might have been flagged as a real regression and delayed the release. With the skill, the company says, a "grounded, auditable verdict" arrived in under 30 minutes. The days-long figure is Microsoft's estimate of a hypothetical, not a measured baseline from earlier evaluations.

Five lessons Microsoft draws

The post closes its main section with five working principles. The first is to encode methodology rather than procedures, so that the agent can adapt when a column is missing or a query times out. The second is to capture operational knowledge, since every optimization rule "was learned through failure" - the switch to T-Digest took days of debugging, and the sampling protocol came after the team nearly shipped a false alarm.

Third come guardrails against misinterpretation. The skill names the signals that block a ship decision: fatal log counts, drops in the end-of-session unload signal, and Web Vitals regressions. Fourth, the file is meant to improve itself, with each new timeout pattern or missing interpretation rule written back into it after a run.

Fifth is verification. "The skill produces a report, not a decision," the post says, and "the ship/no-ship decision remains human."

What comes next

Microsoft lists three extensions under consideration. Automated regression bisection would diff source code between versions after a regression is confirmed and map the problem to specific changes. Continuous monitoring would run the evaluation on a schedule to catch problems early in the beta cycle. Cross-version trending would build historical baselines, so that gradual metric drift can be detected in addition to version-over-version change.

What the post leaves out

Several details are absent. The post gives no count of evaluations run with the skill, no sample sizes or date range for the 0.8.60-beta comparison, and no baseline values behind the percentage ranges, some of which are wide (data transmitted moved by 4% to 30%, depending on the segment). It does not name the Claude model used, the compute or token cost of a run, or how often the agent's interpretation has been overturned by an engineer. All performance and time-saving figures are Microsoft's own, published on its own blog, and no outside party has verified them.

Why this matters for the marketing community

Clarity is a free analytics product, and its script sits on a large number of pages that carry advertising. Microsoft Advertising began requiring Clarity on third-party publisher placements in November 2025, with non-compliant traffic filtered as nonbillable, according to PPC Land's reporting at the time. A script that publishers are obliged to install makes its own release quality a matter of wider interest than a typical analytics tag. The post indicates the performance lines Microsoft examines before shipping - blocked main-thread time, layout shift, paint timing, end-of-session data delivery - and that disclosure shows what a beta is judged on.

The post also fits a longer pattern in which Clarity has tied itself closely to AI tooling. Microsoft added Copilot to Clarity on April 15, 2025, and released a Model Context Protocol server on June 4, 2025 that lets analysts query Clarity metrics in natural language through Claude and other compatible clients. AI channel groups followed on August 29, 2025, separating traffic from ChatGPT, Claude, Gemini, Copilot and Perplexity. Clarity research published in December 2025 found that AI-referred traffic had grown 155% in eight months.

The AI Visibility section has since filled out quickly: Bot Activity on January 21, 2026, Citations in general availability on May 13, 2026, robots.txt violation detection on June 23, 2026, Topic Insights in July, and a split of AI citations by brand in August. A cadence of that kind depends on a dependable evaluation process underneath, and the September 30 post offers a view of one such process, though it does not itself tie the skill to any particular feature.

A second thread concerns the form of the tool. Skills are spreading through marketing technology as a way to package repeatable expertise for AI agents. Typeface described skills as reusable packages agents apply across campaigns, and the same PPC Land report recorded Salesforce and Anthropic placing 37 prebuilt sales skills inside Claude for pilot customers on August 26, 2026. Clarity's example sits at the engineering end of that spectrum, with a statistical protocol instead of a brand rulebook. The shared feature is the one the post states outright: the value lies in the written methodology, and the model only carries it out.

The "sampling artifact" case also touches anything built on sampled analytics. When two groups are sampled at different rates, a raw count comparison can mislead, and Microsoft's remedy - confirm at a finer grain before acting - is a method readers of any analytics dashboard can recognize, whatever the platform.

Timeline

Summary

Who: Microsoft Clarity, the free behavioral analytics product, through a blog post credited to Manu Nair. The affected parties are Clarity's engineers and the site owners and publishers whose pages run the Clarity script.

What: An account of a reusable Claude Code skill that generates warehouse queries, runs them, applies interpretation rules and writes a ship, investigate or block report for each beta version of the Clarity script. Key technical points: a switch to T-Digest quantile functions that, per Microsoft, moved the timeout rate from 100% to a 100% success rate; a drill-down protocol that exposed a 10x flag increase as a sampling artifact; and stratification by recording mode.

When: The post is dated September 30, 2026. The worked example compares script versions 0.8.59 and 0.8.60-beta, and a full run takes about 20 to 30 minutes.

Where: On the Microsoft Clarity blog, with the work running against Microsoft's internal analytics data warehouse, which holds billions of rows of telemetry from sites using Clarity.

Why: Manual evaluation relied on fragile ad-hoc queries, timed out on large tables and offered no fixed method for sampling differences between versions. Microsoft encoded its methodology in a skill to make the process repeatable and faster, while keeping the final ship decision with a human engineer.