Home / Monitoring

What should a Shopify store do when accessibility scan results change between runs?

Published October 5, 2026

Scan results that shift between runs are usually telling you something true about the store, not something broken about the scanner. Triage the change: real regression, dynamic content, or scan noise, in that order.

Do not trust the delta until you trust the baseline

A scan that reports twelve issues on Monday and nineteen on Wednesday triggers a reasonable question: did the store get worse, or did the scan get creative? Before investigating the new issues, verify the two runs are comparable. Same pages, same viewport, same scan configuration, same time of day. A surprising amount of scan drift comes from comparing runs that were never alike.

Check whether the page set changed. Scans that crawl the sitemap will pick up new pages the store published between runs, and new pages bring new issues. That is not drift; that is coverage working. Similarly, a run that scanned logged-out pages versus one that scanned with a session cookie can produce legitimately different results. The first debugging step is always reading the run configuration, not the issue list.

If the runs are truly comparable and the results still differ, you have a real signal. Something about the page or the scan environment moved. The next step is figuring out which.

Sort the change into three buckets

Bucket one is real regression: the store changed and the issues are genuinely new. Theme updates, app installs, new sections, and content edits all introduce accessibility issues at a steady rate. Confirm by checking the deploy log against the scan timestamp. If the new issues map to a recent change, you have your answer and your fix target.

Bucket two is dynamic content: the page renders differently between runs. Carousels that rotate, popups that fire on timers, A/B tests that swap variants, and personalization that changes content per visitor all make a page a moving target. A scan is a snapshot, and two snapshots of a moving page will differ. The fix here is scan hygiene: exclude or stabilize dynamic regions, run scans at consistent times, and pin A/B tests during the scan window if the differences matter.

Bucket three is scan noise: the tool itself is inconsistent. Some checkers are sensitive to timing, rendering race conditions, or third-party resources that load slowly. An issue that appears and disappears across runs with no store change is noise until proven otherwise. Track flaky issues separately from real ones, because mixing them poisons the trend line the team uses to judge progress.

Make the trend line honest

The value of repeated scanning is the trend: are we getting better or worse over time. A trend line built on noisy deltas is worse than no trend line, because it teaches the team to ignore the dashboard. Three practices keep it honest.

First, dedupe by root cause before counting. Ten instances of the same missing label from one shared snippet are one issue, not ten. Counting instances makes the trend jump every time a template change rolls out, which trains everyone to dismiss spikes. Second, separate new issues from carried-over issues in every report. The team needs to see what this week introduced versus what it inherited. Third, annotate the timeline with deploys and known dynamic-content events. A spike with an annotation is information. A spike without one is an argument.

When a single run looks anomalous, re-run before reacting. One weird scan is a data point; two in a row is a finding. The cost of a re-run is minutes. The cost of chasing a phantom regression is a sprint.

When drift means the scanner is wrong for the job

Persistent, unexplained drift across many pages is a tool-fit problem. Some scanners struggle with heavily scripted Shopify themes, single-page checkout flows, or pages where the critical content loads after the scan finishes. If you have ruled out store changes and dynamic content and the numbers still wobble, evaluate whether the scanner's rendering engine matches how real browsers see the page.

That does not necessarily mean replacing the tool. It can mean changing how you use it: scanning fewer, more stable pages with the automated tool and covering the dynamic surfaces with targeted manual checks. The automated scan is a smoke alarm, not a building inspector. Drift tells you where the alarm is unreliable, which is itself useful information about where to look by hand.