Should you gate Shopify theme deployments on accessibility scan results?
Gate on regressions, not on absolute scores. A deploy gate that blocks every release containing any finding will be bypassed within a month; a gate that blocks only new failures introduced by the release keeps the store improving and the team honest. The difference is whether the gate compares against the last scan or against perfection.
Why absolute gates fail
The absolute gate is simple to set up: scan the staging theme, and if the scanner reports any error, the deploy cannot ship. The first week it feels like progress. Then the holiday sale theme ships with a countdown timer that fails color contrast, marketing needs it live by Friday, and someone discovers the bypass switch. After that, the gate exists on paper and the bypass exists in practice. Absolute gates fail because they treat a legacy finding from 2021 the same as a new finding from this morning's pull request.
The deeper issue is that absolute gates punish the teams doing the most work. The store with 400 legacy findings can never ship anything, while the brand-new store with three findings ships freely. The team maintaining the older store learns that the gate is an obstacle, not a tool, and they route around it. A gate that the team hates is a gate that gets disabled during the one release where it mattered most.
Gate on the diff instead
The regression gate compares the scan of the candidate theme against the scan of the live theme and blocks only on new failures. Same URL, same scanner, same config: if the new theme introduces a finding that the live theme does not have, the deploy stops. If the new theme carries the same findings the live theme already had, the deploy proceeds and the legacy issues stay in the backlog where they belong.
This works because it aligns the gate with the team's incentives. The developer who introduces a keyboard trap gets stopped immediately, when the fix is cheapest. The developer who ships a feature on a page with pre-existing issues is not punished for someone else's debt. Over time the diff gate does something the absolute gate never could: it ratchets the store toward zero findings, because every release can only remove issues, never add them.
Implementing it requires scan stability. The same theme scanned twice should produce the same findings, or the diff is noise. Pin the scanner version, scan the same viewport sizes, and exclude third-party content that changes between runs. If your scans are not deterministic, fix that before you build the gate; a flaky gate trains the team to distrust every gate.
What the gate should not check
Keep best-practice warnings out of the gate. A missing landmark role or a redundant link is worth fixing, but it should not stop a revenue release. The gate's job is to prevent harm, not to enforce taste. Every item in the gate should pass a simple test: would a real user be blocked or misled by this? If the answer is yes, it gates. If the answer is maybe, it goes to the backlog.
Also keep the gate away from content the deployment does not touch. If the release changes the product template, do not block it for a finding on the careers page. Scope the scan diff to the templates in the release. This is more work to configure, but it is the difference between a gate the team trusts and a gate the team blames for every delayed launch.
The override policy matters more than the gate
Someone will always need an override, and the override policy is what determines whether the gate survives. The override should be possible but uncomfortable: it requires a named approver, a written reason, and a ticket for the finding that was waived. The reason is not bureaucracy; it is a record. When the same finding gets waived three releases in a row, the record shows the gate is miscalibrated, and the team can adjust the rule instead of fighting it.
Review the waived findings monthly. Patterns in the overrides are the most honest feedback your gate will ever get. If every override is for the same rule, the rule is wrong for your store. If the overrides are spread across many rules, the team is using the override as a habit, and the approver needs to push back. A gate with a healthy override policy gets stronger every quarter. A gate with no override policy gets deleted the first time it blocks a launch.