Skip to content
All articles

The check was green locally and red in CI. Both were right.

A check reported 4.58:1 on my machine and 4.50:1 on the build server, for the same spot and with no change in between. The fault was not in the code it checked, but in the way it measured.

5 min read
  • Accessibility
  • WCAG
  • CI
  • Measurement

I have a check that measures colour contrast where a standard tool stays silent. Tools like axe need a solid background colour; if the text sits on a gradient, a translucent tint or an image, they report “unknown” and skip the position. On a dark site with glow effects and tinted cards, those are not edge cases. They are half the surface.

The check solves it by capturing every text position twice: once with the text, once without. The per-pixel difference tells it where a letter sits and what colour really lies behind it. That has worked for months. Until CI turned red.

Four point five zero against four point five

The message read 4.50:1 statt 4.5:1. So 4.4999-something, rounded to two places, on a comment inside a code block. The same check on my machine, same file, same commit: 4.58:1. No difference in the code, none in the content, none in the font size.

The first reflex is to touch the threshold. Four hundredths below the requirement smells like rounding, and a two-percent allowance would have turned the run green again immediately. That would have been the actual mistake: a check whose threshold you move as soon as it gets in the way only checks itself from then on.

Anti-aliasing had a vote

The check looked for the worst value across the pixels a letter fully covers. That sounds right, because the text colour is purest there. But what counts as “fully covered” is decided by anti-aliasing, and that differs per renderer. Windows and Linux set the same font at the same size with different coverage values.

On a solid background it makes no difference: the same colour lies behind every pixel, so which one you hit changes nothing. On a gradient, a different colour lies behind each. The reported value became a sample, and which sample it was got decided by a detail of font rendering.

Coverage now answers only the first question

The measurement holds two questions, and I had answered both in the same loop. First: is there text here at all, and at what opacity? Second: how dark is the background in the worst case? Only the first one needs anti-aliasing.

// First question: is there text? That is what covered pixels are for.
for (const pixel of pixelsInBox) {
  if (coverage[pixel] >= threshold) corePixels++;
}

// Second question: how bad does it get? Every background pixel in the
// text’s line boxes counts, whether or not a letter happens to hit it.
// This depends on no anti-aliasing at all.
if (corePixels) {
  for (const pixel of pixelsInBox) {
    const behind = backgroundAt(pixel);
    const expected = blend(textColour, behind, opacity);
    worst = Math.min(worst, contrast(expected, behind));
  }
}
check-contrast.mjs: the worst value comes from the background, not from the letter.

The second pass walks every pixel in the text’s line boxes and computes the nominal colour against the background that actually lies there. Whether a letter hits that pixel no longer matters. The check now measures the worst case instead of a random one.

What became visible afterwards

Three consecutive runs reported the same value three times, where 4.58, 4.61 and 4.59 had stood before. And two real violations stepped out of the noise:

  • Code blocks that scroll sideways carry a brightening strip at their edges as a hint. It sits exactly where text sits. At 20 percent white, an axis label in an architecture diagram fell to 4.10:1, well below the 4.5:1 of WCAG 1.4.3. At 12 percent the hint stays visible and the contrast stays in range.
  • One label was the only one of its kind sitting on a tinted panel, and it landed at 4.54:1, four hundredths above the requirement. It now has its own class, one step brighter; the fifteen labels on solid backgrounds stay as they were.

The weakest spot on the whole site has since been 4.80:1 under Windows and 4.76:1 under Linux. The remaining difference comes from line breaking, no longer from font rendering.

What I took away

A check is code too, and it can be wrong itself. Suspicion almost never falls on it: when it is green you believe it, and when it is red you go looking in the code it checked. That both answers were correct and still different left only one explanation: the question was badly put.

The practical rule I have applied since: if a measurement moves between two environments, the threshold is not too tight. The quantity is badly defined. Fix the definition first, then talk about thresholds.

Evidence

  • Portfolio repo, commit 60a46b9 of 8 August 2026 (the red CI run)
  • Portfolio repo, commit d0efe33 of 8 August 2026 (the rewritten method)
  • scripts/check-contrast.mjs, measures 1,302 text positions across two widths
  • Counter-check: three consecutive runs before and after the rewrite
  • Comparison value from the CI log of the same run under Linux