AccessAdvisor.
An AI accessibility reviewer that reads pull requests, live URLs, and whole repositories for WCAG 2.1 violations — and the evaluation harness that proved its own headline feature was costing it accuracy.
Accessibility bugs ship because audits happen after launch, by hand — long after the code was written. By then the fix is a ticket nobody prioritizes. The fix has to live where developers already are: the pull request.
Detection is easy; trust is not. Ask an LLM about accessibility and it invents rules that sound right. So I measured it — and found my own retrieval pipeline was the thing costing me accuracy.
PR diff · live URL · whole repo
650 WCAG passagesChromaDB · LlamaIndex
Claude structured tool-use
machine-parseable violation objects
inline comment at the exact line
three decisions that make it trustworthy
- Grounding in 650 spec passages. Findings retrieve from the official WCAG 2.1 spec (LlamaIndex + ChromaDB) before the model speaks, so it cites a real success criterion instead of inventing one. This was the original thesis — and the one I later had to test rather than assume.
- Structured tool-use, not prose. Claude returns machine-parseable violation objects through its tool-use API — rule, line, severity, fix — instead of a paragraph I'd have to parse with regex. Parsing prose is where this class of tool usually breaks.
- Inline at the exact line. A GitHub reviewer (PyGithub + NextAuth GitHub OAuth) fetches the PR diff and posts each fix as a comment on the offending line — meeting developers where they already work instead of in a dashboard they'd have to remember to open.
the finding: my own core feature was hurting recall
- I stopped trusting the demo. The tool looked right on the examples I'd hand-picked. That's not evidence. I built a 30-case evaluation harness covering all 78 WCAG 2.1 success criteria, with known-correct answers, so the question became measurable instead of vibes.
- The result contradicted the design. With retrieval on, it caught real violations 60% of the time. With retrieval off, 90%. The RAG pipeline — the thing the whole project was built around — was making the tool worse.
- The cause. A bug in the retrieval step was dropping the correct passage in roughly 1 of every 3 queries, so the model was reasoning from the wrong criterion while sounding just as confident. Grounding wasn't wrong in principle; my implementation was quietly losing the answer.
- Why it matters. Without the harness I'd have shipped a tool that felt rigorous and missed a third of what it should have caught. The measurement was worth more than the feature.
making it cheap enough to actually run
- Filter before you spend. Most files in a repo can't contain a WCAG violation. A pre-filter skips them before any model call — cutting 73.5% of files across 5 real codebases.
- Up to 98.5% cheaper. On a real 30-file review, filtering dropped analysis cost by up to 98.5% — the difference between a tool you run on every PR and one you run once and quietly turn off.
what it does
- Reviews a pull request. Fetches the diff, audits changed markup, posts inline comments at the violating lines.
- Audits a live URL. Point it at any page and it returns grounded findings with the spec citation for each.
- Scans a whole repository. Streams results as it walks the codebase instead of making you wait for one giant report.