home projects AccessAdvisor

AccessAdvisor.

An AI accessibility reviewer that reads pull requests, live URLs, and whole repositories for WCAG 2.1 violations — and the evaluation harness that proved its own headline feature was costing it accuracy.

2026 — present Python · FastAPI · Next.js LlamaIndex · ChromaDB · Claude See the code ↗
the problem

Accessibility bugs ship because audits happen after launch, by hand — long after the code was written. By then the fix is a ticket nobody prioritizes. The fix has to live where developers already are: the pull request.

the hard part

Detection is easy; trust is not. Ask an LLM about accessibility and it invents rules that sound right. So I measured it — and found my own retrieval pipeline was the thing costing me accuracy.

accessadvisor.exe — grounded review pipeline
input

PR diff · live URL · whole repo

retrieve

650 WCAG passagesChromaDB · LlamaIndex

reason

Claude structured tool-use

output

machine-parseable violation objects

deliver

inline comment at the exact line

three decisions that make it trustworthy

  1. Grounding in 650 spec passages. Findings retrieve from the official WCAG 2.1 spec (LlamaIndex + ChromaDB) before the model speaks, so it cites a real success criterion instead of inventing one. This was the original thesis — and the one I later had to test rather than assume.
  2. Structured tool-use, not prose. Claude returns machine-parseable violation objects through its tool-use API — rule, line, severity, fix — instead of a paragraph I'd have to parse with regex. Parsing prose is where this class of tool usually breaks.
  3. Inline at the exact line. A GitHub reviewer (PyGithub + NextAuth GitHub OAuth) fetches the PR diff and posts each fix as a comment on the offending line — meeting developers where they already work instead of in a dashboard they'd have to remember to open.

the finding: my own core feature was hurting recall

  1. I stopped trusting the demo. The tool looked right on the examples I'd hand-picked. That's not evidence. I built a 30-case evaluation harness covering all 78 WCAG 2.1 success criteria, with known-correct answers, so the question became measurable instead of vibes.
  2. The result contradicted the design. With retrieval on, it caught real violations 60% of the time. With retrieval off, 90%. The RAG pipeline — the thing the whole project was built around — was making the tool worse.
  3. The cause. A bug in the retrieval step was dropping the correct passage in roughly 1 of every 3 queries, so the model was reasoning from the wrong criterion while sounding just as confident. Grounding wasn't wrong in principle; my implementation was quietly losing the answer.
  4. Why it matters. Without the harness I'd have shipped a tool that felt rigorous and missed a third of what it should have caught. The measurement was worth more than the feature.

making it cheap enough to actually run

  1. Filter before you spend. Most files in a repo can't contain a WCAG violation. A pre-filter skips them before any model call — cutting 73.5% of files across 5 real codebases.
  2. Up to 98.5% cheaper. On a real 30-file review, filtering dropped analysis cost by up to 98.5% — the difference between a tool you run on every PR and one you run once and quietly turn off.

what it does

  1. Reviews a pull request. Fetches the diff, audits changed markup, posts inline comments at the violating lines.
  2. Audits a live URL. Point it at any page and it returns grounded findings with the spec citation for each.
  3. Scans a whole repository. Streams results as it walks the codebase instead of making you wait for one giant report.
90% of violations caught all 78 WCAG criteria tested 98.5% cheaper per review inline at the exact line
See the code ↗