home experience Visual distractions
Visual distractions.
A custom YOLOv11 detector that scores how distracting a railroad crossing is from Street View imagery — so the crossings most likely to pull a driver's eyes off the tracks can be found without visiting them.
Grade-crossing collisions are often about attention, not visibility — a billboard or a cluttered sight line steals the glance that would have caught the train. Auditing crossings by hand doesn't scale to a rail network.
I led the migration from YOLOv5 to YOLOv11, stood up the training infrastructure myself (GitHub, Drive, Colab), and owned model training across 15 experiments. Then I chased down two separate bugs that were quietly capping the model's accuracy.
multi-source datasetfixed-seed splits
Roboflow annotation
YOLOv11s from COCO weights
mAP50 across 15 experiments
per-crossing distraction risk
the numbers
the 0.42 ceiling nobody could explain
- The symptom. Three independently labeled datasets — with a 40× difference in annotation density between them — all plateaued at the exact same 0.42 mAP50 ceiling. Three very different datasets, one identical number.
- The obvious hypothesis (and why it was wrong). A 40× density gap points straight at the labels. But if sparse vs. dense annotation were the cause, the three datasets should not have hit an identical ceiling — they'd fail differently, not identically.
- The finding. The export pipeline was silently downsampling every image to 640×640 before training, then re-exporting at native resolution downstream — throwing away the exact detail that would have separated the datasets. All three were bottlenecked on resolution, not labels.
- The fix. Re-exported the pipeline at native resolution. Overall mAP50 rose 14%; thin-object AP — the small, distant objects a 640px export loses first — rose 111%.
a separate 28% regression
- The symptom. After folding in a second labeled source, mAP50 fell from 0.506 to 0.361 — a 28% regression on what should have been more data.
- Ablations first. Before touching the data, I ran single-variable ablations across backbone freezing, class weighting, learning rate, and input resolution — ruling out training hyperparameters as the cause one at a time instead of guessing.
- The finding. A 2.3× instance-per-image mismatch between sources. The two datasets had been labeled under incompatible protocols — one annotated exhaustively, the other sparsely — so the sparse labels taught the model that real objects were background.
- Why it matters. The fix wasn't a learning rate. It was recognizing that label semantics have to match before data can be pooled — so I authored the team's labeling standard, turning a one-off fix into a rule the next dataset can't violate.
also shipped
- Training infrastructure from scratch. GitHub + Google Drive + Colab, so runs were reproducible by anyone on the team instead of living on one laptop.
- A source-pool dataset system with fixed-seed splits. Same split every run — otherwise run-to-run comparisons measure luck, not the model.
- Fixed a silent Drive path-resolution failure that had been dropping training data without erroring. Nobody knew training was running short.