home experience Visual distractions

Visual distractions.

A custom YOLOv11 detector that scores how distracting a railroad crossing is from Street View imagery — so the crossings most likely to pull a driver's eyes off the tracks can be found without visiting them.

summer 2026 · externship Collaborative Solutions — MBS program YOLOv11 · PyTorch · Roboflow See the code ↗
the problem

Grade-crossing collisions are often about attention, not visibility — a billboard or a cluttered sight line steals the glance that would have caught the train. Auditing crossings by hand doesn't scale to a rail network.

my role

I led the migration from YOLOv5 to YOLOv11, stood up the training infrastructure myself (GitHub, Drive, Colab), and owned model training across 15 experiments. Then I chased down two separate bugs that were quietly capping the model's accuracy.

training.py — the rebuilt pipeline
source pool

multi-source datasetfixed-seed splits

label

Roboflow annotation

transfer

YOLOv11s from COCO weights

evaluate

mAP50 across 15 experiments

score

per-crossing distraction risk

the numbers

+33%mAP50 vs. the prior benchmark0.406 → 0.606, 15 experiments
+111%thin-object AP after the fixsmall/distant objects gained most
40×annotation-density gap that wasn't the causeceiling was identical anyway
2.3×instance-per-image mismatch foundroot cause of a separate regression

the 0.42 ceiling nobody could explain

  1. The symptom. Three independently labeled datasets — with a 40× difference in annotation density between them — all plateaued at the exact same 0.42 mAP50 ceiling. Three very different datasets, one identical number.
  2. The obvious hypothesis (and why it was wrong). A 40× density gap points straight at the labels. But if sparse vs. dense annotation were the cause, the three datasets should not have hit an identical ceiling — they'd fail differently, not identically.
  3. The finding. The export pipeline was silently downsampling every image to 640×640 before training, then re-exporting at native resolution downstream — throwing away the exact detail that would have separated the datasets. All three were bottlenecked on resolution, not labels.
  4. The fix. Re-exported the pipeline at native resolution. Overall mAP50 rose 14%; thin-object AP — the small, distant objects a 640px export loses first — rose 111%.

a separate 28% regression

  1. The symptom. After folding in a second labeled source, mAP50 fell from 0.506 to 0.361 — a 28% regression on what should have been more data.
  2. Ablations first. Before touching the data, I ran single-variable ablations across backbone freezing, class weighting, learning rate, and input resolution — ruling out training hyperparameters as the cause one at a time instead of guessing.
  3. The finding. A 2.3× instance-per-image mismatch between sources. The two datasets had been labeled under incompatible protocols — one annotated exhaustively, the other sparsely — so the sparse labels taught the model that real objects were background.
  4. Why it matters. The fix wasn't a learning rate. It was recognizing that label semantics have to match before data can be pooled — so I authored the team's labeling standard, turning a one-off fix into a rule the next dataset can't violate.

also shipped

  1. Training infrastructure from scratch. GitHub + Google Drive + Colab, so runs were reproducible by anyone on the team instead of living on one laptop.
  2. A source-pool dataset system with fixed-seed splits. Same split every run — otherwise run-to-run comparisons measure luck, not the model.
  3. Fixed a silent Drive path-resolution failure that had been dropping training data without erroring. Nobody knew training was running short.
YOLOv5 → YOLOv11 +111% thin-object AP 0.42 ceiling, root-caused 28% regression, diagnosed
See the code ↗