Why did nobody tell me my training data for my image model was garbage?
I built a small model to sort photos of my neighbor's chickens and it kept calling every brown bird a hawk. After 3 weeks of tuning and about 40 test runs I finally opened the folder and saw half the hawk pictures were actually blurry shots of a rusted tractor. So the model was never confused, it was just copying my messy labels. Now I make a contact sheet of every image before it goes in. What do you all use to spot bad labels before training?
Hold on, half the hawk pics were a rusted tractor? That is genuinely amazing. So for three weeks you were tweaking settings and rerunning tests while the model was just being a good student and repeating your nonsense back to you. Forty runs, and the whole time the answer was sitting in the folder looking like farm equipment. The mental picture of you staring at a blurry tractor finally going "oh no" is killing me. That is way funnier than any actual model bug.