Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

There are all kinds of interesting thought experiments. What happens when a classifier innocently discovers that the best classification is by race? Do we care? How about if we remove race but it happens to discover that four features which are very strongly correlated to race are the best way to classify?


If that second scenario were to happen, then I think we should take a serious look at why that correlation is occurring rather than just throwing out the data because it's "racist". That we removed the classification and then it was re-discovered by other correlations really should suggest something. On the assumption that it wasn't engineered to be biased and was naturally arrived at by the algorithm itself, then that actually seems like an important data point, and could even be a nice litmus test of how we're addressing racial differences if the models evolve to be more positive over time.


There are ways to account for that. A model can be fit to race, and then you only predict "on top" of race (meaning residuals). You use that model, which is independent of race.


This only works if you're aware that race is even a factor. If you're not aware of the problematic factors, then you can't correct for them.


How do you know you have the correct model, and isn't making the system more racist instead of less?

Machine learning is also very opaque.


But if the racial factors aren't all the same, then that creates an incentive for people to lie about their race.

If you verify the race field, then now you're in the business of enforcing racial definitions.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: