fairness
Arrested by a false match
It is 2020 in Detroit, and a man has spent thirty hours in custody for a crime he did not commit, arrested on the strength of a face-recognition "match" from grainy surveillance footage — a system whose error rates, measured independently, run an order of magnitude higher for dark-skinned faces, deployed by a department treating the match as identification rather than lead. Audit the whole sociotechnical pipeline: the demographic error structure, the compounding of a low base rate with automation bias in the human reviewer, the absent thresholds and absent protocols. Then specify the deployment policy — if any exists — under which such a system is defensible in policing. Get it wrong and the technology scales the misidentification of exactly the communities already most policed, with a scientific veneer.
Who this problem belongs to
The two figures whose methods fit it best, out of 31 in contention.
Gebru co-authored Gender Shades (2018, with Joy Buolamwini), the study that first rigorously quantified exactly this failure: commercial face-recognition systems carrying dramatically higher error rates for darker-skinned faces, especially darker-skinned women, than for lighter-skinned men, measured across standardized benchmarks rather than anecdote. That audit methodology — stratifying accuracy by Fitzpatrick skin type and gender, forcing vendors to confront numbers they had never published — is precisely the demographic error-rate analysis this Detroit case requires, and her subsequent advocacy for moratoria on police use of face recognition engaged directly with the deployment-policy question the problem poses. No other carrier on this list did the specific technical and political work that made this exact failure mode a matter of public record and policy debate.
Kanade's career at Carnegie Mellon's Robotics Institute is built substantially on face detection and recognition as a technical problem: his group's work on facial feature tracking, optical flow, and 3D face modeling from the 1990s onward established core techniques still embedded in modern pipelines, and he understands intimately how such systems degrade under pose, lighting, and image quality — exactly the grainy-surveillance-footage conditions in the Detroit arrest. He did not specifically study demographic disparity in error rates, which was documented later and by other researchers, so the audit's fairness dimension is not his direct contribution, but the underlying vision engineering, why a system says 'match' with unwarranted confidence on degraded imagery, is squarely his territory.
In the mind map
The same ideas, as concepts rather than history — in John's ML knowledge map.
31 figures are scored on this problem. Draw it in a battle to see where you land.