SearcharxivSearch

arXiv subjects

Amy N. Yates

Publications and source records attributed to Amy N. Yates.

3 recordsLinked to original sources

Unlocking the power of partnership: How humans and machines can work together to improve face recognition

Human review of consequential decisions by face recognition algorithms creates a collaborative human-machine system. We establish the circumstances under which combining human and machine face identification decisions improves accuracy. Using data from expert and non-expert face identifiers, we show that the benefits of human-human and human-machine collaborations increase as the difference in baseline accuracy between collaborators decreases. This rule holds across a wide range of baseline abilities, from novices to professional forensic face examiners. An important consequence of the rule is that people who are substantially less accurate than the machine, can actually improve decision accuracy when they collaborate with the machine. In a group of individual people collaborating with a machine, "intelligent human-machine fusion" was implemented by selecting people with the potential to increase collaborative accuracy. Performance with intelligent human-machine fusion was more accurate than either the machine operating alone or fusing all humans with the machine. Eliminating the machine from consideration yielded less predictable results, with average performance at or below intelligent human-machine collaboration. However, intelligent human-machine fusion was consistently more effective at minimizing the impact of low-performing humans on accuracy. The results demonstrate a meaningful role for both humans and machines in assuring accurate face identification.

cs.CV

Human-Machine Comparison for Cross-Race Face Verification: Race Bias at the Upper Limits of Performance?

Face recognition algorithms perform more accurately than humans in some cases, though humans and machines both show race-based accuracy differences. As algorithms continue to improve, it is important to continually assess their race bias relative to humans. We constructed a challenging test of 'cross-race' face verification and used it to compare humans and two state-of-the-art face recognition systems. Pairs of same- and different-identity faces of White and Black individuals were selected to be difficult for humans and an open-source implementation of the ArcFace face recognition algorithm from 2019 (5). Human participants (54 Black; 51 White) judged whether face pairs showed the same identity or different identities on a 7-point Likert-type scale. Two top-performing face recognition systems from the Face Recognition Vendor Test-ongoing performed the same test (7). By design, the test proved challenging for humans as a group, who performed above chance, but far less than perfect. Both state-of-the-art face recognition systems scored perfectly (no errors), consequently with equal accuracy for both races. We conclude that state-of-the-art systems for identity verification between two frontal face images of Black and White individuals can surpass the general population. Whether this result generalizes to challenging in-the-wild images is a pressing concern for deploying face recognition systems in unconstrained environments.

cs.CV

Face Identification Proficiency Test Designed Using Item Response Theory

Measures of face-identification proficiency are essential to ensure accurate and consistent performance by professional forensic face examiners and others who perform face-identification tasks in applied scenarios. Current proficiency tests rely on static sets of stimulus items, and so, cannot be administered validly to the same individual multiple times. To create a proficiency test, a large number of items of "known" difficulty must be assembled. Multiple tests of equal difficulty can be constructed then using subsets of items. We introduce the Triad Identity Matching (TIM) test and evaluate it using Item Response Theory (IRT). Participants view face-image "triads" (N=225) (two images of one identity, one image of a different identity) and select the different identity. In Experiment 1, university students (N=197) showed wide-ranging accuracy on the TIM test, and IRT modeling demonstrated that the TIM items span various difficulty levels. In Experiment 2, we used IRT-based item metrics to partition the test into subsets of specific difficulties. Simulations showed that subsets of the TIM items yielded reliable estimates of subject ability. In Experiments 3a and 3b, we found that the student-derived IRT model reliably evaluated the ability of non-student participants and that ability generalized across different test sessions. In Experiment 3c, we show that TIM test performance correlates with other common face-recognition tests. In summary, the TIM test provides a starting point for developing a framework that is flexible and calibrated to measure proficiency across various ability levels (e.g., professionals or populations with face-processing deficits).

cs.CV