SearcharxivSearch

arXiv subjects

Maria Cuellar

Publications and source records attributed to Maria Cuellar.

11 recordsLinked to original sources

Accuracy and Fairness of Facial Recognition Technology in Low-Quality Police Images: An Experiment With Synthetic Faces

Facial recognition technology (FRT) is increasingly used in criminal investigations, yet most evaluations of its accuracy rely on high-quality images, unlike those often encountered by law enforcement. This study examines how five common forms of image degradation--contrast, brightness, motion blur, pose shift, and resolution--affect FRT accuracy and fairness across demographic groups. Using synthetic faces generated by StyleGAN3 and labeled with FairFace, we simulate degraded images and evaluate performance using Deepface with ArcFace loss in 1:n identification tasks. We perform an experiment and find that false positive rates peak near baseline image quality, while false negatives increase as degradation intensifies--especially with blur and low resolution. Error rates are consistently higher for women and Black individuals, with Black females most affected. These disparities raise concerns about fairness and reliability when FRT is used in real-world investigative contexts. Nevertheless, even under the most challenging conditions and for the most affected subgroups, FRT accuracy remains substantially higher than that of many traditional forensic methods. This suggests that, if appropriately validated and regulated, FRT should be considered a valuable investigative tool. However, algorithmic accuracy alone is not sufficient: we must also evaluate how FRT is used in practice, including user-driven data manipulation. Such cases underscore the need for transparency and oversight in FRT deployment to ensure both fairness and forensic validity.

cs.CV

The Prosecutor's Fallacy and Expert Testimony: A Modern Take Using Likelihood Ratios

Forensic examiners and attorneys need to know how to express evidence in favor or against a prosecutor's hypothesis in a way that avoids the prosecutor's fallacy and follows the modern reporting standards for forensic evidence. This article delves into the inherent conflict between legal and scientific principles, exacerbated by the prevalence of alternative facts in contemporary discourse. Courts grapple with contradictory expert testimonies, leading to a surge in erroneous rulings based on flawed amicus briefs and testimonies, notably the persistent prosecutor's fallacy. The piece underscores the necessity for legal practitioners to navigate this fallacy within the modern forensic science framework, emphasizing the importance of reporting likelihood ratios (LRs) over posterior probabilities. Recognizing the challenge of lay comprehension of LRs, the article calls for updated recommendations to mitigate the prosecutor's fallacy. Its contribution lies in providing a detailed analysis of the fallacy using LRs and advocating for a sound interpretation of evidence. Illustrated through a modified real case, this article serves as a valuable guide for legal professionals, offering insights into avoiding fallacious reasoning in forensic evidence assessment.

stat.AP

How Often are Fingerprints Repeated in the Population? Expanding on Evidence from AI With the Birthday Paradox

The assumption of fingerprint uniqueness is foundational in forensic science and central to criminal identification practices. However, empirical evidence supporting this assumption is limited, and recent findings from artificial intelligence challenge its validity. This paper uses a probabilistic approach to examine whether fingerprint patterns remain unique across large populations. We do this by drawing on Francis Galton's 1892 argument and applying the birthday paradox to estimate the probability of fingerprint repetition. Our findings indicate that there is a 50\% probability of coincidental fingerprint matches in populations of 14 million, rising to near certainty at 40 million, which contradicts the traditional view of fingerprints as unique identifiers. We introduce the concept of a Random Overlap Probability (ROP) to assess the likelihood of fingerprint repetition within specific population sizes. We recommend a shift toward probabilistic models for fingerprint comparisons that account for the likelihood of pattern repetition. This approach could strengthen the reliability and fairness of fingerprint comparisons in the criminal justice system.

stat.AP

Statistical Issues in the Diagnosis of Shaken Baby Syndrome/Abusive Head Trauma

The diagnosis of Shaken Baby Syndrome/Abusive Head Trauma (SBS/AHT) is fraught with controversy due to critical statistical deficiencies in the data underpinning these diagnoses. This paper examines the reliability and scientific foundation of SBS/AHT through a statistical lens, highlighting the lack of independently verified ground truth, contextual biases, data circularity, and diagnostic heterogeneity. These issues render current methodologies inadequate and complicate evaluations of diagnostic accuracy, particularly when legal determinations are integrated into medical assessments. Without empirical evidence validating the specificity of symptoms like subdural hematoma, retinal hemorrhage, and brain swelling, the diagnosis remains untested and its foundational validity unproven. We recommend that physicians focus on reporting observed clinical signs and avoid making determinations of abuse, which should remain within the legal domain. Addressing these challenges requires comprehensive, high-quality data collection encompassing contextual, medical, and legal information to evaluate the accuracy, repeatability, and reproducibility of SBS/AHT diagnoses. These efforts are essential to protect vulnerable children while ensuring fairness and accuracy in legal proceedings involving allegations of abuse.

stat.AP

The Neglected Error: False Negatives and the Case for Validating Eliminations

This article examines the overlooked risk of false negative errors arising from eliminations in forensic firearm comparisons. While recent reforms in forensic science have focused on reducing false positives, eliminations--often based on class characteristics or intuitive judgments--receive little empirical scrutiny despite their potential to exclude true sources. In cases involving a closed pool of suspects, eliminations can function as de facto identifications, introducing serious risk of error. A review of existing validity studies reveals that many report only false positive rates, failing to provide a complete assessment of method accuracy. This asymmetry is reinforced by professional guidelines, such as those from AFTE, and echoed in major government reports, including those from NAS and PCAST. The article argues that eliminations, like identifications, must be validated through rigorous testing and reported with transparent error rates. It further cautions against the use of "common sense" eliminations in the absence of empirical support and highlights the dangers of contextual bias when examiners are aware of investigative constraints. Five policy recommendations are proposed to improve the scientific treatment and legal interpretation of eliminations, including balanced reporting of false positive and false negative rates, validation of intuitive judgments, and clear warnings against using eliminations to infer guilt in closed-pool scenarios. Without reform, eliminations will continue to escape scrutiny, perpetuating unmeasured error and undermining the integrity of forensic conclusions.

stat.AP

Mediated probabilities of causation

We propose a set of causal estimands that we call the "mediated probabilities of causation." These estimands quantify the probabilities that an observed negative outcome was induced via a mediating pathway versus a direct pathway in a stylized setting involving a binary exposure or intervention, a single binary mediator, and a binary outcome. We outline a set of conditions sufficient to identify these effects given observed data, and propose a doubly-robust projection based estimation strategy that allows for the use of flexible non-parametric and machine learning methods for estimation. We argue that these effects may be more relevant than the probability of causation, particularly in settings where we observe both some negative outcome and negative mediating event, and we wish to distinguish between settings where the outcome was induced via the exposure inducing the mediator versus the exposure inducing the outcome directly. We motivate these estimands by discussing applications to legal and medical questions of causal attribution.

stat.ME

Methodological Problems in Every Black-Box Study of Forensic Firearm Comparisons

Reviews conducted by the National Academy of Sciences (2009) and the President's Council of Advisors on Science and Technology (2016) concluded that the field of forensic firearm comparisons has not been demonstrated to be scientifically valid. Scientific validity requires adequately designed studies of firearm examiner performance in terms of accuracy, repeatability, and reproducibility. Researchers have performed ``black-box'' studies with the goal of estimating these performance measures. As statisticians with expertise in experimental design, we conducted a literature search of such studies to date and then evaluated the design and statistical analysis methods used in each study. Our conclusion is that all studies in our literature search have methodological flaws that are so grave that they render the studies invalid, that is, incapable of establishing scientific validity of the field of firearms examination. Notably, error rates among firearms examiners, both collectively and individually, remain unknown. Therefore, statements about the common origin of bullets or cartridge cases that are based on examination of ``individual" characteristics do not have a scientific basis. We provide some recommendations for the design and analysis of future studies.

stat.AP

An algorithm for forensic toolmark comparisons

Forensic toolmark analysis traditionally relies on subjective human judgment, leading to inconsistencies and lack of transparency. The multitude of variables, including angles and directions of mark generation, further complicates comparisons. To address this, we first generate a dataset of 3D toolmarks from various angles and directions using consecutively manufactured slotted screwdrivers. By using PAM clustering, we find that there is clustering by tool rather than angle or direction. Using Known Match and Known Non-Match densities, we establish thresholds for classification. Fitting Beta distributions to the densities, we allow for the derivation of likelihood ratios for new toolmark pairs. With a cross-validated sensitivity of 98% and specificity of 96%, our approach enhances the reliability of toolmark analysis. This approach is applicable to slotted screwdrivers, and for screwdrivers that are made with a similar production method. With data collection of other tools and factors, it could be applied to compare toolmarks of other types. This empirically trained, open-source solution offers forensic examiners a standardized means to objectively compare toolmarks, potentially decreasing the number of miscarriages of justice in the legal system.

cs.CR

Commentary on Guyll et al. (2023): Misuse of Statistical Method Results in Highly Biased Interpretation of Forensic Evidence

Since the National Academy of Sciences released their report outlining paths for improving reliability, standards, and policies in the forensic sciences NAS (2009), there has been heightened interest in evaluating and improving the scientific validity within forensic science disciplines. Guyll et al. (2023) seek to evaluate the validity of forensic cartridge-case comparisons. However, they make a serious statistical error that leads to highly inflated claims about the probability that a cartridge case from a crime scene was fired from a reference gun, typically a gun found in the possession of a defendant. It is urgent to address this error since these claims, which are generally biased against defendants, are being presented by the prosecution in an ongoing homicide case where the defendant faces the possibility of a lengthy prison sentence (DC Superior Court, 2023).

stat.AP

A probabilistic formalization of contextual bias in forensic analysis: Evidence that examiner bias leads to systemic bias in the criminal justice system

Although researchers have found evidence contextual bias in forensic science, the discussion of contextual bias is currently qualitative. We formalize years of empirical research and extend this research by showing quantitatively how biases can be propagated throughout the legal system, all the way up to the final determination of guilt in a criminal trial. We provide a probabilistic framework for describing how information is updated in a forensic analysis setting by using the ratio form of Bayes' rule. We analyze results from empirical studies using our framework and use simulations to demonstrate how bias can be compounded where experiments do not exist. We find that even minor biases in the earlier stages of forensic analysis lead to large, compounded biases in the final determination of guilt in a criminal trial.

stat.AP

A nonparametric projection-based estimator for the probability of causation, with application to water sanitation in Kenya

Current estimation methods for the probability of causation (PC) make strong parametric assumptions or are inefficient. We derive a nonparametric influence-function-based estimator for a projection of PC, which allows for simple interpretation and valid inference by making weak structural assumptions. We apply our estimator to real data from an experiment in Kenya, which found, by estimating the average treatment effect, that protecting water springs reduces childhood disease. However, before scaling up this intervention, it is important to determine whether it was the exposure, and not something else, that caused the outcome. Indeed, we find that some children, who were exposed to a high concentration of bacteria in drinking water and had a diarrheal disease, would likely have contracted the disease absent the exposure since the estimated PC for an average child in this study is 0.12 with a 95% confidence interval of (0.11, 0.13). Our nonparametric method offers researchers a way to estimate PC, which is essential if one wishes to determine not only the average treatment effect, but also whether an exposure likely caused the observed outcome.

stat.AP