SearcharxivSearch

arXiv subjects

Monchai Kooakachai

Publications and source records attributed to Monchai Kooakachai.

2 recordsLinked to original sources

Consistency of Familial DNA Search Results in Southeast Asian Populations

DNA databases are widely used in forensic science to identify unknown offenders. When no exact match is found, familial DNA searches can help by identifying first-degree relatives using likelihood ratios. If multiple subpopulations are relevant, likelihood ratios can be computed separately based on allele frequency estimates. Various strategies exist to combine these ratios, such as averaging allele frequencies or taking the average, maximum, or minimum likelihood ratio. While some comparisons have been made in populations like those in the U.S., their effectiveness in other regions remains unclear. This study evaluates likelihood ratio-based strategies in Southeast Asian populations, specifically Thailand, Malaysia, and Singapore. Our findings align with previous research, showing that statistical power varies across strategies. Among Thai subpopulations, the minimum likelihood ratio strategy is preferred, as it maintains high power while minimizing differences between subpopulations.

stat.AP

Incorporating Naive Bayes Classification to Address Subpopulation Structure in Familial DNA Search

Familial DNA search evaluates the genetic relatedness of two individuals by comparing the likelihood of their observed DNA profiles under two competing hypotheses-the null hypothesis that the individuals are unrelated and the alternative hypothesis that they are related-most commonly through the likelihood ratio (LR). Standard LR-based approaches typically assume a uniform genetic background; however, this assumption is rarely valid due to population substructure, where allele frequencies vary among subpopulations and can bias relationship inference. Existing modifications-such as LR calculations based on average allele frequencies (LRLAF) and strategies using maximum, minimum, or average likelihood ratios (LRMAX, LRMIN, LRAVG)-help mitigate these challenges but remain limited in their ability to fully address subpopulation differences. This study introduces a new LR-based statistic, LRCLASS, which incorporates a classification step using the Naive Bayes classifier to account for nuisance parameters associated with unknown subpopulation origins. In LRCLASS, the two DNA profiles being compared are jointly assigned to a subpopulation group via Naive Bayes before LR computation. Empirical evaluations using Thai population data show that LRCLASS achieves higher statistical power for detecting full-sibling relationships than existing LR-based methods. We further assessed multinomial logistic regression as an alternative classifier and found its performance comparable to that of Naive Bayes, suggesting flexibility in classifier choice. Overall, integrating the Naive Bayes classifier with LR computation offers a robust strategy for addressing population substructure in familial DNA search and highlights the broader potential of combining supervised learning techniques with forensic statistical methodologies to enhance the accuracy and reliability of genetic relationship testing.

stat.AP