SearcharxivSearch

arXiv subjects

Simon D. Nguyen

Publications and source records attributed to Simon D. Nguyen.

3 recordsLinked to original sources

REALITrees: Rashomon Ensemble Active Learning for Interpretable Trees

Active learning reduces labeling costs by selecting samples that maximize information gain. A dominant framework, Query-by-Committee (QBC), typically relies on perturbation-based diversity by inducing model disagreement through random feature subsetting or data blinding. While this approximates one notion of epistemic uncertainty, it sacrifices direct characterization of the plausible hypothesis space. We propose the complementary approach: Rashomon Ensembled Active Learning (REAL) which constructs a committee by exhaustively enumerating the Rashomon Set of all near-optimal models. To address functional redundancy within this set, we adopt a PAC-Bayesian framework using a Gibbs posterior to weight committee members by their empirical risk. Leveraging recent algorithmic advances, we exactly enumerate this set for the class of sparse decision trees. Across synthetic and established active learning baselines, REAL outperforms randomized ensembles, particularly in moderately noisy environments where it strategically leverages expanded model multiplicity to achieve faster convergence.

stat.ML

Adaptive Active Learning for Regression via Reinforcement Learning

Active learning for regression reduces labeling costs by selecting the most informative samples. Improved Greedy Sampling is a prominent method that balances feature-space diversity and output-space uncertainty using a static, multiplicative rule. We propose Weighted improved Greedy Sampling (WiGS), which replaces this framework with a dynamic, additive criterion. We formulate weight selection as a reinforcement learning problem, enabling an agent to adapt the exploration-investigation balance throughout learning. Experiments on 18 benchmark datasets and a synthetic environment show WiGS outperforms iGS and other baseline methods in both accuracy and labeling efficiency, particularly in domains with irregular data density where the baseline's multiplicative rule ignores high-error samples in dense regions.

stat.ML

Evaluating and Ranking Criteria for Mathematics Graduate Education Admissions through the Analytic Hierarchy Process

Graduate school admission committees consider many factors for admission into a mathematics PhD program, and aspiring applicants often wonder which factors are more important. Applicants discerning where to best dedicate their time may ask questions such as, "should I take graduate level courses or participate in research?" This paper seeks to answer such questions by constructing an ordinal ranking of admission criteria and provide insight on why some factors are more valued. Using a conjoint analysis method called Analytic Hierarchy Process from mathematician Thomas Saaty, this paper evaluates the relative influence certain predictors have on admissions into a mathematics PhD program. Additionally, this paper analyzes the difference in factor rankings between varying populations. For instance, results indicate that full professors tend to differ from their associate and assistant peers in their criteria rankings, and top 30 math PhD programs also differ in rankings when compared to programs outside the top 30. It is my hope that the ordinal rankings and its subsequent analyses will not only explain why some factors are valued more by the mathematics community but also provide some guidance to potential graduate school applicants on where to dedicate their time as undergraduates

math.HO