Searcharxiv⌕ Search

arXiv subjects

Fred Lu

Publications and source records attributed to Fred Lu.

23 records · Page 2Linked to original sources

Continuously Generalized Ordinal Regression for Linear and Deep Models

Ordinal regression is a classification task where classes have an order and prediction error increases the further the predicted class is from the true class. The standard approach for modeling ordinal data involves fitting parallel separating hyperplanes that optimize a certain loss function. This assumption offers sample efficient learning via inductive bias, but is often too restrictive in real-world datasets where features may have varying effects across different categories. Allowing class-specific hyperplane slopes creates generalized logistic ordinal regression, increasing the flexibility of the model at a cost to sample efficiency. We explore an extension of the generalized model to the all-thresholds logistic loss and propose a regularization approach that interpolates between these two extremes. Our method, which we term continuously generalized ordinal logistic, significantly outperforms the standard ordinal logistic model over a thorough set of ordinal regression benchmark datasets. We further extend this method to deep learning and show that it achieves competitive or lower prediction error compared to previous models over a range of datasets and modalities. Furthermore, two primary alternative models for deep learning ordinal regression are shown to be special cases of our framework.

cs.LG↗

Deep neural networks with controlled variable selection for the identification of putative causal genetic variants

Deep neural networks (DNN) have been used successfully in many scientific problems for their high prediction accuracy, but their application to genetic studies remains challenging due to their poor interpretability. In this paper, we consider the problem of scalable, robust variable selection in DNN for the identification of putative causal genetic variants in genome sequencing studies. We identified a pronounced randomness in feature selection in DNN due to its stochastic nature, which may hinder interpretability and give rise to misleading results. We propose an interpretable neural network model, stabilized using ensembling, with controlled variable selection for genetic studies. The merit of the proposed method includes: (1) flexible modelling of the non-linear effect of genetic variants to improve statistical power; (2) multiple knockoffs in the input layer to rigorously control false discovery rate; (3) hierarchical layers to substantially reduce the number of weight parameters and activations to improve computational efficiency; (4) de-randomized feature selection to stabilize identified signals. We evaluated the proposed method in extensive simulation studies and applied it to the analysis of Alzheimer disease genetics. We showed that the proposed method, when compared to conventional linear and nonlinear methods, can lead to substantially more discoveries.

cs.LG↗

Evaluating the Disentanglement of Deep Generative Models through Manifold Topology

Learning disentangled representations is regarded as a fundamental task for improving the generalization, robustness, and interpretability of generative models. However, measuring disentanglement has been challenging and inconsistent, often dependent on an ad-hoc external model or specific to a certain dataset. To address this, we present a method for quantifying disentanglement that only uses the generative model, by measuring the topological similarity of conditional submanifolds in the learned representation. This method showcases both unsupervised and supervised variants. To illustrate the effectiveness and applicability of our method, we empirically evaluate several state-of-the-art models across multiple datasets. We find that our method ranks models similarly to existing methods. We make ourcode publicly available at https://github.com/stanfordmlgroup/disentanglement.

stat.ML↗

Sub-national levels and trends in contraceptive prevalence, unmet need, and demand for family planning in Nigeria with survey uncertainty

Ambitious global goals have been established to provide universal access to affordable modern contraceptive methods. The UN's sustainable development goal 3.7.1 proposes satisfying the demand for family planning (FP) services by increasing the proportion of women of reproductive age using modern methods. To measure progress toward such goals in populous countries like Nigeria, it's essential to characterize the current levels and trends of FP indicators such as unmet need and modern contraceptive prevalence rates (mCPR). Moreover, the substantial heterogeneity across Nigeria and scale of programmatic implementation requires a sub-national resolution of these FP indicators. However, significant challenges face estimating FP indicators sub-nationally in Nigeria. In this article, we develop a robust, data-driven model to utilize all available surveys to estimate the levels and trends of FP indicators in Nigerian states for all women and by age-parity demographic subgroups. We estimate that overall rates and trends of mCPR and unmet need have remained low in Nigeria: the average annual rate of change for mCPR by state is 0.5% (0.4%,0.6%) from 2012-2017. Unmet need by age-parity demographic groups varied significantly across Nigeria; parous women express much higher rates of unmet need than nulliparous women. Our hierarchical Bayesian model incorporates data from a diverse set of survey instruments, accounts for survey uncertainty, leverages spatio-temporal smoothing, and produces probabilistic estimates with uncertainty intervals. Our flexible modeling framework directly informs programmatic decision-making by identifying age-parity-state subgroups with large rates of unmet need, highlights conflicting trends across survey instruments, and holistically interprets direct survey estimates.

stat.AP↗

Advances in using Internet searches to track dengue

Dengue is a mosquito-borne disease that threatens more than half of the world's population. Despite being endemic to over 100 countries, government-led efforts and mechanisms to timely identify and track the emergence of new infections are still lacking in many affected areas. Multiple methodologies that leverage the use of Internet-based data sources have been proposed as a way to complement dengue surveillance efforts. Among these, the trends in dengue-related Google searches have been shown to correlate with dengue activity. We extend a methodological framework, initially proposed and validated for flu surveillance, to produce near real-time estimates of dengue cases in five countries/regions: Mexico, Brazil, Thailand, Singapore and Taiwan. Our result shows that our modeling framework can be used to improve the tracking of dengue activity in multiple locations around the world.

stat.AP↗