Searcharxiv⌕ Search

arXiv subjects

Md. Noor-E-Alam

Publications and source records attributed to Md. Noor-E-Alam.

At least 19 recordsLinked to original sources

A Novel Computational Framework for Causal Inference: Tree-Based Discretization with ILP-Based Matching

Causal inference is essential for data-driven decision-making, as it aims to uncover causal relationships from observational data. However, identifying causality remains challenging due to the potential for confounding and the distinction between correlation and causation. While recent advances in causal machine learning and matching algorithms have improved estimation accuracy, these methods often face trade-offs between interpretability and computational efficiency. This paper proposes a novel approach that combines a tree-based discretization technique, tailored for causal inference, with an integer linear programming-based matching algorithm. The discretization ensures approximately linear relationships for control datasets within strata, enabling effective matching, while the optimization framework optimizes for global balance. The resulting algorithm yields computational efficiency and less biased ATT estimates compared to state-of-the-art algorithms. Empirical evaluations demonstrate the proposed method's practical advantages over existing techniques in causal inference scenarios.

stat.ML↗

Copula-Based Endogeneity Correction for Doubly Robust Estimation of Treatment Effect

Doubly Robust (DR) estimation of treatment effect relies on an untestable assumption that is the absence of unobserved confounding. This assumption is par- ticularly problematic in the context of healthcare research, where variables like pre- scription refill rates serve as proxies for unobserved behaviors such as medication adherence. These proxy variables are often endogenous, exhibiting correlation with the regression error term due to unmeasured confounding or measurement error. We propose a copula-corrected doubly robust estimator that addresses endogeneity in both the treatment and outcome models without requiring instrumental variables. Gaussian copulas model the joint distribution of endogenous covariates and the error term, enabling consistent estimation while preserving the doubly robust property that requires correct specification of either the treatment or outcome model, not both. Monte Carlo simulations demonstrate that naive DR estimation exhibits substantial bias under endogeneity, whereas our corrected estimator recovers unbiased treatment effects across different data-generating processes. We apply our method to examine the effect of nutritional counseling on blood pressure using the National Health and Nutrition Examination Survey (NHANES) data. Naive DR estimation suggests counseling is associated with increased blood pressure. After copula correction, this effect becomes statistically insignificant, consistent with literature showing modest effects of nutri- Counseling in reducing blood pressure. Our methodology provides researchers with a practical tool for obtaining treatment effects in the presence of endogeneity.

stat.ME↗

A Two-Stage Interpretable Matching Framework for Causal Inference

Matching in causal inference from observational data aims to construct treatment and control groups with similar distributions of covariates, thereby reducing confounding and ensuring an unbiased estimation of treatment effects. This matched sample closely mimics a randomized controlled trial (RCT), thus improving the quality of causal estimates. We introduce a novel Two-stage Interpretable Matching (TIM) framework for transparent and interpretable covariate matching. In the first stage, we perform exact matching across all available covariates. For treatment and control units without an exact match in the first stage, we proceed to the second stage. Here, we iteratively refine the matching process by removing the least significant confounder in each iteration and attempting exact matching on the remaining covariates. We learn a distance metric for the dropped covariates to quantify closeness to the treatment unit(s) within the corresponding strata. We used these high- quality matches to estimate the conditional average treatment effects (CATEs). To validate TIM, we conducted experiments on synthetic datasets with varying association structures and correlations. We assessed its performance by measuring bias in CATE estimation and evaluating multivariate overlap between treatment and control groups before and after matching. Additionally, we apply TIM to a real-world healthcare dataset from the Centers for Disease Control and Prevention (CDC) to estimate the causal effect of high cholesterol on diabetes. Our results demonstrate that TIM improves CATE estimates, increases multivariate overlap, and scales effectively to high-dimensional data, making it a robust tool for causal inference in observational data.

cs.AI↗

Optimizing Feature Selection in Causal Inference: A Three-Stage Computational Framework for Unbiased Estimation

Feature selection is an important but challenging task in causal inference for obtaining unbiased estimates of causal quantities. Properly selected features in causal inference not only significantly reduce the time required to implement a matching algorithm but, more importantly, can also reduce the bias and variance when estimating causal quantities. When feature selection techniques are applied in causal inference, the crucial criterion is to select variables that, when used for matching, can achieve an unbiased and robust estimation of causal quantities. Recent research suggests that balancing only on treatment-associated variables introduces bias while balancing on spurious variables increases variance. To address this issue, we propose an enhanced three-stage framework that shows a significant improvement in selecting the desired subset of variables compared to the existing state-of-the-art feature selection framework for causal inference, resulting in lower bias and variance in estimating the causal quantity. We evaluated our proposed framework using a state-of-the-art synthetic data across various settings and observed superior performance within a feasible computation time, ensuring scalability for large-scale datasets. Finally, to demonstrate the applicability of our proposed methodology using large-scale real-world data, we evaluated an important US healthcare policy related to the opioid epidemic crisis: whether opioid use disorder has a causal relationship with suicidal behavior.

stat.ME↗

A Two-Stage Feature Selection Approach for Robust Evaluation of Treatment Effects in High-Dimensional Observational Data

A Randomized Control Trial (RCT) is considered as the gold standard for evaluating the effect of any intervention or treatment. However, its feasibility is often hindered by ethical, economical, and legal considerations, making observational data a valuable alternative for drawing causal conclusions. Nevertheless, healthcare observational data presents a difficult challenge due to its high dimensionality, requiring careful consideration to ensure unbiased, reliable, and robust causal inferences. To overcome this challenge, in this study, we propose a novel two-stage feature selection technique called, Outcome Adaptive Elastic Net (OAENet), explicitly designed for making robust causal inference decisions using matching techniques. OAENet offers several key advantages over existing methods: superior performance on correlated and high-dimensional data compared to the existing methods and the ability to select specific sets of variables (including confounders and variables associated only with the outcome). This ensures robustness and facilitates an unbiased estimate of the causal effect. Numerical experiments on simulated data demonstrate that OAENet significantly outperforms state-of-the-art methods by either producing a higher-quality estimate or a comparable estimate in significantly less time. To illustrate the applicability of OAENet, we employ large-scale US healthcare data to estimate the effect of Opioid Use Disorder (OUD) on suicidal behavior. When compared to competing methods, OAENet closely aligns with existing literature on the relationship between OUD and suicidal behavior. Performance on both simulated and real-world data highlights that OAENet notably enhances the accuracy of estimating treatment effects or evaluating policy decision-making with causal inference.

cs.LG↗

A Nationwide Multi-Location Multi-Resource Stochastic Programming Based Energy Planning Framework

The global increase in energy consumption and demand has forced many countries to transition into including more diverse energy sources in their electricity market. To efficiently utilize the available fuel resources, all energy sources must be optimized simultaneously. However, the inherent variability in variable renewable energy generators makes deterministic models ineffective. On the other hand, comprehensive stochastic models, including all sources of generation across a nation, can become computationally intractable. This work proposes a comprehensive national energy planning framework from a policymaker's perspective, which is generalizable to any country, region, or any group of countries in energy trade agreements. Given its relative land area and energy consumption globally, the United States is selected as a case study. A two-stage stochastic programming approach is adopted, and a scenario-based Benders decomposition modeling approach is employed to achieve computational efficiency for the large-scale model. Data is obtained from the U.S. Energy Information Administration's online data collection. Scenarios for the uncertain parameters are developed using a k-means clustering algorithm. Various cases are compared from a financial perspective to inform policymaking by identifying implementable cost-reduction strategies. The findings underscore the importance of promoting resource coordination and demonstrate the impact of increasing interchange and renewable energy capacity on overall gains. Moreover, the study highlights the value of stochastic optimization modeling compared to deterministic modeling to make energy planning decisions.

math.OC↗

Optimizing Return and Secure Disposal of Prescription Opioids to Reduce the Diversion to Secondary Users and Black Market

Opioid Use Disorder (OUD) has reached an epidemic level in the US. Diversion of unused prescription opioids to secondary users and black market significantly contributes to the abuse and misuse of these highly addictive drugs, leading to the increased risk of OUD and accidental opioid overdose within communities. Hence, it is critical to design effective strategies to reduce the non-medical use of opioids that can occur via diversion at the patient level. In this paper, we aim to address this critical public health problem by designing strategies for the return and safe disposal of unused prescription opioids. We propose a data-driven optimization framework to determine the optimal incentive disbursement plans and locations of easily accessible opioid disposal kiosks to motivate prescription opioid users of diverse profiles in returning their unused opioids. We develop a Mixed-Integer Non-Linear Programming (MINLP) model to solve the decision problem, followed by a reformulation scheme using Benders Decomposition that results in a computationally efficient solution. We present a case study to show the benefits and usability of the model using a dataset created from Massachusetts All Payer Claims Data (MA APCD). Our proposed model allows the policymakers to estimate and include a penalty cost considering the economic and healthcare burden associated with prescription opioid diversion. Our numerical experiments demonstrate the ability of model and usefulness in determining optimal locations of opioid disposal kiosks and incentive disbursement plans for maximizing the disposal of unused opioids. The proposed optimization framework offers various trade-off strategies that can help government agencies design pragmatic policies for reducing the diversion of unused prescription opioids.

math.OC↗

A robust approach to quantifying uncertainty in matching problems of causal inference

Unquantified sources of uncertainty in observational causal analyses can break the integrity of the results. One would never want another analyst to repeat a calculation with the same dataset, using a seemingly identical procedure, only to find a different conclusion. However, as we show in this work, there is a typical source of uncertainty that is essentially never considered in observational causal studies: the choice of match assignment for matched groups, that is, which unit is matched to which other unit before a hypothesis test is conducted. The choice of match assignment is anything but innocuous, and can have a surprisingly large influence on the causal conclusions. Given that a vast number of causal inference studies test hypotheses on treatment effects after treatment cases are matched with similar control cases, we should find a way to quantify how much this extra source of uncertainty impacts results. What we would really like to be able to report is that \emph{no matter} which match assignment is made, as long as the match is sufficiently good, then the hypothesis test result still holds. In this paper, we provide methodology based on discrete optimization to create robust tests that explicitly account for this possibility. We formulate robust tests for binary and continuous data based on common test statistics as integer linear programs solvable with common methodologies. We study the finite-sample behavior of our test statistic in the discrete-data case. We apply our methods to simulated and real-world datasets and show that they can produce useful results in practical applied settings.

stat.ME↗

A Robust Optimization Framework for Two-Echelon Vehicle and UAV Routing for Post-Disaster Humanitarian Logistics Operations

Providing first aid and other supplies (e.g., epi-pens, medical supplies, dry food, water) during and after a disaster is always challenging. The complexity of these operations increases when the transportation, power, and communications networks fail, leaving people stranded and unable to communicate their locations and needs. The advent of emerging technologies like uncrewed autonomous vehicles can help humanitarian logistics providers reach otherwise stranded populations after transportation network failures. However, due to the failures in telecommunication infrastructure, demand for emergency aid can become uncertain. To address the challenges of delivering emergency aid to trapped populations with failing infrastructure networks, we propose a novel robust computational framework for a two-echelon vehicle routing problem that uses uncrewed autonomous vehicles, or drones, for the deliveries. We formulate the problem as a two-stage robust optimization model to handle demand uncertainty. Then, we propose a column-and-constraint generation approach for worst-case demand scenario generation for a given set of truck and drone routes. Moreover, we develop a decomposition scheme inspired by the column generation approach to heuristically generate drone routes for a set of demand scenarios. Finally, we combine the heuristic decomposition scheme within the column-andconstraint generation approach to determine robust routes for both trucks and drones, the time that affected communities are served, and the quantities of aid materials delivered. To validate our proposed computational framework, we use a simulated dataset that aims to recreate emergency aid requests in different areas of Puerto Rico after Hurricane Maria in 2017.

math.OC↗

Computational Approaches for Solving Two-Echelon Vehicle and UAV Routing Problems for Post-Disaster Humanitarian Operations

Humanitarian logistics service providers have two major responsibilities immediately after a disaster: locating trapped people and routing aid to them. These difficult operations are further hindered by failures in the transportation and telecommunications networks, which are often rendered unusable by the disaster at hand. In this work, we propose a two-echelon vehicle routing framework for performing these operations using aerial uncrewed autonomous vehicles (UAVs or drones) to address the issues associated with these failures. In our proposed framework, we assume that ground vehicles cannot reach the trapped population directly, but they can only transport drones from a depot to some intermediate locations. The drones launched from these locations serve to both identify demands for medical and other aids (e.g., epi-pens, medical supplies, dry food, water) and make deliveries to satisfy them. Specifically, we present a decision framework, in which the resulting optimization problem is formulated as a two-echelon vehicle routing problem with trucks as the first echelon vehicles and for the second echelon vehicles, we consider two types of drones. Hotspot drones have the capability of providing a cell phone and internet reception and hence are used to capture demands. Delivery drones are subsequently employed to satisfy the observed demand. To handle demand uncertainty, we decompose the decision problem into two stages: providing telecommunications capabilities in the first stage thereby capturing demand precisely, and satisfying the resulting demands in the second stage. To solve the resulting models, we propose efficient computational approaches by designing a decomposition algorithm with column generation (CG)-based heuristics to identify optimal drone routes.

math.OC↗

A Computational Framework for Solving Nonlinear Binary OptimizationProblems in Robust Causal Inference

Identifying cause-effect relations among variables is a key step in the decision-making process. While causal inference requires randomized experiments, researchers and policymakers are increasingly using observational studies to test causal hypotheses due to the wide availability of observational data and the infeasibility of experiments. The matching method is the most used technique to make causal inference from observational data. However, the pair assignment process in one-to-one matching creates uncertainty in the inference because of different choices made by the experimenter. Recently, discrete optimization models are proposed to tackle such uncertainty. Although a robust inference is possible with discrete optimization models, they produce nonlinear problems and lack scalability. In this work, we propose greedy algorithms to solve the robust causal inference test instances from observational data with continuous outcomes. We propose a unique framework to reformulate the nonlinear binary optimization problems as feasibility problems. By leveraging the structure of the feasibility formulation, we develop greedy schemes that are efficient in solving robust test problems. In many cases, the proposed algorithms achieve global optimal solutions. We perform experiments on three real-world datasets to demonstrate the effectiveness of the proposed algorithms and compare our result with the state-of-the-art solver. Our experiments show that the proposed algorithms significantly outperform the exact method in terms of computation time while achieving the same conclusion for causal tests. Both numerical experiments and complexity analysis demonstrate that the proposed algorithms ensure the scalability required for harnessing the power of big data in the decision-making process.

math.OC↗

Two-Stage Stochastic Optimization Frameworks to Aid in Decision-Making Under Uncertainty for Variable Resource Generators Participating in a Sequential Energy Market

Decisions for a variable renewable resource generators commitment in the energy market are typically made in advance when little information is obtainable about wind availability and market prices. Much research has been published recommending various frameworks for addressing this issue. However, these frameworks are limited as they do not consider all markets a producer can participate in. Moreover, current stochastic programming models do not allow for uncertainty data to be updated as more accurate information becomes available. This work proposes two decision-making frameworks for a wind energy generator participating in day-ahead, intraday, reserve, and balancing markets. The first framework is a two-stage stochastic convex optimization approach, where both scenario-independent and scenario-dependent decisions are made concurrently. The second framework is a series of four two-stage stochastic optimization models wherein the results from each model feed into each subsequent model allowing for scenarios to be updated as more information becomes available to the decision-maker. In the simulation experiments, the multi-phase framework performs better than the single-phase in every run, and results in an average profit increase of 7%. The proposed optimization frameworks aid in better decision-making while addressing uncertainty related to variable resource generators and maximize the return on investment.

math.OC↗

Heavy Ball Momentum Induced Sampling Kaczmarz Motzkin Methods for Linear Feasibility Problems

The recently proposed Sampling Kaczmarz Motzkin (SKM) algorithm performs well in comparison with the state-of-the-art methods in solving large-scale Linear Feasibility (LF) problems. To explore the concept of momentum in the context of solving LF problems, in this work, we propose a momentum induced algorithm called Momentum Sampling Kaczmarz Motzkin (MSKM). The MSKM algorithm is developed by integrating the heavy ball momentum to the SKM algorithm. We provide a rigorous convergence analysis of the proposed MSKM algorithm from which we obtain convergence results of several Kaczmarz type methods for solving LF problems. Moreover, under somewhat weaker conditions, we establish a sub-linear convergence rate for the so-called Cesaro average of the sequence generated by the MSKM algorithm. We then back up the theoretical results via thorough numerical experiments on artificial and real datasets. For a fair comparison, we test our proposed method in comparison with the SKM method on a wide variety of test instances: 1) randomly generated instances, 2) Netlib LPs and 3) linear classification test instances. We also compare the proposed method with the traditional Interior Point Method (IPM) and Active Set Method (ASM) on Netlib LPs. The proposed momentum induced algorithm significantly outperforms the basic SKM method (with no momentum) on all of the considered test instances. Furthermore, the proposed algorithm also performs well in comparison with IPM and ASM algorithms. Finally, we propose a stochastic version of the MSKM algorithm called Stochastic-Momentum Sampling Kaczmarz Motzkin (SSKM) to better handle large-scale real-world data. We conclude our work with a rigorous theoretical convergence analysis of the proposed SSKM algorithm.

math.OC↗

Sketch & Project Methods for Linear Feasibility Problems: Greedy Sampling & Momentum

We develop two greedy sampling rules for the Sketch & Project method for solving linear feasibility problems. The proposed greedy sampling rules generalize the existing max-distance sampling rule and uniform sampling rule and generate faster variants of Sketch & Project methods. We also introduce greedy capped sampling rules that improve the existing capped sampling rules. Moreover, we incorporate the so-called heavy ball momentum technique to the proposed greedy Sketch & Project method. By varying the parameters such as sampling rules, sketching vectors; we recover several well-known algorithms as special cases, including Randomized Kaczmarz (RK), Motzkin Relaxation (MR), Sampling Kaczmarz Motzkin (SKM). We also obtain several new methods such as Randomized Coordinate Descent, Sampling Coordinate Descent, Capped Coordinate Descent, etc. for solving linear feasibility problems. We provide global linear convergence results for both the basic greedy method and the greedy method with momentum. Under weaker conditions, we prove $\mathcal{O}(\frac{1}{k})$ convergence rate for the Cesaro average of sequences generated by both methods. We extend the so-called certificate of feasibility result for the proposed momentum method that generalizes several existing results. To back up the proposed theoretical results, we carry out comprehensive numerical experiments on randomly generated test instances as well as sparse real-world test instances. The proposed greedy sampling methods significantly outperform the existing sampling methods. And finally, the momentum variants designed in this work extend the computational performance of the Sketch & Project methods for all of the sampling rules.

math.NA↗

Sampling Kaczmarz Motzkin Method for Linear Feasibility Problems: Generalization & Acceleration

Randomized Kaczmarz (RK), Motzkin Method (MM) and Sampling Kaczmarz Motzkin (SKM) algorithms are commonly used iterative techniques for solving a system of linear inequalities (i.e., $Ax \leq b$). As linear systems of equations represent a modeling paradigm for solving many optimization problems, these randomized and iterative techniques are gaining popularity among researchers in different domains. In this work, we propose a Generalized Sampling Kaczmarz Motzkin (GSKM) method that unifies the iterative methods into a single framework. In addition to the general framework, we propose a Nesterov type acceleration scheme in the SKM method called as Probably Accelerated Sampling Kaczmarz Motzkin (PASKM). We prove the convergence theorems for both GSKM and PASKM algorithms in the $L_2$ norm perspective with respect to the proposed sampling distribution. Furthermore, we prove sub-linear convergence for the Cesaro average of iterates for the proposed GSKM and PASKM algorithms.From the convergence theorem of the GSKM algorithm, we find the convergence results of several well-known algorithms like the Kaczmarz method, Motzkin method and SKM algorithm. We perform thorough numerical experiments using both randomly generated and real-world (classification with support vector machine and Netlib LP) test instances to demonstrate the efficiency of the proposed methods. We compare the proposed algorithms with SKM, Interior Point Method (IPM) and Active Set Method (ASM) in terms of computation time and solution quality. In the majority of the problem instances, the proposed generalized and accelerated algorithms significantly outperform the state-of-the-art methods.

math.OC↗

A Big Data Analytics Framework to Predict the Risk of Opioid Use Disorder

Overdose related to prescription opioids have reached an epidemic level in the US, creating an unprecedented national crisis. This has been exacerbated partly due to the lack of tools for physicians to help predict the risk of whether a patient will develop opioid use disorder. Little is known about how machine learning can be applied to a big-data platform to ensure an informed, sustained and judicious prescribing of opioids, in particular for commercially insured population. This study explores Massachusetts All Payer Claims Data, a de-identified healthcare dataset, and proposes a machine learning framework to examine how naïve users develop opioid use disorder. We perform several feature selections techniques to identify influential demographic and clinical features associated with opioid use disorder from a class imbalanced analytic sample. We then compare the predictive power of four well-known machine learning algorithms: Logistic Regression, Random Forest, Decision Tree, and Gradient Boosting to predict the risk of opioid use disorder. The study results show that the Random Forest model outperforms the other three algorithms while determining the features, some of which are consistent with prior clinical findings. Moreover, alongside the higher predictive accuracy, the proposed framework is capable of extracting some risk factors that will add significant knowledge to what is already known in the extant literature. We anticipate that this study will help healthcare practitioners improve the current prescribing practice of opioids and contribute to curb the increasing rate of opioid addiction and overdose.

stat.AP↗

A Primal-Dual Interior Point Method for a Novel Type-2 Second Order Cone Optimization Problem

In this paper, we define a new, special second order cone as a type-$k$ second order cone. We focus on the case of $k=2$, which can be viewed as SOCO with an additional {\em complicating variable}. For this new problem, we develop the necessary prerequisites, based on previous work for traditional SOCO. We then develop a primal-dual interior point algorithm for solving a type-2 second order conic optimization (SOCO) problem, based on a family of kernel functions suitable for this type-2 SOCO. We finally derive the following iteration bound for our framework: \[\frac{L^γ}{θκγ} \left[2N ψ\left( \frac{\varrho \left(τ/4N\right)}{\sqrt{1-θ}}\right)\right]^γ\log \frac{3N}ε.\]

math.OC↗

Accelerated Sampling Kaczmarz Motzkin Algorithm for The Linear Feasibility Problem

The Sampling Kaczmarz Motzkin (SKM) algorithm is a generalized method for solving large scale linear systems of inequalities. Having its root in the relaxation method of Agmon, Schoenberg, and Motzkin and the randomized Kaczmarz method, SKM outperforms the state of the art methods in solving large-scale Linear Feasibility (LF) problems. Motivated by SKM's success, in this work, we propose an Accelerated Sampling Kaczmarz Motzkin (ASKM) algorithm which achieves better convergence compared to the standard SKM algorithm on ill conditioned problems. We provide a thorough convergence analysis for the proposed accelerated algorithm and validate the results with various numerical experiments. We compare the performance and effectiveness of ASKM algorithm with SKM, Interior Point Method (IPM) and Active Set Method (ASM) on randomly generated instances as well as Netlib LPs. In most of the test instances, the proposed ASKM algorithm outperforms the other state of the art methods.

math.OC↗