Searcharxiv⌕ Search

arXiv · 2610.09420

Exact and Efficient Methods for Identifying Carcinogenic Multi-Hit Gene Combinations

Abstract

Cancer is driven by an estimated two to nine gene mutations, known as multi-hit combinations. Due to the large number of genes, identifying these multi-hit combinations presents a computationally challenging problem, which has previously required supercomputers to enumerate all or most gene combinations. Recent work formulated this classification task as the Multi-Hit Cancer Driver Set Cover Problem (MHCDSCP) and solved it via column generation on a single commodity CPU, where a pricing problem identifies promising gene combinations (arXiv:2602.22551). The main bottleneck is the pricing problem that has to be solved repeatedly. To solve the pricing problem to optimality, we propose a dedicated depth first search method that enumerates all possible gene combinations. Although enumerating gene combinations with depth first search have been performed on supercomputers with around 150,000 CPU cores (arXiv:2603.16721), our depth first search algorithm solves the pricing problem to optimality within seconds on a single CPU by leveraging dual information to prune large parts of the search space. By eliminating this bottleneck, our approach yields the first exact method for the MHCDSCP, which is guaranteed to find a global optimal solution.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Rick S. H. Willemsen, Teresa L. F. Ho, Tenindra Abeywickrama. 2026-10-07. Exact and Efficient Methods for Identifying Carcinogenic Multi-Hit Gene Combinations. https://arxiv.org/abs/2610.09420

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Perturbed Iterate SGD for Lipschitz Continuous Loss Functions with Numerical Error and Adaptive Step Sizes

Motivated by neural network training in finite-precision arithmetic environments, this work studies the convergence of perturbed iterate SGD using adaptive step sizes in an environment with numerical error. Considering a general stochastic Lipschitz continuous loss function, an asymptotic convergence result to a Clarke stationary point is proven as well as the non-asymptotic convergence to an approximate stationary point in expectation. It is assumed that only an approximation of the loss function's stochastic gradient can be computed, in addition to error in computing the SGD step itself.

math.OC↗

Controllability Allocation Scores for Targeted Network Intervention

We introduce the controllability allocation score (CAS), a framework for determining how intervention intensity should be distributed among prescribed candidate input directions with respect to designated target variables, together with the target controllability score (TCS) as its nodewise specialization. We establish existence of the CASs and develop a general uniqueness theory based on restricted injectivity of the allocation-to-Gramian map, including generic uniqueness with respect to the time horizon. We show that restricting attention to target variables can fundamentally alter the optimal intervention allocation compared with the standard full-state setting. To enable scalability, we develop a general surrogate-Gramian framework and derive objective-performance guarantees from relative Gramian errors without requiring uniqueness or closeness of the CASs. For the TCS specialization, we further construct a target-only reduced virtual system and derive explicit bounds showing how the approximation error depends on the coupling between target and non-target nodes and on the dynamical growth rate. Experiments on human brain networks show that the reduced formulation accurately approximates the TCS at short horizons, whereas the two controllability criteria exhibit markedly different approximation accuracy at long horizons.

math.OC↗

Solving the Offline and Online Min-Max Problem of Non-smooth Submodular-Concave Functions: A Zeroth-Order Approach

We consider max-min and min-max problems with objective functions that are possibly non-smooth, submodular with respect to the minimiser and concave with respect to the maximiser. We investigate the performance of a zeroth-order method applied to this problem. The method is based on the subgradient of the Lovász extension of the objective function with respect to the minimiser and based on Gaussian smoothing to estimate the smoothed function gradient with respect to the maximiser. In expectation sense, we prove the convergence of the algorithm to an $ε$-saddle point in the offline case. Moreover, we show that, in the expectation sense, in the online setting, the algorithm achieves $O(\sqrt{N(1+\bar{P}_N)})$ online duality gap, where $N$ is the number of iterations and $\bar{P}_N$ is the path length of the sequence of optimal decisions. The complexity analysis and hyperparameter selection are presented for all the cases. The theoretical results are illustrated via numerical examples.

math.OC↗