SearcharxivSearch

arXiv subjects

Soham Mukherjee

Publications and source records attributed to Soham Mukherjee.

18 recordsLinked to original sources

Latent variable estimation with composite Hilbert space Gaussian processes

We develop a scalable class of models for latent variable estimation using composite Gaussian processes, with a focus on derivative Gaussian processes. We jointly model multiple data sources as outputs to improve the accuracy of latent variable inference under a single probabilistic framework. Similarly specified exact Gaussian processes scale poorly with large datasets. To overcome this, we extend the recently developed Hilbert space approximation methods for Gaussian processes to obtain a reduced-rank representation of the composite covariance function through its spectral decomposition. Specifically, we derive and analyze the spectral decomposition of derivative covariance functions and further study their properties theoretically. Using these spectral decompositions, our methods easily scale up to data scenarios involving thousands of samples. We validate our methods in terms of latent variable estimation accuracy, uncertainty calibration, and inference speed across diverse simulation scenarios. Finally, using a real world case study from single-cell biology, we demonstrate the potential of our models in estimating latent cellular ordering given gene expression levels, thus enhancing our understanding of the underlying biological process.

stat.ME

Hilbert space methods for approximating multi-output latent variable Gaussian processes

Gaussian processes are a powerful class of non-linear models, but have limited applicability for larger datasets due to their high computational complexity. In such cases, approximate methods are required, for example, the recently developed class of Hilbert space Gaussian processes. They have been shown to significantly reduce computation time while retaining most of the favorable properties of exact Gaussian processes. However, Hilbert space approximations have so far only been developed for uni-dimensional outputs and manifest (known) inputs. Thus, we generalize Hilbert space methods to multi-output and latent input settings. Through extensive simulations, we show that the developed approximate Gaussian processes are indeed not only faster, but also provide similar or even better uncertainty calibration and accuracy of latent variable estimates compared to exact Gaussian processes. While not necessarily faster than alternative Gaussian process approximations, our new models provide better calibration and estimation accuracy, thus striking an excellent balance between trustworthiness and speed. We additionally illustrate our methods on a real-world case study from single cell biology.

stat.ME

D-GRIL: End-to-End Topological Learning with 2-parameter Persistence

End-to-end topological learning using 1-parameter persistence is well-known. We show that the framework can be enhanced using 2-parameter persistence by adopting a recently introduced 2-parameter persistence based vectorization technique called GRIL. We establish a theoretical foundation of differentiating GRIL producing D-GRIL. We show that D-GRIL can be used to learn a bifiltration function on standard benchmark graph datasets. Further, we exhibit that this framework can be applied in the context of bio-activity prediction in drug discovery.

cs.LG

TopoX: A Suite of Python Packages for Machine Learning on Topological Domains

We introduce TopoX, a Python software suite that provides reliable and user-friendly building blocks for computing and machine learning on topological domains that extend graphs: hypergraphs, simplicial, cellular, path and combinatorial complexes. TopoX consists of three packages: TopoNetX facilitates constructing and computing on these domains, including working with nodes, edges and higher-order cells; TopoEmbedX provides methods to embed topological domains into vector spaces, akin to popular graph-based embedding algorithms such as node2vec; TopoModelX is built on top of PyTorch and offers a comprehensive toolbox of higher-order message passing functions for neural networks on topological domains. The extensively documented and unit-tested source code of TopoX is available under MIT license at https://pyt-team.github.io/}{https://pyt-team.github.io/.

cs.LG

An Atmospheric Correction Integrated LULC Segmentation Model for High-Resolution Satellite Imagery

The integration of fine-scale multispectral imagery with deep learning models has revolutionized land use and land cover (LULC) classification. However, the atmospheric effects present in Top-of-Atmosphere sensor measured Digital Number values must be corrected to retrieve accurate Bottom-of-Atmosphere surface reflectance for reliable analysis. This study employs look-up-table-based radiative transfer simulations to estimate the atmospheric path reflectance and transmittance for atmospherically correcting high-resolution CARTOSAT-3 Multispectral (MX) imagery for several Indian cities. The corrected surface reflectance data were subsequently used in supervised and semi-supervised segmentation models, demonstrating stability in multi-class (buildings, roads, trees and water bodies) LULC segmentation accuracy, particularly in scenarios with sparsely labelled data.

cs.CV

GEFL: Extended Filtration Learning for Graph Classification

Extended persistence is a technique from topological data analysis to obtain global multiscale topological information from a graph. This includes information about connected components and cycles that are captured by the so-called persistence barcodes. We introduce extended persistence into a supervised learning framework for graph classification. Global topological information, in the form of a barcode with four different types of bars and their explicit cycle representatives, is combined into the model by the readout function which is computed by extended persistence. The entire model is end-to-end differentiable. We use a link-cut tree data structure and parallelism to lower the complexity of computing extended persistence, obtaining a speedup of more than 60x over the state-of-the-art for extended persistence computation. This makes extended persistence feasible for machine learning. We show that, under certain conditions, extended persistence surpasses both the WL[1] graph isomorphism test and 0-dimensional barcodes in terms of expressivity because it adds more global (topological) information. In particular, arbitrarily long cycles can be represented, which is difficult for finite receptive field message passing graph neural networks. Furthermore, we show the effectiveness of our method on real world datasets compared to many existing recent graph representation learning methods.

cs.LG

Charge Transport and Defects in Sulfur-Deficient Chalcogenide Perovskite BaZrS$_3$

Exploring the conduction mechanism in the chalcogenide perovskite BaZrS$_3$ is of significant interest due to its potential suitability as a top absorber layer in silicon-based tandem solar cells and other optoelectronic applications. Theoretical and experimental studies anticipate native ambipolar doping in BaZrS$_3$, although experimental validation remains limited. This study reveals a transition from highly insulating behavior to n-type conductivity in BaZrS$_3$ through annealing in an S-poor environment. BaZrS$_3$ thin films are synthesized $\textit{via}$ a two step process: co-sputtering of Ba-Zr followed by sulfurization at 600 $^{\circ}$C, and subsequent annealing in high vacuum. UV-Vis measurement reveal a red-shift in the absorption edge concurrent with sample color darkening after annealing. The increase in defect density with vacuum annealing, coupled with low activation energy and n-type character of defects, strongly suggests that sulfur vacancies (V$_{\mathrm{S}}$) are responsible, in agreement with theoretical predictions. The shift of the Fermi level towards conduction band minimum, quantified by Hard X-ray Photoelectron Spectroscopy (Ga K$α$, 9.25 keV), further corroborates the induced n-type of conductivity in annealed samples. Our findings indicate that vacuum annealing induces V$_{\mathrm{S}}$ defects that dominate the charge transport, thereby making BaZrS$_3$ an n-type semiconductor under S-poor conditions. This study offers crucial insights into understanding the defect properties of BaZrS$_3$, facilitating further improvements for its use in solar cell applications.

cond-mat.mtrl-sci

DGP-LVM: Derivative Gaussian process latent variable models

We develop a framework for derivative Gaussian process latent variable models (DGP-LVMs) that can handle multi-dimensional output data using modified derivative covariance functions. The modifications account for complexities in the underlying data generating process such as scaled derivatives, varying information across multiple output dimensions as well as interactions between outputs. Further, our framework provides uncertainty estimates for each latent variable samples using Bayesian inference. Through extensive simulations, we demonstrate that latent variable estimation accuracy can be drastically increased by including derivative information due to our proposed covariance function modifications. The developments are motivated by a concrete biological research problem involving the estimation of the unobserved cellular ordering from single-cell RNA (scRNA) sequencing data for gene expression and its corresponding derivative information known as RNA velocity. Since the RNA velocity is only an estimate of the exact derivative information, the derivative covariance functions need to account for potential scale differences. In a real-world case study, we illustrate the application of DGP-LVMs to such scRNA sequencing data. While motivated by this biological problem, our framework is generally applicable to all kinds of latent variable estimation problems involving derivative information irrespective of the field of study.

stat.ME

ICML 2023 Topological Deep Learning Challenge : Design and Results

This paper presents the computational challenge on topological deep learning that was hosted within the ICML 2023 Workshop on Topology and Geometry in Machine Learning. The competition asked participants to provide open-source implementations of topological neural networks from the literature by contributing to the python packages TopoNetX (data processing) and TopoModelX (deep learning). The challenge attracted twenty-eight qualifying submissions in its two-month duration. This paper describes the design of the challenge and summarizes its main findings.

cs.LG

Controlling the Manifold of Polariton States Through Molecular Disorder

Exciton polaritons, arising from the interaction of electronic transitions with confined electromagnetic fields, have emerged as a powerful tool to manipulate the properties of organic materials. However, standard experimental and theoretical approaches overlook the significant energetic disorder present in most materials now studied. Using the conjugated polymer P3HT as a model platform, we systematically tune the degree of energetic disorder and observe a corresponding redistribution of photonic character within the polariton manifold. Based on these subtle spectral features, we develop a more generalized approach to describe strong light-matter coupling in disordered systems that captures the key spectroscopic observables and provides a description of the rich manifold of states intermediate between bright and dark. Applied to a wide range of organic systems, our method challenges prevailing notions about ultrastrong coupling and whether it can be achieved with broad, disordered absorbers.

cond-mat.mtrl-sci

GRIL: A $2$-parameter Persistence Based Vectorization for Machine Learning

$1$-parameter persistent homology, a cornerstone in Topological Data Analysis (TDA), studies the evolution of topological features such as connected components and cycles hidden in data. It has been applied to enhance the representation power of deep learning models, such as Graph Neural Networks (GNNs). To enrich the representations of topological features, here we propose to study $2$-parameter persistence modules induced by bi-filtration functions. In order to incorporate these representations into machine learning models, we introduce a novel vector representation called Generalized Rank Invariant Landscape (GRIL) for $2$-parameter persistence modules. We show that this vector representation is $1$-Lipschitz stable and differentiable with respect to underlying filtration functions and can be easily integrated into machine learning models to augment encoding topological features. We present an algorithm to compute the vector representation efficiently. We also test our methods on synthetic and benchmark graph datasets, and compare the results with previous vector representations of $1$-parameter and $2$-parameter persistence modules. Further, we augment GNNs with GRIL features and observe an increase in performance indicating that GRIL can capture additional features enriching GNNs. We make the complete code for the proposed method available at https://github.com/soham0209/mpml-graph.

cs.LG

Topological Deep Learning: Going Beyond Graph Data

Topological deep learning is a rapidly growing field that pertains to the development of deep learning models for data supported on topological domains such as simplicial complexes, cell complexes, and hypergraphs, which generalize many domains encountered in scientific computations. In this paper, we present a unifying deep learning framework built upon a richer data structure that includes widely adopted topological domains. Specifically, we first introduce combinatorial complexes, a novel type of topological domain. Combinatorial complexes can be seen as generalizations of graphs that maintain certain desirable properties. Similar to hypergraphs, combinatorial complexes impose no constraints on the set of relations. In addition, combinatorial complexes permit the construction of hierarchical higher-order relations, analogous to those found in simplicial and cell complexes. Thus, combinatorial complexes generalize and combine useful traits of both hypergraphs and cell complexes, which have emerged as two promising abstractions that facilitate the generalization of graph neural networks to topological spaces. Second, building upon combinatorial complexes and their rich combinatorial and algebraic structure, we develop a general class of message-passing combinatorial complex neural networks (CCNNs), focusing primarily on attention-based CCNNs. We characterize permutation and orientation equivariances of CCNNs, and discuss pooling and unpooling operations within CCNNs in detail. Third, we evaluate the performance of CCNNs on tasks related to mesh shape analysis and graph learning. Our experiments demonstrate that CCNNs have competitive performance as compared to state-of-the-art deep learning models specifically tailored to the same tasks. Our findings demonstrate the advantages of incorporating higher-order relations into deep learning models in different applications.

cs.LG

Determining clinically relevant features in cytometry data using persistent homology

Cytometry experiments yield high-dimensional point cloud data that is difficult to interpret manually. Boolean gating techniques coupled with comparisons of relative abundances of cellular subsets is the current standard for cytometry data analysis. However, this approach is unable to capture more subtle topological features hidden in data, especially if those features are further masked by data transforms or significant batch effects or donor-to-donor variations in clinical data. Analysis of publicly available cytometry data describing non-naïve CD8+ T cells in COVID-19 patients and healthy controls shows that systematic structural differences exist between single cell protein expressions in COVID-19 patients and healthy controls. We identify proteins of interest by a decision-tree based classifier, sample points randomly and compute persistence diagrams from these sampled points. The resulting persistence diagrams identify regions in cytometry datasets of varying density and identify protruded structures such as `elbows'. We compute Wasserstein distances between these persistence diagrams for random pairs of healthy controls and COVID-19 patients and find that systematic structural differences exist between COVID-19 patients and healthy controls in the expression data for T-bet, Eomes, and Ki-67. Further analysis shows that expression of T-bet and Eomes are significantly downregulated in COVID-19 patient non-naïve CD8+ T cells compared to healthy controls. This counter-intuitive finding may indicate that canonical effector CD8+ T cells are less prevalent in COVID-19 patients than healthy controls. This method is applicable to any cytometry dataset for discovering novel insights through topological data analysis which may be difficult to ascertain otherwise with a standard gating strategy or existing bioinformatic tools.

q-bio.QM

Conformally curved initial data for charged, spinning black hole binaries on arbitrary orbits

We present a method to construct conformally curved initial data for charged black hole binaries with spin on arbitrary orbits. We generalize the superposed Kerr-Schild, extended conformal thin sandwich construction from [Lovelace et al., Phys. Rev. D {78}, 084017 (2008)] to use Kerr-Newman metrics for the superposed black holes and to solve the electromagnetic constraint equations. We implement the construction in the pseudospectral code SGRID. The code thus provides a complementary and completely independent excision-based construction, compared to the existing charged black hole initial data constructed using the puncture method [Bozzola and Paschalidis, Phys. Rev. D {99}, 104044 (2019)]. It also provides an independent implementation (with some small changes) of the Lovelace et al. vacuum construction. We construct initial data for different configurations of orbiting binaries, e.g., with black holes that are highly charged or rapidly spinning (90 and 80 percent of the extremal values, respectively, for this initial test, though the code should be able to produce data with even higher values of these parameters using higher resolutions), as well as for generic spinning, charged black holes. We carry out exploratory evolutions with the finite difference, moving punctures codes BAM (in the vacuum case) and HAD (for head-on collisions including charge), filling inside the excision surfaces. In the charged case, evolutions of these initial data provide a proxy for binary black hole waveforms in modified theories of gravity. Moreover, the generalization of the construction to Einstein-Maxwell-dilaton theory should be straightforward.

gr-qc

Bayesian Analysis of Stochastic Volatility Model using Finite Gaussian Mixtures with Unknown Number of Components

Financial studies require volatility based models which provides useful insights on risks related to investments. Stochastic volatility models are one of the most popular approaches to model volatility in such studies. The asset returns under study may come in multiple clusters which are not captured well assuming standard distributions. Mixture distributions are more appropriate in such situations. In this work, an algorithm is demonstrated which is capable of studying finite mixtures but with unknown number of components. This algorithm uses a Birth-Death process to adjust the number of components in the mixture distribution and the weights are assigned accordingly. This mixture distribution specification is then used for asset returns and a semi-parametric stochastic volatility model is fitted in a Bayesian framework. A specific case of Gaussian mixtures is studied. Using appropriate prior specification, Gibbs sampling method is used to generate posterior chains and assess model convergence. A case study of stock return data for State Bank of India is used to illustrate the methodology.

stat.ME

A-site Cation Influence on the Conduction Band of Lead Bromide Perovskites

Hot carrier solar cells hold promise for exceeding the Shockley-Queisser limit. Slow hot carrier cooling is one of the most intriguing properties of lead halide perovskites and distinguishes this class of materials from competing materials used in solar cells. Here we use the element selectivity of high-resolution X-ray spectroscopy to uncover a previously hidden feature in the conduction band states, the σ-π energy separation, and find that it is strongly influenced by the strength of electronic coupling between the A-cation and bromide-lead sublattice. Our finding provides an alternative mechanism to the commonly discussed polaronic screening and hot phonon bottleneck carrier cooling mechanisms. Our work emphasizes the optoelectronic role of the A-cation, provides a comprehensive view of A-cation effects in the electronic and crystal structures, and outlines a broadly applicable spectroscopic approach for assessing the impact of chemical alterations of the A-cation on halide and potentially non-halide perovskite electronic structure.

cond-mat.mtrl-sci

On the Hierarchical Community Structure of Practical Boolean Formulas

Modern CDCL SAT solvers easily solve industrial instances containing tens of millions of variables and clauses, despite the theoretical intractability of the SAT problem. This gap between practice and theory is a central problem in solver research. It is believed that SAT solvers exploit structure inherent in industrial instances, and hence there have been numerous attempts over the last 25 years at characterizing this structure via parameters. These can be classified as rigorous, i.e., they serve as a basis for complexity-theoretic upper bounds (e.g., backdoors), or correlative, i.e., they correlate well with solver run time and are observed in industrial instances (e.g., community structure). Unfortunately, no parameter proposed to date has been shown to be both strongly correlative and rigorous over a large fraction of industrial instances. Given the sheer difficulty of the problem, we aim for an intermediate goal of proposing a set of parameters that is strongly correlative and has good theoretical properties. Specifically, we propose parameters based on a graph partitioning called Hierarchical Community Structure (HCS), which captures the recursive community structure of a graph of a Boolean formula. We show that HCS parameters are strongly correlative with solver run time using an Empirical Hardness Model, and further build a classifier based on HCS parameters that distinguishes between easy industrial and hard random/crafted instances with very high accuracy. We further strengthen our hypotheses via scaling studies. On the theoretical side, we show that counterexamples which plagued community structure do not apply to HCS, and that there is a subset of HCS parameters such that restricting them limits the size of embeddable expanders.

cs.LO

Towards improved lossy image compression: Human image reconstruction with public-domain images

Lossy image compression has been studied extensively in the context of typical loss functions such as RMSE, MS-SSIM, etc. However, compression at low bitrates generally produces unsatisfying results. Furthermore, the availability of massive public image datasets appears to have hardly been exploited in image compression. Here, we present a paradigm for eliciting human image reconstruction in order to perform lossy image compression. In this paradigm, one human describes images to a second human, whose task is to reconstruct the target image using publicly available images and text instructions. The resulting reconstructions are then evaluated by human raters on the Amazon Mechanical Turk platform and compared to reconstructions obtained using state-of-the-art compressor WebP. Our results suggest that prioritizing semantic visual elements may be key to achieving significant improvements in image compression, and that our paradigm can be used to develop a more human-centric loss function. The images, results and additional data are available at https://compression.stanford.edu/human-compression

eess.IV