Searcharxiv⌕ Search

arXiv subjects

Snehanshu Saha

Publications and source records attributed to Snehanshu Saha.

At least 19 recordsLinked to original sources

A Generalisation Signal Need Not Be a Model-Selection Signal

Model selection in computational biology often relies on validation data drawn from the training regime, even when deployment lies outside it. When validation no longer preserves which model is best, a natural alternative is to rank candidates using properties of the trained network itself. We test this idea using a novel, forward-only proxy motivated by the norm of the Hessian, alongside common Hessian measures, across molecular property, protein fitness, and drug-response tasks. Contrary to our hypothesis, geometry does not become more useful as validation Spearman correlation deteriorates: augmenting validation helps some shifts but significantly harms others. More surprisingly, the proxy still correlates with generalisation gap on most tasks even when Hessian trace and top-eigenvalue relationships are weak or reversed, yet this signal does not reliably identify the deployment-best model. A curvature bound need not preserve cross-model rankings, and low geometric scores can even favour collapsed predictors. Thus, a generalisation signal need not be a model-selection signal.

cs.LG↗

Decoupling candidate dual AGN from chance superpositions in the GOTHIC survey via a deep-learning framework

Dual active galactic nuclei (DAGN) mark a critical phase in the evolution of merging galaxies and the pairing of supermassive black holes, yet they remain difficult to identify in large imaging surveys because of projection effects and limited spatial resolution. Compact foreground stars and unresolved substructure can mimic dual nuclei through chance superposition, complicating automated detection. We revisit the 46,061 galaxies flagged but rejected as DAGN candidates by the GOTHIC pipeline, primarily because the two nuclei fell within the SDSS fibre aperture or exceeded its separation threshold. We train a supervised deep-learning framework based on the YOLOv11 oriented-bounding-box architecture on annotated SDSS imaging to separate genuine dual nuclei from foreground stellar contaminants and other spurious alignments. The final model attains a validation precision of 0.919, recall of 0.905, and $F_1$ of 0.912 for the dual-nuclei class, and yields 29,605 dual-nucleus candidates after removing star-dominated and blended detections. Structured visual inspection indicates that $54.5$--$62\%$ are consistent with genuine dual nuclei, implying $\sim(1.4$--$1.8)\times10^{4}$ plausible systems. Cross-calibrating the YOLO separation against the deterministic GOTHIC centroid measurement and restricting to the compact regime ($d \le 6.87''$) gives a conservative subset of $\sim 13{,}672$ candidates, reaching calibrated separations of $\sim 0.56''$. Spectroscopy of the most compact ($\le 1$~kpc) systems shows they are dominated by passive, absorption-line galaxies with no resolved double-peaked emission, so confirmation requires higher-resolution follow-up. The catalogue is a statistically refined list of candidates, not confirmed DAGN. Nonetheless, deep-learning detection substantially reduces contamination and expands the plausible DAGN census.

astro-ph.GA↗

The Cosmic Ultraviolet Background at the Galactic Poles

We have used archival GALEX data to separate the cosmic ultraviolet background at the Galactic Poles into two components: the dust scattered light and an offset. We have modeled the dust-scattered light using a single- scattering model finding 1 sigma limits of 0.54 -- 0.71 for the albedo (a) and 0.74 -- 0.83 for the phase function asymmetry factor (g) at 1530 Å and 0.66 -- 0.73 for a and 0.71 -- 0.77 for g at 2360 Å, that is, the grains are moderately reflective and highly forward-scattering. The offsets are 277 -- 284 photon units at 1530 Å and 513 -- 520 photon units at 2360 Å. We have estimated other Galactic and extragalactic contributors to the offset finding that 161 +- 18 photon units is unaccounted for at 1530 Å and 335 +- 38 at 2360 Å. The offsets are constant over these regions with variances of 20 -- 30 photon units.

astro-ph.GA↗

Investigating the Spectral Properties of Dual Nuclei in Galaxy Mergers from the GOTHIC survey: Supermassive Black Hole Growth, metal enrichment and Dual AGN

Dual nuclei systems are galaxy merger remnants or closely merging galaxies that have two distinct stellar cores separated by ~ 10pc to 10kpc. They are important laboratories for probing the co-evolution of stellar populations, galaxy dynamics, and central black holes during the hierarchical assembly of galaxies. In this study, we present a spectroscopic analysis of a sample of dual nuclei from the GOTHIC survey, using the penalized pixel-fitting (pPXF) code. The sample consists of star forming nuclei pairs, dual active galactic nuclei (DAGN) and mixed pairs. Using the SDSS spectra, we extracted stellar kinematics, emission line fluxes, the star formation history, metallicity of the nuclei, and derived important properties such as the supermassive black hole (SMBH) masses, accretion rates and SMBH ratios. We compared different properties of the nuclei in the dual systems, such as stellar velocity dispersion, stellar masses, black hole masses, age and metallicity. Our results show that the SMBH masses are higher for BHs in galaxy mergers compared to single nuclei for a given stellar mass, thus revealing that SMBHs grow during the galaxy merging process and not only due to the merger of SMBHs. Our study provides new observational constraints on the dynamical and evolutionary states of dual-nuclei systems, offering a deeper understanding of the role these systems play in galaxy evolution and central black hole growth.

astro-ph.GA↗

Altruistic Ride Sharing: A Framework for Fair and Sustainable Urban Mobility via Peer-to-Peer Incentives

Urban mobility systems face persistent challenges of congestion, underutilized vehicles, and rising emissions driven by private point-to-point commuting. Although ride-sharing platforms exist, their profit-driven incentive structures often fail to align individual participation with broader community benefit. We introduce Altruistic Ride Sharing (ARS), a decentralized peer-to-peer mobility framework in which commuters alternate between driver and rider roles using altruism points, a non-monetary credit mechanism that rewards providing rides and discourages persistent free-riding. To enable scalable coordination among agents, ARS formulates ride-sharing as a multi-agent reinforcement learning problem and introduces ORACLE (One-Network Actor-Critic for Learning in Cooperative Environments), a shared-parameter learning architecture for decentralized rider selection. We evaluate ARS using real-world New York City Taxi and Limousine Commission (TLC) trajectory data under varying agent populations and behavioral dynamics. Across simulations, ARS reduces total travel distance and associated carbon emissions by approximately 20%, reduces urban traffic density by up to 30%, and doubles vehicle utilization relative to no-sharing baselines while maintaining balanced participation across agents. These results demonstrate that altruism-based incentives combined with decentralized learning can provide a scalable and equitable alternative to profit-driven ride-sharing systems.

cs.MA↗

Multi-Agent Training-free Urban Food Delivery System using Resilient UMST Network

Delivery systems have become a core part of urban life, supporting the demand for food, medicine, and other goods. Yet traditional logistics networks remain fragile, often struggling to adapt to road closures, accidents, and shifting demand. Online Food Delivery (OFD) platforms now represent a cornerstone of urban logistics, with the global market projected to grow to over 500 billion USD by 2030. Designing delivery networks that are efficient and resilient remains a major challenge: fully connected graphs provide flexibility but are computationally infeasible at scale, while single Minimum Spanning Trees (MSTs) are efficient but easily disrupted. We propose the Union of Minimum Spanning Trees (UMST) approach to construct delivery networks that are sparse yet robust. UMST generates multiple MSTs through randomized edge perturbations and unites them, producing graphs with far fewer edges than fully connected networks while maintaining multiple alternative routes between delivery hotspots. Across multiple U.S. cities, UMST achieves 20--40$\times$ fewer edges than fully connected graphs while enabling substantial order bundling with 75--83% participation rates. Compared to learning-based baselines including MADDPG and Graph Neural Networks, UMST delivers competitive performance (88-96% success rates, 44-53% distance savings) without requiring training, achieving 30$\times$ faster execution while maintaining interpretable routing structures. Its combination of structural efficiency and operational flexibility offers a scalable and resilient foundation for urban delivery networks.

cs.MA↗

Matching High-Dimensional Geometric Quantiles for Test-Time Adaptation of Transformers and Convolutional Networks Alike

Test-time adaptation (TTA) refers to adapting a classifier for the test data when the probability distribution of the test data slightly differs from that of the training data of the model. To the best of our knowledge, most of the existing TTA approaches modify the weights of the classifier relying heavily on the architecture. It is unclear as to how these approaches are extendable to generic architectures. In this article, we propose an architecture-agnostic approach to TTA by adding an adapter network pre-processing the input images suitable to the classifier. This adapter is trained using the proposed quantile loss. Unlike existing approaches, we correct for the distribution shift by matching high-dimensional geometric quantiles. We prove theoretically that under suitable conditions minimizing quantile loss can learn the optimal adapter. We validate our approach on CIFAR10-C, CIFAR100-C and TinyImageNet-C by training both classic convolutional and transformer networks on CIFAR10, CIFAR100 and TinyImageNet datasets.

cs.LG↗

Investigating the Bulge Morphology of Dual AGN Host Galaxies from the GOTHIC survey

We present a structural analysis of bulges in dual active galactic nuclei (AGN) host galaxies. Dual AGN arise in galaxy mergers where both supermassive black holes (SMBHs) are actively accreting. The AGN are typically embedded in compact bulges, which appear as luminous nuclei in optical images. Galaxy mergers can result in bulge growth, often via star formation. The bulges can be disky (pseudobulges), classical bulges, or belong to elliptical galaxies. Using SDSS DR18 gri images and GALFIT modelling, we performed 2D decomposition for 131 dual AGN bulges (comprising 61 galaxy pairs and 3 galaxy triplets) identified in the GOTHIC survey. We derived sérsic indices, luminosities, masses, and scalelengths of the bulges. Most bulges (105/131) are classical, with sérsic indices lying between $n=2$ and $n=8$. Among these, 64% are elliptical galaxies, while the remainder are classical bulges in disc galaxies. Only $\sim$20% of the sample exhibit pseudobulges. Bulge masses span $1.5\times10^9$ to $1.4\times10^{12}\,M_\odot$, with the most massive systems being ellipticals. Galaxy type matching shows that elliptical--elliptical (E--E) and elliptical--disc (E--D) mergers dominate over disc--disc (D--D) mergers. At least one galaxy in two-thirds of the dual AGN systems is elliptical and only $\sim$30% involve two disc galaxies. Although our sample is limited, our results suggest that dual AGN preferentially occur in evolved, red, quenched systems, that typically form via major mergers. They are predominantly hosted in classical bulges or elliptical galaxies rather than star-forming disc galaxies.

astro-ph.GA↗

Deconvolution for Large Astronomical Surveys: A Study of the Scaled Gradient Projection Method on Zwicky Transient Facility Data

Ground-based astronomical observations will continue to produce resolution-limited images due to atmospheric seeing. Deconvolution reverses such effects and thus can benefit extracted science in multifaceted ways. We apply the Scaled Gradient Projection (SGP) algorithm for the single-band deconvolution of several observed images from the Zwicky Transient Facility and mainly discuss the performance on stellar sources. The method shows good photometric flux preservation, which deteriorates for fainter sources but significantly reduces flux uncertainties even for the faintest sources. Deconvolved sources have a well-defined Full-Width-at-Half-Maximum (FWHM) of roughly one pixel (one arcsecond for ZTF) regardless of the observed seeing. Detection after deconvolution results in catalogs with $\gtrsim$99.6% completeness relative to detections in the observed images. A few observed sources that could not be detected in the deconvolved image are found near saturated sources, whereas for others, the deconvolved counterparts are detected when slightly different detection parameters are used. The deconvolution reveals new faint sources previously undetectable, which are confirmed by crossmatching with the deeper DESI Legacy DR10 and with Pan-STARRS1 through forced photometry. The method could identify examples of serendipitous potential deblends that exceeded SExtractor's deblending capabilities, with as extreme as $Δm \approx 3$ and separations as small as one arcsecond between the deblended components. Our survey-agnostic approach is better and eight times faster than Richardson-Lucy deconvolution and could be a reliable method for incorporation into survey pipelines.

astro-ph.IM↗

Benchmarking Anomaly Detection Algorithms: Deep Learning and Beyond

Detection of anomalous situations for complex mission-critical systems hold paramount importance when their service continuity needs to be ensured. A major challenge in detecting anomalies from the operational data arises due to the imbalanced class distribution problem since the anomalies are supposed to be rare events. This paper evaluates a diverse array of Machine Learning (ML)-based anomaly detection algorithms through a comprehensive benchmark study. The paper contributes significantly by conducting an unbiased comparison of various anomaly detection algorithms, spanning classical ML, including various tree-based approaches to Deep Learning (DL) and outlier detection methods. The inclusion of 104 publicly available enhances the diversity of the study, allowing a more realistic evaluation of algorithm performance and emphasizing the importance of adaptability to real-world scenarios. The paper evaluates the general notion of DL as a universal solution, showing that, while powerful, it is not always the best fit for every scenario. The findings reveal that recently proposed tree-based evolutionary algorithms match DL methods and sometimes outperform them in many instances of univariate data where the size of the data is small and number of anomalies are less than 10%. Specifically, tree-based approaches successfully detect singleton anomalies in datasets where DL falls short. To the best of the authors' knowledge, such a study on a large number of state-of-the-art algorithms using diverse data sets, with the objective of guiding researchers and practitioners in making informed algorithmic choices, has not been attempted earlier.

cs.LG↗

A Radon-Nikodým Perspective on Anomaly Detection: Theory and Implications

Which principle underpins the design of an effective anomaly detection loss function? The answer lies in the concept of Radon-Nikodým theorem, a fundamental concept in measure theory. The key insight from this article is: Multiplying the vanilla loss function with the Radon-Nikodým derivative improves the performance across the board. We refer to this as RN-Loss. We prove this using the setting of PAC (Probably Approximately Correct) learnability. Depending on the context a Radon-Nikodým derivative takes different forms. In the simplest case of supervised anomaly detection, Radon-Nikodým derivative takes the form of a simple weighted loss. In the case of unsupervised anomaly detection (with distributional assumptions), Radon-Nikodým derivative takes the form of the popular cluster based local outlier factor. We evaluate our algorithm on 96 datasets, including univariate and multivariate data from diverse domains, including healthcare, cybersecurity, and finance. We show that RN-Derivative algorithms outperform state-of-the-art methods on 68% of Multivariate datasets (based on F1 scores) and also achieves peak F1-scores on 72% of time series (Univariate) datasets.

cs.LG↗

Quantile Activation: Correcting a Failure Mode of ML Models

Standard ML models fail to infer the context distribution and suitably adapt. For instance, the learning fails when the underlying distribution is actually a mixture of distributions with contradictory labels. Learning also fails if there is a shift between train and test distributions. Standard neural network architectures like MLPs or CNNs are not equipped to handle this. In this article, we propose a simple activation function, quantile activation (QAct), that addresses this problem without significantly increasing computational costs. The core idea is to "adapt" the outputs of each neuron to its context distribution. The proposed quantile activation (QAct) outputs the relative quantile position of neuron activations within their context distribution, diverging from the direct numerical outputs common in traditional networks. A specific case of the above failure mode is when there is an inherent distribution shift, i.e the test distribution differs slightly from the train distribution. We validate the proposed activation function under covariate shifts, using datasets designed to test robustness against distortions. Our results demonstrate significantly better generalization across distortions compared to conventional classifiers and other adaptive methods, across various architectures. Although this paper presents a proof of concept, we find that this approach unexpectedly outperforms DINOv2 (small), despite DINOv2 being trained with a much larger network and dataset.

cs.LG↗

A Granger-Causal Perspective on Gradient Descent with Application to Pruning

Stochastic Gradient Descent (SGD) is the main approach to optimizing neural networks. Several generalization properties of deep networks, such as convergence to a flatter minima, are believed to arise from SGD. This article explores the causality aspect of gradient descent. Specifically, we show that the gradient descent procedure has an implicit granger-causal relationship between the reduction in loss and a change in parameters. By suitable modifications, we make this causal relationship explicit. A causal approach to gradient descent has many significant applications which allow greater control. In this article, we illustrate the significance of the causal approach using the application of Pruning. The causal approach to pruning has several interesting properties - (i) We observe a phase shift as the percentage of pruned parameters increase. Such phase shift is indicative of an optimal pruning strategy. (ii) After pruning, we see that minima becomes "flatter", explaining the increase in accuracy after pruning weights.

cs.LG↗

QuantProb: Generalizing Probabilities along with Predictions for a Pre-trained Classifier

Quantification of Uncertainty in predictions is a challenging problem. In the classification settings, although deep learning based models generalize well, class probabilities often lack reliability. Calibration errors are used to quantify uncertainty, and several methods exist to minimize calibration error. We argue that between the choice of having a minimum calibration error on original distribution which increases across distortions or having a (possibly slightly higher) calibration error which is constant across distortions, we prefer the latter We hypothesize that the reason for unreliability of deep networks is - The way neural networks are currently trained, the probabilities do not generalize across small distortions. We observe that quantile based approaches can potentially solve this problem. We propose an innovative approach to decouple the construction of quantile representations from the loss function allowing us to compute quantile based probabilities without disturbing the original network. We achieve this by establishing a novel duality property between quantiles and probabilities, and an ability to obtain quantile probabilities from any pre-trained classifier. While post-hoc calibration techniques successfully minimize calibration errors, they do not preserve robustness to distortions. We show that, Quantile probabilities (QuantProb), obtained from Quantile representations, preserve the calibration errors across distortions, since quantile probabilities generalize better than the naive Softmax probabilities.

cs.LG↗

DeliverAI: Reinforcement Learning Based Distributed Path-Sharing Network for Food Deliveries

Delivery of items from the producer to the consumer has experienced significant growth over the past decade and has been greatly fueled by the recent pandemic. Amazon Fresh, Shopify, UberEats, InstaCart, and DoorDash are rapidly growing and are sharing the same business model of consumer items or food delivery. Existing food delivery methods are sub-optimal because each delivery is individually optimized to go directly from the producer to the consumer via the shortest time path. We observe a significant scope for reducing the costs associated with completing deliveries under the current model. We model our food delivery problem as a multi-objective optimization, where consumer satisfaction and delivery costs, both, need to be optimized. Taking inspiration from the success of ride-sharing in the taxi industry, we propose DeliverAI - a reinforcement learning-based path-sharing algorithm. Unlike previous attempts for path-sharing, DeliverAI can provide real-time, time-efficient decision-making using a Reinforcement learning-enabled agent system. Our novel agent interaction scheme leverages path-sharing among deliveries to reduce the total distance traveled while keeping the delivery completion time under check. We generate and test our methodology vigorously on a simulation setup using real data from the city of Chicago. Our results show that DeliverAI can reduce the delivery fleet size by 12\%, the distance traveled by 13%, and achieve 50% higher fleet utilization compared to the baselines.

cs.LG↗

Strong convexity-guided hyper-parameter optimization for flatter losses

We propose a novel white-box approach to hyper-parameter optimization. Motivated by recent work establishing a relationship between flat minima and generalization, we first establish a relationship between the strong convexity of the loss and its flatness. Based on this, we seek to find hyper-parameter configurations that improve flatness by minimizing the strong convexity of the loss. By using the structure of the underlying neural network, we derive closed-form equations to approximate the strong convexity parameter, and attempt to find hyper-parameters that minimize it in a randomized fashion. Through experiments on 14 classification datasets, we show that our method achieves strong performance at a fraction of the runtime.

cs.LG↗

A study of two periodogram algorithms for improving the detection of small transiting planets

The sensitivities of two periodograms are compared for weak signal planet detection in transit surveys: the widely used Box-Least Squares (BLS) algorithm following light curve detrending and the Transit Comb Filter (TCF) algorithm following autoregressive ARIMA modeling. Small depth transits are injected into light curves with different simulated noise characteristics. Two measures of spectral peak significance are examined: the periodogram signal-to-noise ratio (SNR) and a False Alarm Probability (FAP) based on the generalized extreme value distribution. The relative performance of the BLS and TCF algorithms for small planet detection is examined for a range of light curve characteristics, including orbital period, transit duration, depth, number of transits, and type of noise. We find that the TCF periodogram applied to ARIMA fit residuals with the SNR detection metric is preferred when short-memory autocorrelation is present in the detrended light curve and even when the light curve noise had white Gaussian noise. BLS is more sensitive to small planets only under limited circumstances with the FAP metric. BLS periodogram characteristics are inferior when autocorrelated noise is present due to heteroscedastic noise and false period detection. Application of these methods to TESS light curves with known small exoplanets confirms our simulation results. The study ends with a decision tree that advises transit survey scientists on procedures to detect small planets most efficiently. The use of ARIMA detrending and TCF periodograms can significantly improve the sensitivity of any transit survey with regularly spaced cadence.

astro-ph.EP↗

A novel RNA pseudouridine site prediction model using Utility Kernel and data-driven parameters

RNA protein Interactions (RPIs) play an important role in biological systems. Recently, we have enumerated the RPIs at the residue level and have elucidated the minimum structural unit (MSU) in these interactions to be a stretch of five residues (Nucleotides/amino acids). Pseudouridine is the most frequent modification in RNA. The conversion of uridine to pseudouridine involves interactions between pseudouridine synthase and RNA. The existing models to predict the pseudouridine sites in a given RNA sequence mainly depend on user-defined features such as mono and dinucleotide composition/propensities of RNA sequences. Predicting pseudouridine sites is a non-linear classification problem with limited data points. Deep Learning models are efficient discriminators when the data set size is reasonably large and fail when there is a paucity of data ($<1000$ samples). To mitigate this problem, we propose a Support Vector Machine (SVM) Kernel based on utility theory from Economics, and using data-driven parameters (i.e. MSU) as features. For this purpose, we have used position-specific tri/quad/pentanucleotide composition/propensity (PSPC/PSPP) besides nucleotide and dineculeotide composition as features. SVMs are known to work well in small data regimes and kernels in SVM are designed to classify non-linear data. The proposed model outperforms the existing state-of-the-art models significantly (10%-15% on average).

q-bio.BM↗