SearcharxivSearch

arXiv subjects

Sudhir Kumar

Publications and source records attributed to Sudhir Kumar.

At least 19 recordsLinked to original sources

ESL-PSC Toolkit: a graphical software environment for linking shared genetic changes to convergent phenotypes

Convergent evolution provides a useful framework for testing whether independent origins of similar traits share common genetic mechanisms. Evolutionary Sparse Learning with Paired Species Contrast (ESL-PSC) is an approach to identify genes and sites associated with convergent traits from aligned sequences by fitting sparse predictive models to phylogenetically informed species contrasts. However, practical use of ESL-PSC currently requires substantial command-line fluency for data assembly, species-pair design, execution, and output interpretation. Here we present an integrated ESL-PSC analysis environment (ESL-PSC Toolkit) centered on a graphical user interface (GUI). ESL-PSC Toolkit is designed to assist users from experimental design through data interpretation without requiring extensive technical expertise. It supports guided input validation, interactive tree-based pair selection, command preview, live execution, post-run exploration of ranked genes and aligned sites, a complementary substitution-counting method, and analysis of continuous quantitative convergent traits. The computational backend has been reimplemented in Rust with many performance optimizations and parallelism, greatly reducing runtime for most analyses and enabling cross-platform packaged distributions. Downloadable GUI and CLI toolkit software packages for Mac, Windows, and Linux are available at https://github.com/John-Allard/ESL-PSC/releases/latest.

q-bio.PE

HyperEvoGen: Exploring deep phylogeny using non-Euclidean variational inference

Homologous proteins evolve from a common ancestral sequence, constrained by intricate patterns of co-evolving residues. Accurate reconstruction of evolutionary histories remains a challenge, primarily due to the inability of the existing approaches to capture long-range coevolutionary ties and lack of a precise metric to represent the evolutionary distance between sequences. Standard approaches are based on p-distance or substitution-corrected measures such as Jukes-Cantor. These methods saturate in cases of deep evolutionary divergence, losing all evolutionary signal after enough time. We present HyperEvoGen, a Poincar\'e variational autoencoder with adversarial training, hyperbolic latent geometry, and a compound loss function that learns evolutionarily meaningful representations from single-family alignments. The arrangement of protein sequences in HyperEvoGen's hyperbolic embedding aims to preserve phylogenetic structure and produce latent distances which scale with true evolutionary divergence. HyperEvoGen enables fast, scalable modeling of protein evolution while preserving hierarchical relatedness in a geometry-aware representation. On Potts-coupled simulation benchmarks, it produces more accurate ancestral reconstructions than conventional baselines, and offers higher-quality sequence generation with less training time than Potts models. This combination of accuracy and throughput supports large-family evolutionary studies and accelerates design-oriented applications.

q-bio.QM

Investigating Temporal Features in Swift GRB Afterglows: A Comparative Study of UVOT and XRT Data

This study presents a statistical analysis of optical light curves (LCs) of 200 UVOT-detected GRBs from 2005 to 2018. We have categorised these LCs based on their distinct morphological features, including early flares, bumps, breaks, plateaus, etc. Additionally, to compare features across different wavelengths, we have also included XRT LCs in our sample. The early observation capability of UVOT has allowed us to identify very early flares in 21 GRBs preceding the normal decay or bump, consistent with predictions of external reverse or internal shock. The decay indices of optical LCs following a simple power-law (PL) are shallower than corresponding X-ray LCs, indicative of a spectral break between two wavelengths. Not all LCs with PL decay align with the forward shock model and require additional components such as energy injection or a structured jet. Further, plateaus in the optical LCs are primarily consistent with energy injection from the central engine to the external medium. However, in four cases, plateaus followed by steep decay may have an internal origin. The optical luminosity observed during the plateau is tightly correlated with the break time, indicative of a magnetar as their possible central engine. For LCs with early bumps, the peak position, correlations between the parameters, and observed achromaticity allowed us to constrain their origin as the onset of afterglow, off-axis jet, late re-brightening, etc. In conclusion, the ensemble of observed features is explained through diverse physical mechanisms or emissions observed from different outflow locations and, in turn, diversity among possible progenitors.

astro-ph.HE

Treemble: A Graphical Tool to Generate Newick Strings from Phylogenetic Tree Images

Phylogenetic trees are ubiquitous and central to biology, but most published trees are available only as visual diagrams and not in the machine-readable newick format. There are thus thousands of published trees in the scientific literature that are unavailable for follow-up analyses, comparisons, supertree construction, etc. Experts can easily read such diagrams, but the manual construction of a newick string is prohibitively laborious. Previous attempts to semi-automate the reading of tree images relied on image processing techniques. These quickly encounter difficulties with typical published tree diagrams that contain various graphical elements that overlap the branches, such as error bars on internal nodes. Here we introduce Treemble, a user-friendly desktop application for generating newick strings from tree images. The user simply clicks to mark node locations, and Treemble algorithmically assembles the tree from the node coordinates alone. Tip nodes can be automatically detected and marked. Treemble also facilitates the automatic reading of tip name labels and can handle both rectangular and circular trees. Treemble is a native desktop application for both MacOS and Windows, and is freely available and fully documented at treemble.org.

q-bio.PE

Health Sentinel: An AI Pipeline For Real-time Disease Outbreak Detection

Early detection of disease outbreaks is crucial to ensure timely intervention by the health authorities. Due to the challenges associated with traditional indicator-based surveillance, monitoring informal sources such as online media has become increasingly popular. However, owing to the number of online articles getting published everyday, manual screening of the articles is impractical. To address this, we propose Health Sentinel. It is a multi-stage information extraction pipeline that uses a combination of ML and non-ML methods to extract events-structured information concerning disease outbreaks or other unusual health events-from online articles. The extracted events are made available to the Media Scanning and Verification Cell (MSVC) at the National Centre for Disease Control (NCDC), Delhi for analysis, interpretation and further dissemination to local agencies for timely intervention. From April 2022 till date, Health Sentinel has processed over 300 million news articles and identified over 95,000 unique health events across India of which over 3,500 events were shortlisted by the public health experts at NCDC as potential outbreaks.

cs.CL

Harnessing Layer-Controlled Two-dimensional Semiconductors for Photoelectrochemical Energy Storage via Quantum Capacitance and Band Nesting

Two-dimensional (2D) transition metal dichalcogenides like molybdenum diselenide (MoSe$_2$) have shown great potential in optoelectronics and energy storage due to their layer-dependent bandgap. However, producing high-quality 2D MoSe$_2$ layers in a scalable and controlled manner remains challenging. Traditional methods, such as hydrothermal and liquid-phase exfoliation, lack precision and understanding at the nanoscale, limiting further applications. Atmospheric pressure chemical vapor deposition (APCVD) offers a scalable solution for growing high-quality, large-area, layer-controlled 2D MoSe$_2$. Despite this, the photoelectrochemical performance of APCVD-grown 2D MoSe$_2$, particularly in energy storage, has not been extensively explored. This study addresses this by examining MoSe$_2$'s layer-dependent quantum capacitance and photo-induced charge storage properties. Using a three-electrode setup in 0.5M H$_2$SO$_4$, we observed a layer-dependent increase in areal capacitance under both dark and illuminated conditions. A six-layer MoSe$_2$ film exhibited the highest capacitance, reaching $96 \mu\mathrm{F/cm^2}$ in the dark and $115 \mu\mathrm{F/cm^2}$ under illumination at a current density of $5 \mu\mathrm{A/cm^2}$. Density Functional Theory (DFT) and Many-Body Perturbation Theory calculations reveal that Van Hove singularities and band nesting significantly enhance optical absorption and quantum capacitance. These results highlight APCVD-grown 2D MoSe$_2$'s potential as light-responsive, high-performance energy storage electrodes, paving the way for innovative energy storage systems.

cond-mat.mtrl-sci

R3F: An R package for evolutionary dates, rates, and priors using the relative rate framework

The relative rate framework (RRF) can estimate divergence times from branch lengths in a phylogeny, which is the theoretical basis of the RelTime method frequently applied, a relaxed clock approach for molecular dating that scales well for large phylogenies. The use of RRF has also enabled the development of computationally efficient and accurate methods for testing the autocorrelation of lineage rates in a phylogeny (CorrTest) and selecting data-driven parameters of the birth-death speciation model (ddBD), which can be used to specify priors in Bayesian molecular dating. We have developed R3F, an R package implementing RRF to estimate divergence times, infer lineage rates, conduct CorrTest, and build a ddBD tree prior for Bayesian dating in molecular phylogenies. Here, we describe R3F functionality and explain how to interpret and use its outputs in other visualization software and packages, such as MEGA, ggtree, and FigTree. Ultimately, R3F is intended to enable the dating of the Tree of Life with greater accuracy and precision, which would have important implications for studies of organism evolution, diversification dynamics, phylogeography, and biogeography. Availability and Implementation: The source codes and related instructions for installing and implementing R3F are available from GitHub (https://github.com/cathyqqtao/R3F).

q-bio.PE

MyESL: Sparse learning in molecular evolution and phylogenetic analysis

Evolutionary sparse learning (ESL) uses a supervised machine learning approach, Least Absolute Shrinkage and Selection Operator (LASSO), to build models explaining the relationship between a hypothesis and the variation across genomic features (e.g., sites) in sequence alignments. ESL employs sparsity between and within the groups of genomic features (e.g., genomic loci) by using sparse-group LASSO. Although some software packages are available for performing sparse group LASSO, we found them less well-suited for processing and analyzing genome-scale data containing millions of features, such as bases. MyESL software fills the need for open-source software for conducting ESL analyses with facilities to pre-process the input hypotheses and large alignments, make LASSO flexible and computationally efficient, and post-process the output model to produce different metrics useful in functional or evolutionary genomics. MyESL can take phylogenetic trees and sequence alignments as input and transform them into numeric responses and features, respecetively. The model outputs are processed into user-friendly text and graphical files. The computational core of MyESL is written in C++, which offers model building with or without group sparsity, while the pre- and post-processing of inputs and model outputs is performed using customized functions written in Python. One of its applications in phylogenomics showcases the utility of MyESL. Our analysis of empirical genome-scale datasets shows that MyESL can build evolutionary models quickly and efficiently on a personal desktop, while other computational packages were unable due to their prohibitive requirements of computational resources and time. MyESL is available for Python environments on Linux and distributed as a standalone application for Windows and macOS. It is available from https://github.com/kumarlabgit/MyESL.

q-bio.PE

Prompt and afterglow analysis of the Fermi-LAT detected GRB 230812B

Prompt emission of GRB 230812B stands out as one of the most luminous events observed by both the Fermi-GBM and LAT. Prompt emission spectral analysis (both time-integrated and resolved) of this burst supports an additional thermal component together with a non-thermal, indicating the hybrid jet composition. The spectral parameters alpha, Ep, and kT of the best-fit Band+Blackbody model show a tacking behaviour with the intensity. Further, the low energy afterglow emission is consistent with the synchrotron emission from the external forward shock in the ISM medium. LAT detected very high energy emission (VHE) deviating from the synchrotron mechanism, possibly originating from the Lorentz boosting of prompt emission photons by accelerated electrons in the external shock via Inverse Compton (IC) or Synchrotron Self Compton (SSC) emission mechanisms. The comparison of the prompt and afterglow emission properties of this burst revealed that, unlike the bright prompt emission, the afterglow of GRB 230812B is fainter than the other SN-detected bright bursts (GRB 130427A and GRB 171010A) at a similar redshift.

astro-ph.HE

Exploring Origin of Ultra-Long Gamma-ray Bursts: Lessons from GRB 221009A

The brightest Gamma-ray burst (GRB) ever, GRB 221009A, displays ultra-long GRB (ULGRB) characteristics, with a prompt emission duration exceeding 1000 s. To constrain the origin and central engine of this unique burst, we analyze its prompt and afterglow characteristics and compare them to the established set of similar GRBs. To achieve this, we statistically examine a nearly complete sample of Swift-detected GRBs with measured redshifts. Categorizing the sample to Bronze, Silver, and Gold by fitting a Gaussian function to the log-normal of T$_{90}$ duration distribution and considering three sub-samples respectively to 1, 2, and 3 times of the standard deviation to the mean value. GRB 221009A falls into the Gold sub-sample. Our analysis of prompt emission and afterglow characteristics aims to identify trends between the three burst groups. Notably, the Gold sub-sample (a higher likelihood of being ULGRB candidates) suggests a collapsar scenario with a hyper-accreting black hole as a potential central engine, while a few GRBs (GRB 060218, GRB 091024A, and GRB 100316D) in our Gold sub-sample favor a magnetar. Late-time near-IR (NIR) observations from 3.6m Devasthal Optical Telescope (DOT) rule out the presence of any bright supernova associated with GRB 221009A in the Gold sub-sample. To further constrain the physical properties of ULGRB progenitors, we employ the tool MESA to simulate the evolution of low-metallicity massive stars with different initial rotations. The outcomes suggest that rotating ($\Omega \geq 0.2\,\Omega_{\rm c}$) massive stars could potentially be the progenitors of ULGRBs within the considered parameters and initial inputs to MESA.

astro-ph.HE

Nanomolecular OLED Pixelization Enabling Electroluminescent Metasurfaces

Miniaturization of light-emitting diodes (LEDs) can enable high-resolution augmented and virtual reality displays and on-chip light sources for ultra-broadband chiplet communication. However, unlike silicon scaling in electronic integrated circuits, patterning of inorganic III-V semiconductors in LEDs considerably compromises device efficiencies at submicrometer scales. Here, we present the scalable fabrication of nanoscale organic LEDs (nano-OLEDs), with the highest array density (>84,000 pixels per inch) and the smallest pixel size (~100 nm) ever reported to date. Direct nanomolecular patterning of organic semiconductors is realized by self-aligned evaporation through nanoapertures fabricated on a free-standing silicon nitride film adhering to the substrate. The average external quantum efficiencies (EQEs) extracted from a nano-OLED device of more than 4 megapixels reach up to 10%. At the subwavelength scale, individual pixels act as electroluminescent meta-atoms forming metasurfaces that directly convert electricity into modulated light. The diffractive coupling between nano-pixels enables control over the far-field emission properties, including directionality and polarization. The results presented here lay the foundation for bright surface light sources of dimension smaller than the Abbe diffraction limit, offering new technological platforms for super-resolution imaging, spectroscopy, sensing, and hybrid integrated photonics.

physics.optics

Performance analysis of InAlN/GaN HEMT and optimization for high frequency applications

An InAlN/GaN HEMT device was studied using extensive temperature dependent DC IV measurements and CV measurements. Barrier traps in the InAlN layer were characterized using transient analysis. Forward gate current was modelled using analytical equations. RF performance of the device was also studied and device parameters were extracted following small signal equivalent circuit model. Extensive simulations in Silvaco TCAD were also carried out by varying stem height, gate length and incorporating back barrier to optimize the suitability of this device in Ku-band by reducing the detrimental Short Channel Effects (SCEs). In this paper a novel structure i.e., a short length T gate with recess, on thin GaN buffer to achieve high cut-off frequency (f$_T$) and high maximum oscillating frequency (f$_{max}$) apt for Ku-band applications is also proposed.

physics.app-ph

Investigation of RF performance of Ku-band GaN HEMT device and an in-depth analysis of short channel effects

In this paper, we have characterized an AlGaN/GaN High Electron Mobility Transistor (HEMT) with a short gate length (Lg $\approx$ 0.15$\mu$m). We have studied the effect of short gate length on the small signal parameters, linearity parameters and gm-gd ratio in GaN HEMT devices. To understand how scaling results in the variation of the above-mentioned parameters a comparative study with higher gate length devices on similar heterostructure is also presented here. We have scaled down the gate length but the barrier thickness(t$_{bar}$) remained same which affects the aspect ratio (L$_{g}$/t$_{bar}$) of the device and its inseparable consequences are the prominent short channel effects (SCEs) barring the optimum output performance of the device. These interesting phenomena were studied in detail and explored over a temperature range of -40$^\circ$C to 80$^\circ$C. To the best of our knowledge this paper explores temperature dependence of SCEs of GaN HEMT for the first time. With an approach to reduce the impact of SCEs a simulation study in Silvaco TCAD was carried out and it is observed that a recessed gate structure on conventional heterostructure successfully reduces SCEs and improves RF performance of the device. This work gives an overall view of gate length scaling on conventional AlGaN/GaN HEMTs.

physics.app-ph

Deep Low-Shot Learning for Biological Image Classification and Visualization from Limited Training Samples

Predictive modeling is useful but very challenging in biological image analysis due to the high cost of obtaining and labeling training data. For example, in the study of gene interaction and regulation in Drosophila embryogenesis, the analysis is most biologically meaningful when in situ hybridization (ISH) gene expression pattern images from the same developmental stage are compared. However, labeling training data with precise stages is very time-consuming even for evelopmental biologists. Thus, a critical challenge is how to build accurate computational models for precise developmental stage classification from limited training samples. In addition, identification and visualization of developmental landmarks are required to enable biologists to interpret prediction results and calibrate models. To address these challenges, we propose a deep two-step low-shot learning framework to accurately classify ISH images using limited training images. Specifically, to enable accurate model training on limited training samples, we formulate the task as a deep low-shot learning problem and develop a novel two-step learning approach, including data-level learning and feature-level learning. We use a deep residual network as our base model and achieve improved performance in the precise stage prediction task of ISH images. Furthermore, the deep model can be interpreted by computing saliency maps, which consist of pixel-wise contributions of an image to its prediction result. In our task, saliency maps are used to assist the identification and visualization of developmental landmarks. Our experimental results show that the proposed model can not only make accurate predictions, but also yield biologically meaningful interpretations. We anticipate our methods to be easily generalizable to other biological image classification tasks with small training datasets.

cs.LG

Image-based phenotyping of diverse Rice (Oryza Sativa L.) Genotypes

Development of either drought-resistant or drought-tolerant varieties in rice (Oryza sativa L.), especially for high yield in the context of climate change, is a crucial task across the world. The need for high yielding rice varieties is a prime concern for developing nations like India, China, and other Asian-African countries where rice is a primary staple food. The present investigation is carried out for discriminating drought tolerant, and susceptible genotypes. A total of 150 genotypes were grown under controlled conditions to evaluate at High Throughput Plant Phenomics facility, Nanaji Deshmukh Plant Phenomics Centre, Indian Council of Agricultural Research-Indian Agricultural Research Institute, New Delhi. A subset of 10 genotypes is taken out of 150 for the current investigation. To discriminate against the genotypes, we considered features such as the number of leaves per plant, the convex hull and convex hull area of a plant-convex hull formed by joining the tips of the leaves, the number of leaves per unit convex hull of a plant, canopy spread - vertical spread, and horizontal spread of a plant. We trained You Only Look Once (YOLO) deep learning algorithm for leaves tips detection and to estimate the number of leaves in a rice plant. With this proposed framework, we screened the genotypes based on selected traits. These genotypes were further grouped among different groupings of drought-tolerant and drought susceptible genotypes using the Ward method of clustering.

cs.CV

$c^+$GAN: Complementary Fashion Item Recommendation

We present a conditional generative adversarial model to draw realistic samples from paired fashion clothing distribution and provide real samples to pair with arbitrary fashion units. More concretely, given an image of a shirt, obtained from a fashion magazine, a brochure or even any random click on ones phone, we draw realistic samples from a parameterized conditional distribution learned as a conditional generative adversarial network ($c^+$GAN) to generate the possible pants which can go with the shirt. We start with a classical cGAN model as proposed by Mirza and Osindero [arXiv:1411.1784] and modify both the generator and discriminator to work on captured-in-the-wild data with no human alignment. We gather a dataset from web crawled data, systematically develop a method which counters the problems inherent to such data, and finally present plausible results based on our technique. We propose simple ideas to evaluate how these techniques can conquer the cognitive gap that exists when arbitrary clothing articles need to be paired with another relevant article, based on similarity of search results.

cs.CV

A Review of Localization and Tracking Algorithms in Wireless Sensor Networks

In this paper, a comprehensive survey of the pioneer as well as the state of-the-art localization and tracking methods in the wireless sensor networks is presented. Localization is mostly applicable for the static sensor nodes, whereas, tracking for the mobile sensor nodes. The localization algorithms are broadly classified as range-based and range-free methods. The estimated range (distance) between an anchor and an unknown node is highly erroneous in an indoor scenario. This limitation can be handled up to a large extent by employing a large number of existing access points (APs) in the range free localization method. Recent works emphasize on the use multi-sensor data like magnetic, inertial, compass, gyroscope, ultrasound, infrared, visual and/or odometer to improve the localization accuracy further. Additionally, tracking method does the future prediction of location based on the past location history. A smooth trajectory is noted even if some of the received measurements are erroneous. Real experimental set-ups such as National Instruments (NI) wireless sensor nodes, Crossbow motes and hand-held devices for carrying out the localization and tracking are also highlighted herein.

cs.NI

Significance of Mobility on Received Signal Strength: An Experimental Investigation

In this paper, estimation of mobility using received signal strength is presented. In contrast to standard methods, speed can be inferred without the use of any additional hardware like accelerometer, gyroscope or position estimator. The strength of Wi-Fi signal is considered herein to compute the time-domain features such as mean, minimum, maximum, and autocorrelation. The experiments are carried out in different environments like academic area, residential area and in open space. The complexity of the algorithm in training and testing phase are quadratic and linear with the number of Wi-Fi samples respectively. The experimental results indicate that the average error in the estimated speed is 12 % when the maximum signal strength features are taken into account. The proposed method is cost-effective and having a low complexity with reasonable accuracy in a Wi-Fi or cellular environment. Additionally, the proposed method is scalable that is the performance is not affected in a multi-smartphones scenario.

cs.NI