SearcharxivSearch

arXiv subjects

Divyansh Srivastava

Publications and source records attributed to Divyansh Srivastava.

10 recordsLinked to original sources

Joint Observations of PTPS Targets (JOTA). Combined spectroscopic and photometric analysis of 16 SB1 systems

Context. The Pennsylvania-Toruń Planet Search, which operates on a sample of $\sim$1000 northern-hemisphere stars, has been underway since 2004. Stars with radial velocity (RV) amplitudes exceeding 2 km s$^{-1}$ represent a backup subprogram sample dedicated to monitoring binary stars with multiple instruments. Aims. We used almost 21 years of combined RV measurements to search for Doppler signals consistent with stellar or brown dwarf companions and to produce a catalog of both known and previously unpublished binary stars in our planet-search sample. Methods. We analyzed the combined RV measurements, searching for stellar companions and obtaining orbital solutions for both known and new binary systems. We also searched for periods in available long-term photometric monitoring data from All Sky Automated Survey (ASAS), and applied Gaia astrometric data whenever available. Results. We report the results of long-term RV monitoring of 16 systems: new detections of low-mass companions to 11 stars and updated orbital elements based on combined sets of our new and literature data for five systems. For two objects (TYC 3318-00789-1 and TYC 3318-01538-1), we find evidence of activity or the influence of another unseen companion, possibly due to line-profile variations or unresolved spectral contamination. We present true masses for three systems in our sample: TYC 1931-1040-1, TYC 3451- 1449-1, and TYC 3667-1636-1. We obtained these masses using Gaia DR3 non-single star data.

astro-ph.SR

Estimating stellar metallicities from Gaia DR3 XP data using LAMOST DR10

Gaia DR3 provides astrophysical parameters for hundreds of millions of stars, but the metallicities [M/H] from its GSP-Phot module suffer from systematic biases. We estimate stellar metallicities from Gaia DR3 data using the homogeneous spectroscopic iron abundances [Fe/H] of LAMOST DR10 as training labels. We cross-matched LAMOST DR10 with Gaia DR3 and trained a gradient-boosted decision-tree regressor (XGBoost) on 1.20 million AFGK stars using only Gaia-derived inputs and proxies. We validated the estimates on held-out LAMOST stars, GALAH DR4, APOGEE DR17, and 46 open clusters, and applied the model to measure the radial metallicity gradient of the Milky Way disk. On the held-out test set, the model achieves a mean absolute error of 0.052 dex and $R^2=0.94$ with negligible bias, compared with 0.242 dex for GSP-Phot on the same stars. The estimates transfer well to external surveys, with mean absolute errors of 0.066 dex for GALAH and 0.068 dex for APOGEE. For open clusters, the median difference between our estimated [Fe/H] and spectroscopic values is 0.041 dex, smaller than both GSP-Phot (0.248 dex) and a previous APOGEE-trained XGBoost model (0.067 dex). Applied to the Galactic disk, our model recovers a broken thin-disk radial gradient, with inner and outer slopes of $+0.119$ and $-0.058\,\mathrm{dex\,kpc^{-1}}$, respectively, and a break near 5.9 kpc, as well as an open-cluster gradient of $-0.066\,\mathrm{dex\,kpc^{-1}}$; both agree with previous high-resolution spectroscopic studies. Our [Fe/H] estimates are accurate to 0.05-0.07 dex for AFGK stars with $[\mathrm{Fe/H}]\gtrsim-2.5$; below this limit, the predictions should be treated as lower bounds. The catalogue and trained model are publicly available on Zenodo and are suitable for chemical studies of the Milky Way.

astro-ph.GA

Agentic AI-based Framework for Mitigating Premature Diagnostic Handoff and Silent Hallucination in Healthcare Applications

Recent advances in Large Language Models (LLMs) and multi-agent systems have driven the rise of Agentic AI, showing promise for medical reasoning. However, open-ended conversational agents remain prone to two critical failure modes: premature diagnostic handoff and silent clinical hallucinations that may go undetected before reaching the patient. In this work, we propose a multi-agent framework that addresses both issues by replacing ``LLM-as-a-judge'' routing with deterministic orchestration constraints. The framework incorporates two safety mechanisms. First, a neuro-symbolic state-tracking gate enforces completeness of the OLDCARTS clinical protocol (Onset, Location, Duration, Character, Aggravating/Alleviating factors, Radiation, Timing, and Severity) by blocking diagnostic transitions until all required dimensions are collected. Second, an epistemic uncertainty quantification (UQ) gate computes semantic entropy (H) across K=5 independent diagnostic samples to identify and intercept divergent outputs before delivery. We evaluate the system using simulated patient agents powered by the llama-3.1-70b-instruct model on 150 test cases. The full architecture achieves 49.3% diagnostic precision, representing an absolute improvement of 11.3 percentage points over an unconstrained baseline. Additionally, we observe a statistically significant negative correlation (r = -0.181, p < 0.05) between OLDCARTS completeness (σ) and semantic entropy (H), suggesting that structured information gathering is associated with reduced diagnostic uncertainty.

cs.AI

DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation

Decoder-only autoregressive image generation typically relies on fixed-length tokenization schemes whose token counts grow quadratically with resolution, substantially increasing the computational and memory demands of attention. We present DPAR, a novel decoder-only autoregressive model that dynamically aggregates image tokens into a variable number of patches for efficient image generation. Our work is the first to demonstrate that next-token prediction entropy from a lightweight and unsupervised autoregressive model provides a reliable criterion for merging tokens into larger patches based on information content. DPAR makes minimal modifications to the standard decoder architecture, ensuring compatibility with multimodal generation frameworks and allocating more compute to generation of high-information image regions. Further, we demonstrate that training with dynamically sized patches yields representations that are robust to patch boundaries, allowing DPAR to scale to larger patch sizes at inference. DPAR reduces token count by 1.81x and 2.06x on Imagenet 256 and 384 generation resolution respectively, leading to a reduction of up to 40% FLOPs in training costs. Further, our method exhibits faster convergence and improves FID by up to 27.1% relative to baseline models.

cs.CV

Vision Transformer for Transient Noise Classification

Transient noise (glitches) in LIGO data hinders the detection of gravitational waves (GW). The Gravity Spy project has categorized these noise events into various classes. With the O3 run, there is the inclusion of two additional noise classes and thus a need to train new models for effective classification. We aim to classify glitches in LIGO data into 22 existing classes from the first run plus 2 additional noise classes from O3a using the Vision Transformer (ViT) model. We train a pre-trained Vision Transformer (ViT-B/32) model on a combined dataset consisting of the Gravity Spy dataset with the additional two classes from the LIGO O3a run. We achieve a classification efficiency of 92.26%, demonstrating the potential of Vision Transformer to improve the accuracy of gravitational wave detection by effectively distinguishing transient noise. Key words: gravitational waves --vision transformer --machine learning

cs.CV

Low-mass companions to nine stars

We present an independent spectroscopic and radial velocity analysis for nine stars from the Pennsylvania-Toruń Planet Search. For BD+24 4697, we present an updated true companion's mass (0.16$\pm$0.02 \, M$_{\odot}$), as well as evidence of stellar activity. For BD+54 1640 and BD+65 1241 we present true masses of companions, $m = 0.15 \pm 0.04\,M_\odot$ and $m = 0.091 \pm 0.005\,M_\odot$, respectively. For BD+63 974 and BD+69 935 we find low mass companions with $m \sin i = 0.046 \pm 0.001\,M_\odot$ and $m \sin i = 0.090 \pm 0.005\,M_\odot$. For BD+52 1281, BD+54 1382, TYC 2704-2680-1, and TYC 3525-02043-1 we present evidence of low-mass companions with $m \sin i$ of 0.115 $\pm 0.006\,M_\odot$, 0.083 $\pm 0.007\,M_\odot$, 0.279 $\pm 0.009\,M_\odot$, and $0.064 \pm 0.006\,M_\odot$, respectively. Consequently, BD+54 1382, BD+63 974, BD+65 1241, BD+69 935 and TYC 3525-02043-1 appear to be Brown Dwarf host candidates.

astro-ph.SR

OverLayBench: A Benchmark for Layout-to-Image Generation with Dense Overlaps

Despite steady progress in layout-to-image generation, current methods still struggle with layouts containing significant overlap between bounding boxes. We identify two primary challenges: (1) large overlapping regions and (2) overlapping instances with minimal semantic distinction. Through both qualitative examples and quantitative analysis, we demonstrate how these factors degrade generation quality. To systematically assess this issue, we introduce OverLayScore, a novel metric that quantifies the complexity of overlapping bounding boxes. Our analysis reveals that existing benchmarks are biased toward simpler cases with low OverLayScore values, limiting their effectiveness in evaluating model performance under more challenging conditions. To bridge this gap, we present OverLayBench, a new benchmark featuring high-quality annotations and a balanced distribution across different levels of OverLayScore. As an initial step toward improving performance on complex overlaps, we also propose CreatiLayout-AM, a model fine-tuned on a curated amodal mask dataset. Together, our contributions lay the groundwork for more robust layout-to-image generation under realistic and challenging scenarios. Project link: https://mlpc-ucsd.github.io/OverLayBench.

cs.CV

Lay-Your-Scene: Natural Scene Layout Generation with Diffusion Transformers

We present Lay-Your-Scene (shorthand LayouSyn), a novel text-to-layout generation pipeline for natural scenes. Prior scene layout generation methods are either closed-vocabulary or use proprietary large language models for open-vocabulary generation, limiting their modeling capabilities and broader applicability in controllable image generation. In this work, we propose to use lightweight open-source language models to obtain scene elements from text prompts and a novel aspect-aware diffusion Transformer architecture trained in an open-vocabulary manner for conditional layout generation. Extensive experiments demonstrate that LayouSyn outperforms existing methods and achieves state-of-the-art performance on challenging spatial and numerical reasoning benchmarks. Additionally, we present two applications of LayouSyn. First, we show that coarse initialization from large language models can be seamlessly combined with our method to achieve better results. Second, we present a pipeline for adding objects to images, demonstrating the potential of LayouSyn in image editing applications.

cs.CV

VLG-CBM: Training Concept Bottleneck Models with Vision-Language Guidance

Concept Bottleneck Models (CBMs) provide interpretable prediction by introducing an intermediate Concept Bottleneck Layer (CBL), which encodes human-understandable concepts to explain models' decision. Recent works proposed to utilize Large Language Models and pre-trained Vision-Language Models to automate the training of CBMs, making it more scalable and automated. However, existing approaches still fall short in two aspects: First, the concepts predicted by CBL often mismatch the input image, raising doubts about the faithfulness of interpretation. Second, it has been shown that concept values encode unintended information: even a set of random concepts could achieve comparable test accuracy to state-of-the-art CBMs. To address these critical limitations, in this work, we propose a novel framework called Vision-Language-Guided Concept Bottleneck Model (VLG-CBM) to enable faithful interpretability with the benefits of boosted performance. Our method leverages off-the-shelf open-domain grounded object detectors to provide visually grounded concept annotation, which largely enhances the faithfulness of concept prediction while further improving the model performance. In addition, we propose a new metric called Number of Effective Concepts (NEC) to control the information leakage and provide better interpretability. Extensive evaluations across five standard benchmarks show that our method, VLG-CBM, outperforms existing methods by at least 4.27% and up to 51.09% on Accuracy at NEC=5 (denoted as ANEC-5), and by at least 0.45% and up to 29.78% on average accuracy (denoted as ANEC-avg), while preserving both faithfulness and interpretability of the learned concepts as demonstrated in extensive experiments.

cs.CV

Corrupting Neuron Explanations of Deep Visual Features

The inability of DNNs to explain their black-box behavior has led to a recent surge of explainability methods. However, there are growing concerns that these explainability methods are not robust and trustworthy. In this work, we perform the first robustness analysis of Neuron Explanation Methods under a unified pipeline and show that these explanations can be significantly corrupted by random noises and well-designed perturbations added to their probing data. We find that even adding small random noise with a standard deviation of 0.02 can already change the assigned concepts of up to 28% neurons in the deeper layers. Furthermore, we devise a novel corruption algorithm and show that our algorithm can manipulate the explanation of more than 80% neurons by poisoning less than 10% of probing data. This raises the concern of trusting Neuron Explanation Methods in real-life safety and fairness critical applications.

cs.LG