SearcharxivSearch

arXiv subjects

Mainak Singha

Publications and source records attributed to Mainak Singha.

At least 19 recordsLinked to original sources

The Accretion Explorer Interferometer (AEI) Phase I NASA Innovative Advanced Concepts Final Report

We must create superb X-ray images to understand the detailed physical processes behind some of the most powerful astronomical objects. The need to achieve this capability has been known for decades. But as time proceeds, X-ray astronomy falls further behind other wavebands that are steadily increasing their imaging capacity. Radio astronomy in particular has reached an angular resolution on the order of micro arcseconds via aperture synthesis interferometry, using the interference of electromagnetic waves from many small telescopes together to simulate having a much larger telescope. Developing an equivalent high-resolution capability in the X-ray band would be a game changer for high-energy astrophysics. We will understand how supermassive black holes grow and evolve. We will learn what powers astrophysical jets. We will learn how young, active stars affect the habitability of their planets. Technologically, our NIAC study has shown that the Accretion Explorer Interferometer (AEI) concept, unlike the original MAXIM concept, is more feasible in operation, being only 2 km long, versus approximately 450 km. Our study has also shown that satellite station keeping is possible, leveraging from LISA pathfinder technology, and using large mirror flats plus an X-ray beamsplitter for enabling technology is feasible.

astro-ph.IM

DistMoE: Private-data Rehearsal-free Routing in Mixture-of-Experts for Distributed Instruction Tuning

Multimodal Large Language Models (MLLMs) have shown strong multimodal instruction-following ability, but adapting them to diverse visual-language domains typically assumes centralized data access and costly joint training. This is restrictive when data is distributed across private, domain-specific, or permission-limited clients. To this end, we propose DistMoE, a mixture-of-experts (MoE) approach for distributed visual instruction tuning. In each layer of the language decoder it augments the public feedforward network (FFN) with a client-specific private FFN expert, with the goal to acquire domain-specific knowledge. However, independent expert training causes the private FFNs to learn representation of different scale and magnitudes, making merging the experts difficult. To reduce client-specific drift, we introduce a public-anchored expert composition stage that updates only routers and lightweight private projection adapters on a mix of local client data and public data, via an isotropic regularization loss, therefore making it cross-client rehearsal-free composition. During inference, DistMoE performs modular routing over public and private experts, enabling token-wise domain composition without explicit domain labels. Experiments across diverse visual-language benchmarks show that DistMoE enables flexible expert reuse, effective domain adaptation, and competitive performance while preserving modular control over client-specific knowledge. Codes are available at https://github.com/mainaksingha01/DistMoE.

cs.CV

BioVLM: Routing Prompts, Not Parameters, for Cross-Modality Generalization in Biomedical VLMs

Pretrained biomedical vision-language models (VLMs) such as BioMedCLIP perform well on average but often degrade on challenging modalities where inter-class margins are small and acquisition-specific variations are pronounced, especially under few-shot supervision and when modality priors differ from pretraining corpora substantially. We propose BioVLM, a prompt-learning framework that improves cross-domain generalization without extensive backbone fine-tuning. BioVLM learns a diverse prompt bank and introduces dynamic prompt selection: for each input, it selects the most discriminative prompts via a low-entropy criterion on the predictive distribution, effectively coupling sparse few-shot evidence with rich LLM semantic priors. To strengthen this coupling, we distill high-confidence LLM-derived attributes and enforce robust knowledge transfer through strong/weak augmentation consistency. At test time, BioVLM adapts by choosing modality-appropriate prompts, enabling transfer to unseen categories and domains, while keeping training lightweight and inference efficient. On 11 MedMNIST+ 2D datasets, BioVLM achieves new state of the art across three distinct generalization settings. Codes are available at https://github.com/mainaksingha01/BioVLM.

cs.CV

GeoMeld: Toward Semantically Grounded Foundation Models for Remote Sensing

Effective foundation modeling in remote sensing requires spatially aligned heterogeneous modalities coupled with semantically grounded supervision, yet such resources remain limited at scale. We present GeoMeld, a large-scale multimodal dataset with approximately 2.5 million spatially aligned samples. The dataset spans diverse modalities and resolutions and is constructed under a unified alignment protocol for modality-aware representation learning. GeoMeld provides semantically grounded language supervision through an agentic captioning framework that synthesizes and verifies annotations from spectral signals, terrain statistics, and structured geographic metadata, encoding measurable cross-modality relationships within textual descriptions. To leverage this dataset, we introduce GeoMeld-FM, a pretraining framework that combines multi-pretext masked autoencoding over aligned modalities, JEPA representation learning, and caption-vision contrastive alignment. This joint objective enables the learned representation space to capture both reliable cross-sensor physical consistency and grounded semantics. Experiments demonstrate consistent gains in downstream transfer and cross-sensor robustness. Together, GeoMeld and GeoMeld-FM establish a scalable reference framework for semantically grounded multi-modal foundation modeling in remote sensing.

cs.CV

CLIPoint3D: Language-Grounded Few-Shot Unsupervised 3D Point Cloud Domain Adaptation

Recent vision-language models (VLMs) such as CLIP demonstrate impressive cross-modal reasoning, extending beyond images to 3D perception. Yet, these models remain fragile under domain shifts, especially when adapting from synthetic to real-world point clouds. Conventional 3D domain adaptation approaches rely on heavy trainable encoders, yielding strong accuracy but at the cost of efficiency. We introduce CLIPoint3D, the first framework for few-shot unsupervised 3D point cloud domain adaptation built upon CLIP. Our approach projects 3D samples into multiple depth maps and exploits the frozen CLIP backbone, refined through a knowledge-driven prompt tuning scheme that integrates high-level language priors with geometric cues from a lightweight 3D encoder. To adapt task-specific features effectively, we apply parameter-efficient fine-tuning to CLIP's encoders and design an entropy-guided view sampling strategy for selecting confident projections. Furthermore, an optimal transport-based alignment loss and an uncertainty-aware prototype alignment loss collaboratively bridge source-target distribution gaps while maintaining class separability. Extensive experiments on PointDA-10 and GraspNetPC-10 benchmarks show that CLIPoint3D achieves consistent 3-16% accuracy gains over both CLIP-based and conventional encoder-based baselines. Project page: https://sarthakm320.github.io/CLIPoint3D.

cs.CV

bi-modal textual prompt learning for vision-language models in remote sensing

Prompt learning (PL) has emerged as an effective strategy to adapt vision-language models (VLMs), such as CLIP, for downstream tasks under limited supervision. While PL has demonstrated strong generalization on natural image datasets, its transferability to remote sensing (RS) imagery remains underexplored. RS data present unique challenges, including multi-label scenes, high intra-class variability, and diverse spatial resolutions, that hinder the direct applicability of existing PL methods. In particular, current prompt-based approaches often struggle to identify dominant semantic cues and fail to generalize to novel classes in RS scenarios. To address these challenges, we propose BiMoRS, a lightweight bi-modal prompt learning framework tailored for RS tasks. BiMoRS employs a frozen image captioning model (e.g., BLIP-2) to extract textual semantic summaries from RS images. These captions are tokenized using a BERT tokenizer and fused with high-level visual features from the CLIP encoder. A lightweight cross-attention module then conditions a learnable query prompt on the fused textual-visual representation, yielding contextualized prompts without altering the CLIP backbone. We evaluate BiMoRS on four RS datasets across three domain generalization (DG) tasks and observe consistent performance gains, outperforming strong baselines by up to 2% on average. Codes are available at https://github.com/ipankhi/BiMoRS.

cs.CV

The Need for Ultra High Resolution X-ray Imaging

This paper discusses the broad science case for obtaining milliarcsecond to microarcsecond astronomical imaging resolution in the soft to medium-energy X-ray band (~0.5 to ~8 keV). Astronomy across much of the electromagnetic spectrum has been fundamentally transformed with a rapid increase in ground-based and space-based capabilities to examine celestial objects on small scales that relate directly to their relevant physical processes. X-ray imaging capabilities, however, have fallen far behind observations at longer wavelengths. As such, without decisive advances in X-ray imaging, we will be unable to uncover key phenomena on the smallest astrophysical scales, leaving entire classes of high-energy discoveries beyond our reach. Here we describe several science goals for which high quality X-ray imaging is crucial and the status of some current technologies or mission concepts that would be required for these advances. In particular, we discuss the Accretion Explorer, a mission architecture under current study for a dispersed aperture X-ray interferometer.

astro-ph.HE

MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP

In this paper, we propose a novel multimodal framework, Multimodal Language-Guided Network (MMLGNet), to align heterogeneous remote sensing modalities like Hyperspectral Imaging (HSI) and LiDAR with natural language semantics using vision-language models such as CLIP. With the increasing availability of multimodal Earth observation data, there is a growing need for methods that effectively fuse spectral, spatial, and geometric information while enabling semantic-level understanding. MMLGNet employs modality-specific encoders and aligns visual features with handcrafted textual embeddings in a shared latent space via bi-directional contrastive learning. Inspired by CLIP's training paradigm, our approach bridges the gap between high-dimensional remote sensing data and language-guided interpretation. Notably, MMLGNet achieves strong performance with simple CNN-based encoders, outperforming several established multimodal visual-only methods on two benchmark datasets, demonstrating the significant benefit of language supervision. Codes are available at https://github.com/AdityaChaudhary2913/CLIP_HSI.

cs.CV

Reconstruction Guided Few-shot Network For Remote Sensing Image Classification

Few-shot remote sensing image classification is challenging due to limited labeled samples and high variability in land-cover types. We propose a reconstruction-guided few-shot network (RGFS-Net) that enhances generalization to unseen classes while preserving consistency for seen categories. Our method incorporates a masked image reconstruction task, where parts of the input are occluded and reconstructed to encourage semantically rich feature learning. This auxiliary task strengthens spatial understanding and improves class discrimination under low-data settings. We evaluated the efficacy of EuroSAT and PatternNet datasets under 1-shot and 5-shot protocols, our approach consistently outperforms existing baselines. The proposed method is simple, effective, and compatible with standard backbones, offering a robust solution for few-shot remote sensing classification. Codes are available at https://github.com/stark0908/RGFS.

cs.CV

SDHSI-Net: Learning Better Representations for Hyperspectral Images via Self-Distillation

Hyperspectral image (HSI) classification presents unique challenges due to its high spectral dimensionality and limited labeled data. Traditional deep learning models often suffer from overfitting and high computational costs. Self-distillation (SD), a variant of knowledge distillation where a network learns from its own predictions, has recently emerged as a promising strategy to enhance model performance without requiring external teacher networks. In this work, we explore the application of SD to HSI by treating earlier outputs as soft targets, thereby enforcing consistency between intermediate and final predictions. This process improves intra-class compactness and inter-class separability in the learned feature space. Our approach is validated on two benchmark HSI datasets and demonstrates significant improvements in classification accuracy and robustness, highlighting the effectiveness of SD for spectral-spatial learning. Codes are available at https://github.com/Prachet-Dev-Singh/SDHSI.

cs.CV

How (Mis)calibrated is your Federated CLIP and what to do about it?

Vision-language models (VLMs) such as CLIP are increasingly adapted across decentralized data silos, yet the reliability of their predictions under federated learning (FL) remains largely unexplored. In this work, we present a systematic study of calibration in federated CLIP under non-IID client distributions. Our experiments reveal that widely used prompt-tuning methods consistently degrade calibration, often yielding substantially higher calibration error despite competitive recognition performance, while explicit training-time calibration regularizers provide only limited improvements. Motivated by these findings, we identify the choice of fine-tuning parameterization as a critical factor governing calibration and conduct a controlled comparison between prompt tuning and five backbone fine-tuning (BFT) strategies: AdaptFormer, LayerNorm, LoRA, VeRA, and DoRA. We find that BFT methods generally offer a more favorable accuracy-calibration trade-off than prompt tuning, although their benefits are not universal. Through extensive analysis, we show that calibration behavior is closely linked to the geometry of federated updates, residual parameterization, and the resulting client and logit drift. Across in-distribution, domain-generalization, and base-to-new evaluation settings, our results establish fine-tuning parameterization as a central design choice for building accurate and reliable federated CLIP models. Codes are available at https://github.com/mainaksingha01/FL2oRA.

cs.CV

Detecting AI Hallucinations in Finance: An Information-Theoretic Method Cuts Hallucination Rate by 92%

Large language models (LLMs) produce fluent but unsupported answers - hallucinations - limiting safe deployment in high-stakes domains. We propose ECLIPSE, a framework that treats hallucination as a mismatch between a model's semantic entropy and the capacity of available evidence. We combine entropy estimation via multi-sample clustering with a novel perplexity decomposition that measures how models use retrieved evidence. We prove that under mild conditions, the resulting entropy-capacity objective is strictly convex with a unique stable optimum. We evaluate on a controlled financial question answering dataset with GPT-3.5-turbo (n=200 balanced samples with synthetic hallucinations), where ECLIPSE achieves ROC AUC of 0.89 and average precision of 0.90, substantially outperforming a semantic entropy-only baseline (AUC 0.50). A controlled ablation with Claude-3-Haiku, which lacks token-level log probabilities, shows AUC dropping to 0.59 with coefficient magnitudes decreasing by 95% - demonstrating that ECLIPSE is a logprob-native mechanism whose effectiveness depends on calibrated token-level uncertainties. The perplexity decomposition features exhibit the largest learned coefficients, confirming that evidence utilization is central to hallucination detection. We position this work as a controlled mechanism study; broader validation across domains and naturally occurring hallucinations remains future work.

cs.LG

Hidden Order in Trades Predicts the Size of Price Moves

Financial markets exhibit an apparent paradox: while directional price movements remain largely unpredictable--consistent with weak-form efficiency--the magnitude of price changes displays systematic structure. Here we demonstrate that real-time order-flow entropy, computed from a 15-state Markov transition matrix at second resolution, predicts the magnitude of intraday returns without providing directional information. Analysis of 38.5 million SPY trades over 36 trading days reveals that conditioning on entropy below the 5th percentile increases subsequent 5-minute absolute returns by a factor of 2.89 (t = 12.41, p < 0.0001), while directional accuracy remains at 45.0%--statistically indistinguishable from chance (p = 0.12). This decoupling arises from a fundamental symmetry: entropy is invariant under sign permutation, detecting the presence of informed trading without revealing its direction. Walk-forward validation across five non-overlapping test periods confirms out-of-sample predictability, and label-permutation placebo tests yield z = 14.4 against the null. These findings suggest that information-theoretic measures may serve as volatility state variables in market microstructure, though the limited sample (36 days, single instrument) requires extended validation.

q-fin.TR

Discovery of a 13-Sharpe OOS Factor: Drift Regimes Unlock Hidden Cross-Sectional Predictability

We document a high-performing cross-sectional equity factor that achieves out-of-sample Sharpe ratios above 13 through regime-conditional signal activation. The strategy combines value and short-term reversal signals only during stock-specific drift regimes, defined as periods when individual stocks show more than 60 percent positive days in trailing 63-day windows. Under these conditions, the factor delivers annualized returns of 158.6 percent with 12.0 percent volatility and a maximum drawdown of minus 11.9 percent. Using rigorous walk-forward validation across 20 years of S&P 500 data (2004 to 2024), we show performance roughly 13 times stronger than market benchmarks on a risk-adjusted basis, produced entirely out-of-sample with frozen parameters. The factor passes extensive robustness tests, including 1,000 randomization trials with p-values below 0.001, and maintains Sharpe ratios above 7 even under 30 percent parameter perturbations. Exposure to standard risk factors is negligible, with total R-squared values below 3 percent. We provide mechanistic evidence that drift regimes reshape market microstructure by amplifying behavioral biases, altering liquidity patterns, and creating conditions where cross-sectional price discovery becomes systematically exploitable. Conservative capacity estimates indicate deployable capital of 100 to 500 million dollars before noticeable performance degradation.

q-fin.TR

Forecast-to-Fill: Benchmark-Neutral Alpha and Billion-Dollar Capacity in Gold Futures (2015-2025)

We test whether simple, interpretable state variables-trend and momentum-can generate durable out-of-sample alpha in one of the world's most liquid assets, gold. Using a rolling 10-year training and 6-month testing walk-forward from 2015 to 2025 (2,793 trading days), we convert a smoothed trend-momentum regime signal into volatility-targeted, friction-aware positions through fractional, impact-adjusted Kelly sizing and ATR-based exits. Out of sample, the strategy delivers a Sharpe ratio of 2.88 and a maximum drawdown of 0.52 percent, net of 0.7 basis-point linear cost and a square-root impact term (gamma = 0.02). A regression on spot-gold returns yields a 43 percent annualized return (CAGR approximately 43 percent) and a 37 percent alpha (Sharpe = 2.88, IR = 2.09) at a 15 percent volatility target with beta approximately 0.03, confirming benchmark-neutral performance. Bootstrap confidence intervals ([2.49, 3.27]) and SPA tests (p = 0.000) confirm statistical significance and robustness to latency, reversal, and cost stress. We conclude that forecast-to-fill engineering-linking transparent signals to executable trades with explicit risk, cost, and impact control-can transform modest predictability into allocator-grade, billion-dollar-scalable alpha.

q-fin.TR

Learning Under Laws: A Constraint-Projected Neural PDE Solver that Eliminates Hallucinations

Neural networks can approximate solutions to partial differential equations, but they often break the very laws they are meant to model-creating mass from nowhere, drifting shocks, or violating conservation and entropy. We address this by training within the laws of physics rather than beside them. Our framework, called Constraint-Projected Learning (CPL), keeps every update physically admissible by projecting network outputs onto the intersection of constraint sets defined by conservation, Rankine-Hugoniot balance, entropy, and positivity. The projection is differentiable and adds only about 10% computational overhead, making it fully compatible with back-propagation. We further stabilize training with total-variation damping (TVD) to suppress small oscillations and a rollout curriculum that enforces consistency over long prediction horizons. Together, these mechanisms eliminate both hard and soft violations: conservation holds at machine precision, total-variation growth vanishes, and entropy and error remain bounded. On Burgers and Euler systems, CPL produces stable, physically lawful solutions without loss of accuracy. Instead of hoping neural solvers will respect physics, CPL makes that behavior an intrinsic property of the learning process.

cs.LG

Faint active galactic nuclei supplied 31-75% of hydrogen-ionizing photons at z>5

The origin of the ionizing photons that completed hydrogen reionization remains debated. Using recent JWST and ground-based surveys at 4.5 <= z <= 6.5, we construct a unified rest-UV AGN luminosity function that separates unobscured Type I and obscured Type II populations, and show that "little red dots" and X-ray selected sources are magnitude-filtered subsets of Type I with a mixture fraction eta = 0.10 +/- 0.02. We anchor the Lyman-continuum (LyC) escape fraction to outflow incidence and geometric clearing rather than assuming quasar-like values for all classes, and propagate uncertainties through a joint fit. Integrating over -27 < M_UV < -17, AGN inject Ndot_ion,AGN = (3.77 +1.08/-0.95) x 10^51 s^-1 Mpc^-3, nearly twice earlier estimates and comparable to the Ly-alpha inferred requirement at z ~ 6. When combined with the JWST galaxy UV luminosity function and a harder stellar ionizing efficiency of log10(xi_ion) = 25.7, AGN contribute 31-75% of the total ionizing photons for representative galaxy escape fractions f_esc,gal = 0.03-0.20. The resulting hydrogen photoionization rate, Gamma_HI ~ (0.5-2) x 10^-12 s^-1 at z ~ 5-6, lies squarely within the Ly-alpha forest constraints once mean free paths and IGM clumpiness are accounted for, remaining consistent for combined AGN-galaxy models up to f_esc,gal <= 5%. These results suggest that AGN and galaxies jointly sustained the ionizing background during the final stages of reionization, with AGN remaining a major but not exclusive contributor.

astro-ph.GA

A century of change: new changing-look event in Mrk 1018's past

We investigate the long-term variability of the known Changing Look Active Galactic Nuclei (CL AGN) Mrk 1018, whose second change we discovered as part of the Close AGN Reference Survey (CARS). Collating over a hundred years worth of photometry from scanned photographic plates and five modern surveys we find a historic outburst between ~1935-1960, with variation in Johnson B magnitude of ~0.8 that is consistent with Mrk 1018's brightness before and after its latest changing look event in the early 2010s. Using the combined modern and historic data, a Generalised Lomb-Scargle suggests broad feature with P = 29-47 years. Its width and stability across tests, as well as the turn-on speed and bright phase duration of the historic event suggests a timescale associated with long-term modulation, such as via rapid flickering in the accretion rate caused by the Chaotic Cold Accretion model rather than a strictly periodic CL mechanism driving changes in Mrk 1018. We also use the modern photometry to constrain Mrk 1018's latest turn-off duration to less than ~1.9 years, providing further support for a CL mechanism with rapid transition timescales, such as a changing mode of accretion.

astro-ph.GA