SearcharxivSearch

arXiv subjects

Aditya Chaudhary

Publications and source records attributed to Aditya Chaudhary.

3 recordsLinked to original sources

MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP

In this paper, we propose a novel multimodal framework, Multimodal Language-Guided Network (MMLGNet), to align heterogeneous remote sensing modalities like Hyperspectral Imaging (HSI) and LiDAR with natural language semantics using vision-language models such as CLIP. With the increasing availability of multimodal Earth observation data, there is a growing need for methods that effectively fuse spectral, spatial, and geometric information while enabling semantic-level understanding. MMLGNet employs modality-specific encoders and aligns visual features with handcrafted textual embeddings in a shared latent space via bi-directional contrastive learning. Inspired by CLIP's training paradigm, our approach bridges the gap between high-dimensional remote sensing data and language-guided interpretation. Notably, MMLGNet achieves strong performance with simple CNN-based encoders, outperforming several established multimodal visual-only methods on two benchmark datasets, demonstrating the significant benefit of language supervision. Codes are available at https://github.com/AdityaChaudhary2913/CLIP_HSI.

cs.CV

Two-Stage Vision Transformer for Image Restoration: Colorization Pretraining + Residual Upsampling

In computer vision, Single Image Super-Resolution (SISR) is still a difficult problem. We present ViT-SR, a new technique to improve the performance of a Vision Transformer (ViT) employing a two-stage training strategy. In our method, the model learns rich, generalizable visual representations from the data itself through a self-supervised pretraining phase on a colourization task. The pre-trained model is then adjusted for 4x super-resolution. By predicting the addition of a high-frequency residual image to an initial bicubic interpolation, this design simplifies residual learning. ViT-SR, trained and evaluated on the DIV2K benchmark dataset, achieves an impressive SSIM of 0.712 and PSNR of 22.90 dB. These results demonstrate the efficacy of our two-stage approach and highlight the potential of self-supervised pre-training for complex image restoration tasks. Further improvements may be possible with larger ViT architectures or alternative pretext tasks.

cs.CV

Solar Cells, Lambert W and the LogWright Functions

Algorithms that calculate the current-voltage (I-V) characteristics of a solar cell play an important role in processes that aim to improve the efficiency of a solar cell. I-V characteristics can be obtained from different models used to represent the solar cell, and the single diode model is a simple yet accurate model for common field implementations. However, the I-V characteristics are obtained by solving implicit equations, which involve repeated iterations and inherent errors associated with numerical methods used. Some methods use the Lambert W function to get an exact explicit formula, but often causes numerical overflow problems. The present work discusses an algorithm to calculate I-V characteristics using the LogWright function, a transformation of the Lambert W function, which addresses the problem of arithmetic overflow that occurs in the Lambert W implementation. An implementation of this algorithm is presented and compared against other algorithms in the literature. It is observed that in addition to addressing the numerical overflow problem, the algorithm based on the LogWright function offers speed benefits while retaining high precision.

physics.comp-ph