SearcharxivSearch

arXiv subjects

Samiran Das

Publications and source records attributed to Samiran Das.

8 recordsLinked to original sources

Multi-Scale Fruit Capsules: Dilated Convolutions and Dynamic Routing for In-the-Wild Explainable Fruit Recognition

The same fruit appears in a bunch, unpicked, peeled, bagged in plastic, or sliced on a dish, so automated fruit classification in the wild (AFCW) must absorb wide intra- class and narrow inter-class variability in shape, size, colour and texture. Convolutional networks route information through pooling, which discards the pose and location of the region of interest and therefore generalises poorly across these presentations. We propose FruitCapsNet, a capsule network whose Fruit Capsules replace the standard convolutional front end with dilated convolutions: the receptive field grows exponentially at constant parameter cost, so each capsule encodes multi-scale context before dynamic routing resolves part whole spatial agreement. Hyper-parameters, including the dilation factor, are selected by Bayesian optimisation rather than grid search. On three public datasets (SMP, FruitsGB, Fruits-360) and a new 19-class, 10,639-image in-the-wild dataset (PD-19), FruitCapsNet exceeds ten fine-tuned transfer-learning backbones at one-third the depth, with the largest margin (+2.7% over the nearest competitor) on the hardest set. Grad-CAM saliency propagated from the DigitCaps layer shows that the improvement comes from attributing decisions to whole-fruit regions rather than to object edges, giving post-hoc evidence that the gain is not a dataset artefact.

cs.CV

Histogram Assisted Quality Aware Generative Model for Resolution Invariant NIR Image Colorization

We present HAQAGen, a unified generative model for resolution-invariant NIR-to-RGB colorization that balances chromatic realism with structural fidelity. The proposed model introduces (i) a combined loss term aligning the global color statistics through differentiable histogram matching, perceptual image quality measure, and feature based similarity to preserve texture information, (ii) local hue-saturation priors injected via Spatially Adaptive Denormalization (SPADE) to stabilize chromatic reconstruction, and (iii) texture-aware supervision within a Mamba backbone to preserve fine details. We introduce an adaptive-resolution inference engine that further enables high-resolution translation without sacrificing quality. Our proposed NIR-to-RGB translation model simultaneously enforces global color statistics and local chromatic consistency, while scaling to native resolutions without compromising texture fidelity or generalization. Extensive evaluations on FANVID, OMSIV, VCIP2020, and RGB2NIR using different evaluation metrics demonstrate consistent improvements over state-of-the-art baseline methods. HAQAGen produces images with sharper textures, natural colors, attaining significant gains as per perceptual metrics. These results position HAQAGen as a scalable and effective solution for NIR-to-RGB translation across diverse imaging scenarios. Project Page: https://rajeev-dw9.github.io/HAQAGen/

cs.CV

On VLMs for Diverse Tasks in Multimodal Meme Classification

In this paper, we present a comprehensive and systematic analysis of vision-language models (VLMs) for disparate meme classification tasks. We introduced a novel approach that generates a VLM-based understanding of meme images and fine-tunes the LLMs on textual understanding of the embedded meme text for improving the performance. Our contributions are threefold: (1) Benchmarking VLMs with diverse prompting strategies purposely to each sub-task; (2) Evaluating LoRA fine-tuning across all VLM components to assess performance gains; and (3) Proposing a novel approach where detailed meme interpretations generated by VLMs are used to train smaller language models (LLMs), significantly improving classification. The strategy of combining VLMs with LLMs improved the baseline performance by 8.34%, 3.52% and 26.24% for sarcasm, offensive and sentiment classification, respectively. Our results reveal the strengths and limitations of VLMs and present a novel strategy for meme understanding.

cs.CL

MaskCD: A Remote Sensing Change Detection Network Based on Mask Classification

Change detection (CD) from remote sensing (RS) images using deep learning has been widely investigated in the literature. It is typically regarded as a pixel-wise labeling task that aims to classify each pixel as changed or unchanged. Although per-pixel classification networks in encoder-decoder structures have shown dominance, they still suffer from imprecise boundaries and incomplete object delineation at various scenes. For high-resolution RS images, partly or totally changed objects are more worthy of attention rather than a single pixel. Therefore, we revisit the CD task from the mask prediction and classification perspective and propose MaskCD to detect changed areas by adaptively generating categorized masks from input image pairs. Specifically, it utilizes a cross-level change representation perceiver (CLCRP) to learn multiscale change-aware representations and capture spatiotemporal relations from encoded features by exploiting deformable multihead self-attention (DeformMHSA). Subsequently, a masked-attention-based detection transformers (MA-DETR) decoder is developed to accurately locate and identify changed objects based on masked attention and self-attention mechanisms. It reconstructs the desired changed objects by decoding the pixel-wise representations into learnable mask proposals and making final predictions from these candidates. Experimental results on five benchmark datasets demonstrate the proposed approach outperforms other state-of-the-art models. Codes and pretrained models are available online (https://github.com/EricYu97/MaskCD).

cs.CV

On strong $\mathcal{A}^{\mathcal{I}}$-statistical convergence of sequences in probabilistic metric spaces

In this paper using a non-negative regular summability matrix $\mathcal{A}$ and a non-trivial admissible ideal $\mathcal{I}$ in $\mathbb{N}$ we study some basic properties of strong $\mathcal{A}^{\mathcal{I}}$-statistical convergence and strong $\mathcal{A}^{\mathcal{I}}$-statistical Cauchyness of sequences in probabilistic metric spaces not done earlier. We also introduce strong $\mathcal{A}^{\mathcal{I^*}}$-statistical Cauchyness in probabilistic metric space and study its relationship with strong A$\mathcal{A}^{\mathcal{I}}$-statistical Cauchyness there. Further, we study some basic properties of strong $\mathcal{A}^{\mathcal{I}}$-statistical limit points and strong $\mathcal{A}^{\mathcal{I}}$-statistical cluster points of a sequence in probabilistic metric spaces.

math.FA

On Strong A-statistical Convergence in Probabilistic Metric Spaces

In this paper we study some basic properties of strong A-statistical convergence and strong A-statistical Cauchyness of sequences in probabilistic metric spaces not done earlier. We also study some basic properties of strong A-statistical limit points and strong A-statistical cluster points of a sequence in a probabilistic metric space. Further we also introduce the notion of strong statistically A-summable sequence in a probabilistic metric space and study its relationship with strong A-statistical convergence.

math.FA

On strong {\lambda}-statistical convergence of sequences in probabilistic metric (pm) spaces

In this paper we study some basic properties of strong {\lambda}- statistical convergence of sequences in probabilistic metric (PM) spaces. We also introduce and study the notion of strong {\lambda}-statistically Cauchyness. Further introducing the notions of strong {\lambda}-statistical limit point and strong {\lambda}-statistical cluster point of a sequence in a probabilistic metric (PM) space we examine their interrelationship.

math.FA

Hyperspectral Unmixing by Nuclear Norm Difference Maximization based Dictionary Pruning

Dictionary pruning methods perform unmixing by identifying a smaller subset of active spectral library elements that can represent the image efficiently as a linear combination. This paper presents a new nuclear norm difference based approach for dictionary pruning utilizing the low rank property of hyperspectral data. The proposed workflow calculates the nuclear norm of abundance of the original data assuming the whole spectral library as endmembers. In the next step, the algorithm calculates nuclear norm of abundance after appending a spectral library element with the data. The spectral library elements having the maximum difference in the nuclear norm of the obtained abundance matrices are suitable candidates for being image endmember. The proposed workflow is verified with a large number of synthetic data generated by varying condition as well as some real images.

eess.IV