SearcharxivSearch

arXiv subjects

Arnab Mondal

Publications and source records attributed to Arnab Mondal.

8 recordsLinked to original sources

On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs

Reinforcement learning (RL) finetuning has become a key technique for enhancing large language models (LLMs) on reasoning-intensive tasks, motivating its extension to vision-language models (VLMs). While RL-tuned VLMs improve on visual reasoning benchmarks, they remain vulnerable to weak visual grounding, hallucinations, and over-reliance on textual cues. We show that simple, controlled textual perturbations, including misleading captions or incorrect chain-of-thought (CoT) traces, cause substantial drops in robustness and confidence, and that these effects are more pronounced when CoT consistency is taken into account across open-source multimodal reasoning models. In contrast, closed models exhibit similar failure modes but maintain markedly greater robustness and reasoning consistency, suggesting that the gap reflects a shortcoming in current open-source RL finetuning rather than an inherent limitation of the task. To better understand these vulnerabilities, we further analyze RL finetuning dynamics and uncover an accuracy-faithfulness trade-off: finetuning raises benchmark accuracy, but can simultaneously erode the reliability of the accompanying CoT and its robustness to contextual shifts. Although adversarial augmentation improves robustness, it does not by itself prevent faithfulness drift. Incorporating a faithfulness-aware reward can restore alignment between answers and reasoning, but when paired with augmentation, training risks collapsing onto shortcut strategies and robustness remains elusive. Together, these findings highlight the limitations of accuracy-only evaluations and motivate training and assessment protocols that jointly emphasize correctness, robustness, and the faithfulness of visually grounded reasoning.

cs.LG

2D GaSe-Based Single-Pixel Spectrometer via Electro-Optical Barrier Co-Modulation

Driven by the growing demand for miniaturized spectrometers for in-situ analysis, and point-of-care diagnostics, conventional spectrometers are often constrained by bulky architectures and pathlength-limited spectral resolution. Achieving high-resolution, single-pixel computational spectrometers is therefore critical for the realization of compact, on-chip systems. Here, we report a single-pixel spectrometer enabled by a single 2D material; few-layer GaSe-based photodetector, in which the Schottky barrier height modulation, governed jointly by applied bias and optical excitation, provides an efficient mechanism for spectral encoding without the need for bulky dispersive elements. The device exhibits a high peak-wavelength accuracy of ~0.78 nm across a broad operational bandwidth (300-700 nm) within a compact footprint of ~100 um^2 and resolves closely spaced spectral features with separations down to ~5 nm. The device operates at low bias (+/- 4V) with an ultralow dark current density ~0.3 pA/um^2 at 4V bias. These results establish a simple, scalable route toward compact, cost-effective spectroscopic systems for on-chip spectral sensing and portable hyperspectral imaging applications.

physics.optics

Rendering-Aware Reinforcement Learning for Vector Graphics Generation

Scalable Vector Graphics (SVG) offer a powerful format for representing visual designs as interpretable code. Recent advances in vision-language models (VLMs) have enabled high-quality SVG generation by framing the problem as a code generation task and leveraging large-scale pretraining. VLMs are particularly suitable for this task as they capture both global semantics and fine-grained visual patterns, while transferring knowledge across vision, natural language, and code domains. However, existing VLM approaches often struggle to produce faithful and efficient SVGs because they never observe the rendered images during training. Although differentiable rendering for autoregressive SVG code generation remains unavailable, rendered outputs can still be compared to original inputs, enabling evaluative feedback suitable for reinforcement learning (RL). We introduce RLRF (Reinforcement Learning from Rendering Feedback), an RL method that enhances SVG generation in autoregressive VLMs by leveraging feedback from rendered SVG outputs. Given an input image, the model generates SVG roll-outs that are rendered and compared to the original image to compute a reward. This visual fidelity feedback guides the model toward producing more accurate, efficient, and semantically coherent SVGs. RLRF significantly outperforms supervised fine-tuning, addressing common failure modes and enabling precise, high-quality SVG generation with strong structural understanding and generalization.

cs.CV

The generation of diverse traveling pulses and its solution scheme in an excitable slow-fast dynamics

In this paper, we report on the generation and propagation of traveling pulses in a homogeneous network of diffusively coupled, excitable, slow-fast dynamical neurons. The spatially extended system is modelled using the nearest neighbor coupling theory, in which the diffusion part measures the spatial distribution of the coupling topology. We derive analytically the conditions for traveling wave profiles that allow the construction of the shape of traveling nerve impulses. The analytical and numerical results are used to explore the nature of the propagating pulses. The symmetric or asymmetric nature of the traveling pulses is characterized and the wave velocity is derived as a function of system parameters. Moreover, we present our results for an extended excitable medium by considering a slow-fast biophysical model with a homogeneous, diffusive coupling that can exhibit various traveling pulses. The appearance of series of pulses is an interesting phenomenon from biophysical and dynamical perspective. Varying the perturbation and coupling parameters, we observe the propagation of activities with various amplitude modulations and transition phases of different wave profiles that affect the speed of the pulses in certain parameter regimes. We observe different types of traveling pulses, such as envelope solitons and multi-bump solutions and show how system parameters and the coupling play a major role in the formation of different traveling pulses. Finally, we obtain the conditions for stable and unstable plane waves.

nlin.PS

Non trivial dynamics in the FizHugh-Rinzel model and non-homogeneous oscillatory-excitable reaction-diffusions systems

In this article, we discuss the dynamics of the 3-dimensional FitzHugh-Rinzel (FHR) model and a class of non-homogeneous FitzHugh-Nagumo (Nh-FHN) Reaction-Diffusion systems. The Nh-FHN models can be used to generate relevant wave propagation phenomena in Neuroscience context. This gives raise locally to complex dynamics such as canards, Mixed Mode Oscillations, Hopf-Bifurcations some of which can be observed in the FHR model.

nlin.PS

Spatiotemporal instabilities and pattern formation in systems of diffusively coupled Izhikevich neurons

Neurons are often connected, spatially and temporally, in phenomenal ways that promote wave propagation. Therefore, it is essential to analyze the emergent spatiotemporal patterns to understand the working mechanism of brain activity, especially in cortical areas. Here, we present an explicit mathematical analysis, corroborated by numerical results, to identify and investigate the spatiotemporal, non-uniform, patterns that emerge due to instability in an extended homogeneous 2D spatial domain, using the excitable Izhikevich neuron model. We examine diffusive instability and perform bifurcation and fixed-point analyses to characterize the patterns and their stability. Then, we derive analytically the amplitude equations that establish the activities of reaction-diffusion structures. We report on the emergence of diverse spatial structures including hexagonal and mixed-type patterns by providing a systematic mathematical approach, including variations in correlated oscillations, pattern variations and amplitude fluctuations. Our work shows that the emergence of spatiotemporal behavior, commonly found in excitable systems, has the potential to contribute significantly to the study of diffusively-coupled biophysical systems at large.

nlin.PS

Retinal Vessel Segmentation under Extreme Low Annotation: A Generative Adversarial Network Approach

Contemporary deep learning based medical image segmentation algorithms require hours of annotation labor by domain experts. These data hungry deep models perform sub-optimally in the presence of limited amount of labeled data. In this paper, we present a data efficient learning framework using the recent concept of Generative Adversarial Networks; this allows a deep neural network to perform significantly better than its fully supervised counterpart in low annotation regime. The proposed method is an extension of our previous work with the addition of a new unsupervised adversarial loss and a structured prediction based architecture. To the best of our knowledge, this work is the first demonstration of an adversarial framework based structured prediction model for medical image segmentation. Though generic, we apply our method for segmentation of blood vessels in retinal fundus images. We experiment with extreme low annotation budget (0.8 - 1.6% of contemporary annotation size). On DRIVE and STARE datasets, the proposed method outperforms our previous method and other fully supervised benchmark models by significant margins especially with very low number of annotated examples. In addition, our systematic ablation studies suggest some key recipes for successfully training GAN based semi-supervised algorithms with an encoder-decoder style network architecture.

cs.CV

Low Cost Autonomous Navigation and Control of a Mechanically Balanced Bicycle with Dual Locomotion Mode

On the lines of the huge and varied efforts in the field of automation with respect to technology development and innovation of vehicles to make them run autonomously, this paper presents an innovation to a bicycle. A normal daily use bicycle was modified at low cost such that it runs autonomously, while maintaining its original form i.e. the manual drive. Hence, a bicycle which could be normally driven by any human and with a press of switch could run autonomously according to the needs of the user has been developed.

cs.RO