SearcharxivSearch

arXiv subjects

Tahmid Rahman

Publications and source records attributed to Tahmid Rahman.

2 recordsLinked to original sources

A Heterogeneous Two-Stream Framework for Video Action Recognition with Comparative Fusion Analysis

Most two-stream action recognition networks apply the same convolutional backbone to both RGB and optical flow streams, ignoring the fact that the two modalities have fundamentally different structural properties. Optical flow captures fine-grained motion patterns, while RGB frames carry rich appearance and scene context - treating them identically discards this distinction. We propose DualStreamHybrid, a heterogeneous two-stream architecture that assigns each stream a backbone suited to its input: a pretrained ViT-Tiny/16 for RGB frames, and a MobileNetV2 trained from scratch on a 20-channel stacked optical flow representation. A learned projection layer maps the two differently-sized feature vectors to a common dimensionality before fusion, enabling the two streams to interact without forcing architectural symmetry. We design five fusion strategies within a unified framework - late fusion, concatenation, cross-attention, weighted fusion, and gated fusion - and evaluate them on UCF11 (1,600 videos, 11 classes) and UCF50 (6,681 videos, 50 classes) to study how fusion behaviour scales with dataset size. On UCF11, cross-attention achieves 98.12% test accuracy, outperforming the RGB-only ViT-Tiny baseline of 95.94%, which suggests that explicit inter-modal attention is particularly effective on smaller, less complex datasets. On UCF50, weighted fusion reaches 96.86% and proves the most consistent strategy across both benchmarks. The learned stream weights reveal an interesting pattern: UCF11 sees near-equal modality contribution (RGB: 0.507, flow: 0.493), while UCF50 favours the RGB stream slightly more (RGB: 0.554, flow: 0.446) - arguably reflecting the larger and more visually diverse action space. Taken together, these results suggest that even a lightweight motion stream meaningfully complements a strong appearance encoder, and that the optimal fusion strategy depends on dataset scale.

cs.CV

Wavelength-Accurate Nonlinear Conversion through Wavenumber Selectivity in Photonic Crystal Resonators

Integrated nonlinear wavelength converters transfer optical energy from lasers or quantum emitters to other useful colors, but chromatic dispersion limits the range of achievable wavelength shifts. Moreover, because of geometric dispersion, fabrication tolerances reduce the accuracy with which devices produce specific target wavelengths. Here, we report nonlinear wavelength converters whose operation is not contingent on dispersion engineering; yet, the output wavelengths are controlled with high accuracy. In our scheme, coherent coupling between counter-propagating waves in a photonic crystal microresonator induces a photonic bandgap that isolates (in dispersion space) specific wavenumbers for nonlinear gain. We first demonstrate the wide applicability of this strategy to parametric nonlinear processes, by simulating its use in third harmonic generation, dispersive wave formation in Kerr microcombs, and four-wave mixing Bragg scattering. In experiments, we demonstrate Kerr optical parametric oscillators in which such wavenumber-selective coherent coupling designates the signal mode. As a result, differences between the targeted and realized signal wavelengths are <0.3 percent. Moreover, leveraging the bandgap-protected wavenumber selectivity, we continuously tune the output frequencies by nearly 300 GHz without compromising efficiency. Our results will bring about a paradigm shift in how microresonators are designed for nonlinear optics, and they make headway on the larger problem of building wavelength-accurate light sources using integrated photonics.

physics.optics