SearcharxivSearch

arXiv subjects

Jiaqi Song

Publications and source records attributed to Jiaqi Song.

16 recordsLinked to original sources

Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR

Modern LLM-based ASR systems have established multilingual capability as a standard feature, leveraging large-scale multilingual corpora and LLMs' cross-lingual knowledge to achieve competitive performance across multilingual benchmarks. However, jointly modeling languages with heterogeneous acoustic, phonological, and lexical characteristics inevitably introduces optimization conflicts, undermining language-wise specialization. To address this challenge, we propose Language-Specialized Multi-Teacher On-Policy Distillation (LS-MOPD), which decouples language-specific knowledge acquisition from multilingual capability integration: language-specialized teachers are independently optimized via reinforcement learning (RL), with their expertise then integrated into a generalist multilingual student through language routing and token-level multi-teacher distillation, thereby reducing direct cross-lingual optimization conflicts. We further explore static and dynamic acoustic-prefix configurations to examine how teacher-student prefix consistency influences the efficacy of on-policy distillation. Experiments on benchmarks covering Mandarin, Mandarin subdialects, Cantonese, and English demonstrate that LS-MOPD substantially outperforms RL baselines and surpasses the empirical performance envelope defined by the best-performing RL teachers on nearly all benchmarks, revealing its potential to generalize beyond all teachers in multilingual ASR.

cs.CL

InsightVQA: High-Dimensional Emotion-Cognitive Visual Question Answering Benchmark

Visual emotion understanding requires models not only to recognize emotional states, but also to why they arise and perform higher-level cognitive reasoning. However, existing benchmarks mainly focus on emotion recognition, offering limited support for grounded understanding and response-oriented analysis. To address this gap, we introduce \textbf{InsightVQA}, a large-scale dataset for hierarchical visual question answering on emotion understanding and cognitive reasoning. Building from 351K images collected from six public sources, we apply a rigorous multi-stage filtering pipeline to curate 138K high-confidence images. Each image is annotated at three hierarchical levels: perception QA for emotion and valence recognition, grounded understanding QA constructed from visual trigger extraction through constraint-guided generation, and cognition QA centered on response intent prediction and sequential insight reasoning. In total, InsightVQA contains 725K QA pairs. We further present \textbf{InsightVQA-Bench}, a high-quality evaluation benchmark comprising 30K samples for fine-grained evaluation. To support evaluation, we introduce \textbf{InsightNet}, an emotion-tuned baseline for MLLMs. Results demonstrate that InsightVQA poses significant challenges for grounded emotion understanding and reasoning.

cs.CV

RQ-MoE: Residual Quantization via Mixture of Experts for Efficient Input-Dependent Vector Compression

Vector quantization is a fundamental tool for compressing high-dimensional embeddings, yet existing multi-codebook methods rely on static codebooks that limit expressiveness under heterogeneous data geometry. While recent dynamic quantizers like QINCo adapt codebooks to individual inputs and improve expressiveness, their strict sequential dependencies create decoding bottlenecks. We propose Residual Quantization via Mixture of Experts (RQ-MoE), a framework combining a two-level MoE with dual-stream quantization to enable input-dependent codebook adaptation for efficient vector quantization. RQ-MoE enables dynamic codebook construction and decouples instruction from quantization, facilitating parallel decoding. Theoretically, we show that standard Residual Quantization and QINCo can be recovered as constrained special cases of RQ-MoE, and derive a guideline for setting expert dimensionality in RQ-MoE. Extensive experiments show that RQ-MoE achieves state-of-the-art or on-par performance in reconstruction and retrieval, while providing 6x-14x faster decoding than prior vector quantization methods. The implementation is available at https://github.com/KDEGroup/RQ-MoE.

cs.LG

NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR

Integrating large language models (LLMs) into automatic speech recognition (ASR) has become a mainstream paradigm in recent years. Although existing LLM-based ASR models demonstrate impressive performance on public benchmarks, their training remains predominantly data-driven, leaving key practical challenges insufficiently addressed -- particularly limited downward scalability in resource-constrained deployments and hallucinations under acoustically challenging conditions. To address these issues, we present NIM4-ASR, a production-oriented LLM-based ASR framework optimized for both efficiency and robustness. Grounded in a principled delineation of functional roles between the encoder and the LLM, we redesign the multi-stage training paradigm to align each module with its intended capability boundary. Specifically, we reformulate the pre-training architecture and objective to mitigate the modality gap and improve parameter efficiency; introduce an iterative asynchronous SFT stage to preserve acoustic fidelity and constrain representation drift; and design an ASR-specialized reinforcement learning stage to further enhance recognition quality and robustness. We additionally incorporate a suite of production-oriented optimizations, including robustness under noisy and silent conditions, real-time streaming inference, and hotword customization via retrieval-augmented generation (RAG). Experiments show that NIM4-ASR achieves state-of-the-art performance on multiple public benchmarks with merely 2.3B parameters, while substantially outperforming larger-scale competitors on internal benchmarks -- particularly in entity-intensive real-world scenarios. NIM4-ASR further supports million-scale hotword customization via RAG with sub-millisecond retrieval latency, enabling efficient adaptation to emerging entities and personalized user requirements.

eess.AS

ARGen: Affect-Reinforced Generative Augmentation towards Vision-based Dynamic Emotion Perception

Dynamic facial expression recognition in the wild remains challenging due to data scarcity and long-tail distributions, which hinder models from effectively learning the temporal dynamics of scarce emotions. To address these limitations, we propose ARGen, an Affect-Reinforced Generative Augmentation Framework that enables data-adaptive dynamic expression generation for robust emotion perception. ARGen operates in two stages: Affective Semantic Injection (ASI) and Adaptive Reinforcement Diffusion (ARD). The ASI stage establishes affective knowledge alignment through facial Action Units and employs a retrieval-augmented prompt generation strategy to synthesize consistent and fine-grained affective descriptions via large-scale visual-language models, thereby injecting interpretable emotional priors into the generation process. The ARD stage integrates text-conditioned image-to-video diffusion with reinforcement learning, introducing inter-frame conditional guidance and a multi-objective reward function to jointly optimize expression naturalness, facial integrity, and generative efficiency. Extensive experiments on both generation and recognition tasks verify that ARGen substantially enhances synthesis fidelity and improves recognition performance, establishing an interpretable and generalizable generative augmentation paradigm for vision-based affective computing.

cs.CV

Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs

Integrating large language models (LLMs) into automatic speech recognition (ASR) has become a dominant paradigm. Although recent LLM-based ASR models have shown promising performance on public benchmarks, it remains challenging to balance recognition quality with latency and overhead, while hallucinations further limit real-world deployment. In this study, we revisit LLM-based ASR from an entropy allocation perspective and introduce three metrics to characterize how training paradigms allocate entropy reduction between the speech encoder and the LLM. To remedy entropy-allocation inefficiencies in prevailing approaches, we propose a principled multi-stage training strategy grounded in capability-boundary awareness, optimizing parameter efficiency and hallucination robustness. Specifically, we redesign the pretraining strategy to alleviate the speech-text modality gap, and further introduce an iterative asynchronous SFT stage between alignment and joint SFT to preserve functional decoupling and constrain encoder representation drift. Experiments on Mandarin and English benchmarks show that our method achieves competitive performance with state-of-the-art models using only 2.3B parameters, while also effectively mitigating hallucinations through our decoupling-oriented design.

eess.AS

Compressive hyperspectral phasor imaging with single-pixel detection for spectral tasks

Spectral vision task plays a pivotal role in extracting discriminative spectral-spatial features from high-dimensional data, enabling fine-grained identification beyond human vision. Traditional methods usually involve first collecting rich spectral-spatial information and then using complex algorithms to digitally process it into scene classification and recognition. However, the complexity of processing massive three-dimensional (3D) hyperspectral datasets poses challenges for algorithms. Here, we demonstrate a compressive Hyperspectral Phasor Imaging with Single-pixel detection (HyPIS) that leverages highly compressed spatial-spectral data to achieve spectral task. Two optical encoders are used for wavelength-dependent sine- and cosine-encoding that transforms spectral signals into a two-dimensional (2D) phasor plot. By applying spatial-temporal illumination patterns, a single-pixel detector is enough to reconstruct the phasor image of the object. This allows to directly generate pixel-wise spectral task, bypassing 3D hyperspectral data. Our experiments show that HyPIS can perform real-time classification and recognition tasks of different scenes, reducing the required amount of data by two orders of magnitude, and it can still accurately classify under low light and uneven lighting conditions. This work develops a completely new spectral technology that enables spectral tasks to be performed without obtaining high-resolution hyperspectral datasets, holding promise for spectral applications in mobile devices, robotics, and satellite technologies.

physics.optics

Physics-informed neural network enhanced multispectral single-pixel imaging with a chip spectral sensor

Multispectral imaging (MSI) captures data across multiple spectral bands, offering enhanced informational depth compared to standard RGB imaging and benefiting diverse fields such as agriculture, medical diagnostics, and industrial inspection. Conventional MSI systems, however, suffer from high cost, complexity, and limited performance in low-light conditions. Moreover, data-driven MSI methods depend heavily on large, labeled training datasets and struggle with generalization. In this work, we present a portable multispectral single-pixel imaging (MS-SPI) method that integrates a chip-sized multispectral sensor for system miniaturization and leverages an untrained physics-informed neural network (PINN) to reconstruct high-quality spectral images without the need for labeled training data. The physics-informed structure of the network enables the self-corrected reconstruction of multispectral images directly with the input of raw measurements from the multispectral sensor. Our proof-of-concept prototype achieves the reconstruction of 12-channel high-quality spectral images at the sampling rate of 10%. We also experimentally validate its performance under varying sampling rate conditions, by comparing it with conventional compressive sensing algorithms. Furthermore, we demonstrate the application of this technique to an MSI-based image segmentation task, in which spatial regions are discriminated according to their characteristic spectral signatures. This compact, high-fidelity, and portable approach offers promising pathways to lightweight and cost-effective spectral imaging on mobile platforms.

physics.ins-det

Exploiting scattering-based point spread functions for snapshot 5D and modality-switchable lensless imaging

Snapshot multi-dimensional imaging offers a promising alternative to traditional low-dimensional imaging techniques by enabling the simultaneous capture of spatial, spectral, polarization, and other information in a single shot for improved imaging speed and acquisition efficiency. However, existing snapshot multi-dimensional imaging systems are often hindered by their large size, complexity, and high cost, which constrain their practical applicability. In this work, we propose a compact lensless diffuser camera for snapshot multi-dimensional imaging (Diffuser-mCam), which can reconstruct five-dimensional (5-D) images from a single-shot 2D recording of speckle-like measurement under incoherent illumination. By employing both the scattering medium and the space-division multiplexing strategy to extract high-dimensional optical features, we show that the multi-dimensional data (2D intensity distribution, spectral, polarization, time) of the desired light field can be encoded into a snapshot speckle-like pattern via a diffuser, and subsequently decoded using a compressed sensing algorithm at the sampling rate of 2.5%, eliminating the need for multi-scanning processes. We further demonstrate that our method can be flexibly switched between 5D and selectively reduced-dimensional imaging, providing an efficient way of reducing computational resource demands. Our work presents a compact, cost-effective, and versatile framework for snapshot multi-dimensional imaging and opens up new opportunities for the design of novel imaging systems for applications in areas such as medical imaging, remote sensing, and autonomous systems.

physics.optics

FastAdaSP: Multitask-Adapted Efficient Inference for Large Speech Language Model

In this study, we aim to explore Multitask Speech Language Model (SpeechLM) efficient inference via token reduction. Unlike other modalities such as vision or text, speech has unique temporal dependencies, making previous efficient inference works on other modalities not directly applicable. Furthermore, methods for efficient SpeechLM inference on long sequence and sparse signals remain largely unexplored. Then we propose FastAdaSP, a weighted token merging framework specifically designed for various speech-related tasks to improve the trade-off between efficiency and performance. Experimental results on WavLLM and Qwen-Audio show that our method achieves the state-of-the-art (SOTA) efficiency-performance trade-off compared with other baseline methods. Specifically, FastAdaSP achieved 7x memory efficiency and 1.83x decoding throughput without any degradation on tasks like Emotion Recognition (ER) and Spoken Question Answering (SQA). The code will be available at https://github.com/yichen14/FastAdaSP

eess.AS

SynesLM: A Unified Approach for Audio-visual Speech Recognition and Translation via Language Model and Synthetic Data

In this work, we present SynesLM, an unified model which can perform three multimodal language understanding tasks: audio-visual automatic speech recognition(AV-ASR) and visual-aided speech/machine translation(VST/VMT). Unlike previous research that focused on lip motion as visual cues for speech signals, our work explores more general visual information within entire frames, such as objects and actions. Additionally, we use synthetic image data to enhance the correlation between image and speech data. We benchmark SynesLM against the How2 dataset, demonstrating performance on par with state-of-the-art (SOTA) models dedicated to AV-ASR while maintaining our multitasking framework. Remarkably, for zero-shot AV-ASR, SynesLM achieved SOTA performance by lowering the Word Error Rate (WER) from 43.4% to 39.4% on the VisSpeech Dataset. Furthermore, our results in VST and VMT outperform the previous results, improving the BLEU score to 43.5 from 37.2 for VST, and to 54.8 from 54.4 for VMT.

eess.AS

Multi-photon super-linear image scanning microscopy using upconversion nanoparticles

Super-resolution fluorescence microscopy is of great interest in life science studies for visualizing subcellular structures at the nanometer scale. Among various kinds of super-resolution approaches, image scanning microscopy (ISM) offers a doubled resolution enhancement in a simple and straightforward manner, based on the commonly used confocal microscopes. ISM is also suitable to be integrated with multi-photon microscopy techniques, such as two-photon excitation and second-harmonic generation imaging, for deep tissue imaging, but it remains the twofold limited resolution enhancement and requires expensive femtosecond lasers. Here, we present and experimentally demonstrate the super-linear ISM (SL-ISM) to push the resolution enhancement beyond the factor of two, with a single low-power, continuous-wave, and near-infrared laser, by harnessing the emission nonlinearity within the multiphoton excitation process of lanthanide-doped upconversion nanoparticles (UCNPs). Based on a modified confocal microscope, we achieve a resolution of about 120 nm, 1/8th of the excitation wavelength. Furthermore, we demonstrate a parallel detection strategy of SL-ISM with the multifocal structured excitation pattern, to speed up the acquisition frame rate. This method suggests a new perspective for super-resolution imaging or sensing, multi-photon imaging, and deep-tissue imaging with simple, low-cost, and straightforward implementations.

physics.optics

Miniaturized on-chip spectrometer enabled by electrochromic modulation

Miniaturized on-chip spectrometers with small footprints, lightweight, and low cost are in great demand for portable optical sensing, lab-on-chip systems, and so on. Such miniaturized spectrometers are usually based on engineered spectral response units and then reconstruct unknown spectra with algorithms. However, due to the limited footprints of computational on-chip spectrometers, the recovered spectral resolution is limited by the number of integrated spectral response units/filters. Thus, it is challenging to improve the spectral resolution without increasing the number of used filters. Here we present a computational on-chip spectrometer using electrochromic filters that can be electrochemically modulated to increase the efficient sampling number for higher spectral resolution. These filters are directly integrated on top of the photodetector pixels, and the spectral modulation of the filters results from redox reactions during the dual injection of ions and electrons into the electrochromic material. We experimentally demonstrate that the spectral resolution of the proposed spectrometer can be effectively improved as the number of applied voltages increases. The average difference of the peak wavelengths between the reconstructed and the reference spectra decreases from 14.48 nm to 2.57 nm. We also demonstrate the proposed spectrometer can be worked with only four or two filter units, assisted by electrochromic modulation. This strategy suggests a new way to enhance the performance of miniaturized spectrometers with tunable spectral filters for high resolution, low-cost, and portable spectral sensing, and would also inspire the exploration of other stimulus responses such as photochromic and force-chromic, etc, on computational spectrometers.

physics.optics

Temporal compressive edge imaging enabled by a lensless diffuser camera

Lensless imagers based on diffusers or encoding masks enable high-dimensional imaging from a single shot measurement and have been applied in various applications. However, to further extract image information such as edge detection, conventional post-processing filtering operations are needed after the reconstruction of the original object images in the diffuser imaging systems. Here, we present the concept of a temporal compressive edge detection method based on a lensless diffuser camera, which can directly recover a time sequence of edge images of a moving object from a single-shot measurement, without further post-processing steps. Our approach provides higher image quality during edge detection, compared with the conventional post-processing method. We demonstrate the effectiveness of this approach by both numerical simulation and experiments. The proof-of-concept approach can be further developed with other image post-process operations or versatile computer vision assignments toward task-oriented intelligent lensless imaging systems.

eess.IV

Quantitative and dark field ghost imaging with ultraviolet light

Ultraviolet (UV) imaging enables a diverse array of applications, such as material composition analysis, biological fluorescence imaging, and detecting defects in semiconductor manufacturing. However, scientific-grade UV cameras with high quantum efficiency are expensive and include a complex thermoelectric cooling system. Here, we demonstrate a UV computational ghost imaging (UV-CGI) method to provide a cost-effective UV imaging and detection strategy. By applying spatial-temporal illumination patterns and using a 325 nm laser source, a single-pixel detector is enough to reconstruct the images of objects. To demonstrate its capability for quantitative detection, we use UV-CGI to distinguish four UV-sensitive sunscreen areas with different densities on a sample. Furthermore, we demonstrate dark field UV-CGI in both transmission and reflection schemes. By only collecting the scattered light from objects, we can detect the edges of pure phase objects and small scratches on a compact disc. Our results showcase a feasible low-cost solution for non-destructive UV imaging and detection. By combining it with other imaging techniques, such as hyperspectral imaging or time-resolved imaging, a compact and versatile UV computational imaging platform may be realized for future applications.

physics.optics

Dual-mode adaptive-SVD ghost imaging

In this paper, we present a dual-mode adaptive singular value decomposition ghost imaging (A-SVD GI), which can be easily switched between the modes of imaging and edge detection. It can adaptively localize the foreground pixels via a threshold selection method. Then only the foreground region is illuminated by the singular value decomposition (SVD) - based patterns, consequently retrieving high-quality images with fewer sampling ratios. By changing the selecting range of foreground pixels, the A-SVD GI can be switched to the mode of edge detection to directly reveal the edge of objects, without needing the original image. We investigate the performance of these two modes through both numerical simulations and experiments. We also develop a single-round scheme to halve measurement numbers in experiments, instead of separately illuminating positive and negative patterns in traditional methods. The binarized SVD patterns, generated by the spatial dithering method, are modulated by a digital micromirror device (DMD) to speed up the data acquisition. This dual-mode A-SVD GI can be applied in various applications, such as remote sensing or target recognition, and could be further extended for multi-modality functional imaging/detection.

eess.IV