SearcharxivSearch

arXiv subjects

Meng Cui

Publications and source records attributed to Meng Cui.

9 recordsLinked to original sources

Audio-Visual Class-Incremental Learning for Fish Feeding intensity Assessment in Aquaculture

Fish Feeding Intensity Assessment (FFIA) is crucial in industrial aquaculture management. Recent multi-modal approaches have shown promise in improving FFIA robustness and efficiency. However, these methods face significant challenges when adapting to new fish species or environments due to catastrophic forgetting and the lack of suitable datasets. To address these limitations, we first introduce AV-CIL-FFIA, a new dataset comprising 81,932 labelled audio-visual clips capturing feeding intensities across six different fish species in real aquaculture environments. Then, we pioneer audio-visual class incremental learning (CIL) for FFIA and demonstrate through benchmarking on AV-CIL-FFIA that it significantly outperforms single-modality methods. Existing CIL methods rely heavily on historical data. Exemplar-based approaches store raw samples, creating storage challenges, while exemplar-free methods avoid data storage but struggle to distinguish subtle feeding intensity variations across different fish species. To overcome these limitations, we introduce HAIL-FFIA, a novel audio-visual class-incremental learning framework that bridges this gap with a prototype-based approach that achieves exemplar-free efficiency while preserving essential knowledge through compact feature representations. Specifically, HAIL-FFIA employs hierarchical representation learning with a dual-path knowledge preservation mechanism that separates general intensity knowledge from fish-specific characteristics. Additionally, it features a dynamic modality balancing system that adaptively adjusts the importance of audio versus visual information based on feeding behaviour stages. Experimental results show that HAIL-FFIA is superior to SOTA methods on AV-CIL-FFIA, achieving higher accuracy with lower storage needs while effectively mitigating catastrophic forgetting in incremental fish species learning.

cs.LG

Fish Tracking, Counting, and Behaviour Analysis in Digital Aquaculture: A Comprehensive Survey

Digital aquaculture leverages advanced technologies and data-driven methods, providing substantial benefits over traditional aquaculture practices. This paper presents a comprehensive review of three interconnected digital aquaculture tasks, namely, fish tracking, counting, and behaviour analysis, using a novel and unified approach. Unlike previous reviews which focused on single modalities or individual tasks, we analyse vision-based (i.e. image- and video-based), acoustic-based, and biosensor-based methods across all three tasks. We examine their advantages, limitations, and applications, highlighting recent advancements and identifying critical cross-cutting research gaps. The review also includes emerging ideas such as applying multi-task learning and large language models to address various aspects of fish monitoring, an approach not previously explored in aquaculture literature. We identify the major obstacles hindering research progress in this field, including the scarcity of comprehensive fish datasets and the lack of unified evaluation standards. To overcome the current limitations, we explore the potential of using emerging technologies such as multimodal data fusion and deep learning to improve the accuracy, robustness, and efficiency of integrated fish monitoring systems. In addition, we provide a summary of existing datasets available for fish tracking, counting, and behaviour analysis. This holistic perspective offers a roadmap for future research, emphasizing the need for comprehensive datasets and evaluation standards to facilitate meaningful comparisons between technologies and to promote their practical implementations in real-world settings.

q-bio.QM

Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions

Audio-visual speaker tracking has drawn increasing attention over the past few years due to its academic values and wide applications. Audio and visual modalities can provide complementary information for localization and tracking. With audio and visual information, the Bayesian-based filter and deep learning-based methods can solve the problem of data association, audio-visual fusion and track management. In this paper, we conduct a comprehensive overview of audio-visual speaker tracking. To our knowledge, this is the first extensive survey over the past five years. We introduce the family of Bayesian filters and summarize the methods for obtaining audio-visual measurements. In addition, the existing trackers and their performance on the AV16.3 dataset are summarized. In the past few years, deep learning techniques have thrived, which also boost the development of audio-visual speaker tracking. The influence of deep learning techniques in terms of measurement extraction and state estimation is also discussed. Finally, we discuss the connections between audio-visual speaker tracking and other areas such as speech separation and distributed speaker tracking.

cs.MM

Multimodal Fish Feeding Intensity Assessment in Aquaculture

Fish feeding intensity assessment (FFIA) aims to evaluate fish appetite changes during feeding, which is crucial in industrial aquaculture applications. Existing FFIA methods are limited by their robustness to noise, computational complexity, and the lack of public datasets for developing the models. To address these issues, we first introduce AV-FFIA, a new dataset containing 27,000 labeled audio and video clips that capture different levels of fish feeding intensity. Then, we introduce multi-modal approaches for FFIA by leveraging the models pre-trained on individual modalities and fused with data fusion methods. We perform benchmark studies of these methods on AV-FFIA, and demonstrate the advantages of the multi-modal approach over the single-modality based approach, especially in noisy environments. However, compared to the methods developed for individual modalities, the multimodal approaches may involve higher computational costs due to the need for independent encoders for each modality. To overcome this issue, we further present a novel unified mixed-modality based method for FFIA, termed as U-FFIA. U-FFIA is a single model capable of processing audio, visual, or audio-visual modalities, by leveraging modality dropout during training and knowledge distillation using the models pre-trained with data from single modality. We demonstrate that U-FFIA can achieve performance better than or on par with the state-of-the-art modality-specific FFIA models, with significantly lower computational overhead, enabling robust and efficient FFIA for improved aquaculture management.

cs.SD

WavJourney: Compositional Audio Creation with Large Language Models

Despite breakthroughs in audio generation models, their capabilities are often confined to domain-specific conditions such as speech transcriptions and audio captions. However, real-world audio creation aims to generate harmonious audio containing various elements such as speech, music, and sound effects with controllable conditions, which is challenging to address using existing audio generation systems. We present WavJourney, a novel framework that leverages Large Language Models (LLMs) to connect various audio models for audio creation. WavJourney allows users to create storytelling audio content with diverse audio elements simply from textual descriptions. Specifically, given a text instruction, WavJourney first prompts LLMs to generate an audio script that serves as a structured semantic representation of audio elements. The audio script is then converted into a computer program, where each line of the program calls a task-specific audio generation model or computational operation function. The computer program is then executed to obtain a compositional and interpretable solution for audio creation. Experimental results suggest that WavJourney is capable of synthesizing realistic audio aligned with textually-described semantic, spatial and temporal conditions, achieving state-of-the-art results on text-to-audio generation benchmarks. Additionally, we introduce a new multi-genre story benchmark. Subjective evaluations demonstrate the potential of WavJourney in crafting engaging storytelling audio content from text. We further demonstrate that WavJourney can facilitate human-machine co-creation in multi-round dialogues. To foster future research, the code and synthesized audio are available at: https://audio-agi.github.io/WavJourney_demopage/.

cs.SD

Roadmap on Wavefront Shaping and deep imaging in complex media

The last decade has seen the development of a wide set of tools, such as wavefront shaping, computational or fundamental methods, that allow to understand and control light propagation in a complex medium, such as biological tissues or multimode fibers. A vibrant and diverse community is now working on this field, that has revolutionized the prospect of diffraction-limited imaging at depth in tissues. This roadmap highlights several key aspects of this fast developing field, and some of the challenges and opportunities ahead.

physics.optics

A high throughput (>90%), large compensation range, single-prism femtosecond pulse compressor

We demonstrate a high throughput, large compensation range, single-prism femtosecond pulse compressor, using a single prism and two roof mirrors. The compressor has zero angular dispersion, zero spatial dispersion, zero pulse-front tilt, and unity magnification. The high efficiency is achieved by adopting two roof mirrors as the retroreflectors. We experimentally achieved ~ -14500 fs2 group delay dispersion (GDD) with 30 cm of prism tip-roof mirror prism separation, and ~90.7% system throughput with the current implementation. With better components, the throughput can be even higher.

physics.optics

Breaking the spatial resolution barrier via iterative sound-light interaction in deep tissue microscopy

Optical microscopy has so far been restricted to superficial layers, leaving many important biological questions unanswered. Random scattering causes the ballistic focus, which is conventionally used for image formation, to decay exponentially with depth. Optical imaging beyond the ballistic regime has been demonstrated by hybrid techniques that combine light with the deeper penetration capability of sound waves. Deep inside highly scattering media, the sound focus dimensions restrict the imaging resolutions. Here we show that by iteratively focusing light into an ultrasound focus via phase conjugation, we can fundamentally overcome this resolution barrier in deep tissues and at the same time increase the focus to background ratio. We demonstrate fluorescence microscopy beyond the ballistic regime of light with a threefold improved resolution and a fivefold increase in contrast. This development opens up practical high resolution fluorescence imaging in deep tissues.

physics.optics

Fluorescence microscopy beyond the ballistic regime by ultrasound pulse guided digital phase conjugation

Fluorescence microscopy has revolutionized biomedical research over the past three decades. Its high molecular specificity and unrivaled single molecule level sensitivity have enabled breakthroughs in a variety of research fields. For in vivo applications, its major limitation is the superficial imaging depth as random scattering in biological tissues causes exponential attenuation of the ballistic component of a light wave. Here we present fluorescence microscopy beyond the ballistic regime by combining single cycle pulsed ultrasound modulation and digital optical phase conjugation. We demonstrate near isotropic 3D localized sound-light interaction with an imaging depth as high as thirteen scattering path lengths. With the exceptionally high optical gain provided by the digital optical phase conjugation system, we can deliver sufficient optical power to a focus inside highly scattering media for not only fluorescence microscopy but also a variety of linear and nonlinear spectroscopy measurements. This technology paves the way for many important applications in both fundamental biology research and clinical studies.

physics.optics