SearcharxivSearch

arXiv subjects

Jiayin Zhao

Publications and source records attributed to Jiayin Zhao.

8 recordsLinked to original sources

V2V3D: View-to-View Denoised 3D Reconstruction for Light-Field Microscopy

Light field microscopy (LFM) has gained significant attention due to its ability to capture snapshot-based, large-scale 3D fluorescence images. However, existing LFM reconstruction algorithms are highly sensitive to sensor noise or require hard-to-get ground-truth annotated data for training. To address these challenges, this paper introduces V2V3D, an unsupervised view2view-based framework that establishes a new paradigm for joint optimization of image denoising and 3D reconstruction in a unified architecture. We assume that the LF images are derived from a consistent 3D signal, with the noise in each view being independent. This enables V2V3D to incorporate the principle of noise2noise for effective denoising. To enhance the recovery of high-frequency details, we propose a novel wave-optics-based feature alignment technique, which transforms the point spread function, used for forward propagation in wave optics, into convolution kernels specifically designed for feature alignment. Moreover, we introduce an LFM dataset containing LF images and their corresponding 3D intensity volumes. Extensive experiments demonstrate that our approach achieves high computational efficiency and outperforms the other state-of-the-art methods. These advancements position V2V3D as a promising solution for 3D imaging under challenging conditions.

cs.CV

A Comprehensive Review of Fish Feeding Behavior Analysis in Aquaculture: Tasks, Techniques, and Applications

Fish feeding behavior analysis is a key foundation for intelligent feeding and precision aquaculture management, and plays an important role in improving feed utilization efficiency, reducing production costs, and mitigating environmental burden. Existing reviews mainly focus on specific technical modalities or related applications in smart aquaculture, which makes it difficult to present the overall development of fish feeding behavior analysis in a comprehensive manner. To address these issues, this paper provides a thematic review of fish feeding behavior analysis in aquaculture, and systematically examines its task definition, technical support, and application status. First, from the task perspective, two core subtasks of fish feeding behavior analysis are clearly distinguished, and relevant behavioral characteristics and evaluation metrics are summarized. Second, from the technical perspective, the development trajectories of computer vision, acoustics, sensors, and multimodal fusion technologies are examined, and their advantages, limitations, and applicable scenarios are analyzed. On this basis, the application value of fish feeding behavior analysis in intelligent feeding and aquaculture management is further summarized. Finally, this paper discusses the challenges in robust perception under complex environments, generalization across fish species and farming scenarios, collaborative multimodal modeling and lightweight deployment, closed loop intelligent feeding, coordinated optimization of multiple tasks, and long-term production validation, and outlines future research directions. This review provides a reference for task standardization, technical selection, and engineering application in fish feeding behavior analysis, and offers insights into the development of smart aquaculture and sustainable aquaculture management.

cs.CV

Progressive Multimodal Interaction Network for Reliable Quantification of Fish Feeding Intensity in Aquaculture

Accurate quantification of fish feeding intensity is crucial for precision feeding in aquaculture, as it directly affects feed utilization and farming efficiency. Although multimodal fusion has proven to be an effective solution, existing methods often overlook the inconsistencies in responses and decision conflicts between different modalities, thus limiting the reliability of the quantification results. To address this issue, this paper proposes a Progressive Multimodal Interaction Network (PMIN) that integrates image, audio, and water-wave data for fish feeding intensity quantification. Specifically, a unified feature extraction framework is first constructed to map inputs from different modalities into a structurally consistent feature space, thereby reducing representational discrepancies across modalities. Then, an auxiliary-modality reinforcement primary-modality mechanism is designed to facilitate the fusion of cross-modal information, which is achieved through channel aware recalibration and dual-stage attention interaction. Furthermore, a decision fusion strategy based on adaptive evidence reasoning is introduced to jointly model the confidence, reliability, and conflicts of modality-specific outputs, so as to improve the stability and robustness of the final judgment. Experiments are conducted on a multimodal fish feeding intensity dataset containing 7089 samples. The results show that PMIN has an accuracy of 96.76%, while maintaining relatively low parameter count and computational cost, and its overall performance outperforms both homogeneous and heterogeneous comparison models. Ablation studies, comparative experiments, and real-world application results further validate the effectiveness and superiority of the proposed method. It can provide reliable support for automated feeding monitoring and precise feeding decisions in smart aquaculture.

cs.CV

Measuring elastic properties of granular hydrogels: Effects of capillary interaction and ionic conditions

The elastic properties of granular hydrogels are commonly characterised under wet conditions, yet the influence of capillary interactions remains unclear. In practical applications, hydrogels operate in aqueous environments containing dissolved ionic species, where swelling and elastic behaviour depend sensitively on ionic conditions. In this study, an experimental setup is developed to measure elastic responses of granular hydrogels under wet conditions. This setup directly observes liquid bridges formation and its evolution during compression. Our results show that neglecting capillary contributions leads to a systematic underestimation of the Young's modulus of hydrogels. Such an underestimation due to the capillary interaction increases as the sample size or its intrinsic stiffness decreases. In addition to the swelling ratio, the tested samples were also prepared under controlled salinity levels. The experimentally observed dependence of stiffness on swelling and salinity conditions is well captured by a modified constitutive model. The development of this study offers a robust testing protocol for measuring elastic properties of hydrogels under various environmental conditions.

cond-mat.soft

FMRFT: Fusion Mamba and DETR for Query Time Sequence Intersection Fish Tracking

Early detection of abnormal fish behavior caused by disease or hunger can be achieved through fish tracking using deep learning techniques, which holds significant value for industrial aquaculture. However, underwater reflections and some reasons with fish, such as the high similarity, rapid swimming caused by stimuli and mutual occlusion bring challenges to multi-target tracking of fish. To address these challenges, this paper establishes a complex multi-scenario sturgeon tracking dataset and introduces the FMRFT model, a real-time end-to-end fish tracking solution. The model incorporates the low video memory consumption Mamba In Mamba (MIM) architecture, which facilitates multi-frame temporal memory and feature extraction, thereby addressing the challenges to track multiple fish across frames. Additionally, the FMRFT model with the Query Time Sequence Intersection (QTSI) module effectively manages occluded objects and reduces redundant tracking frames using the superior feature interaction and prior frame processing capabilities of RT-DETR. This combination significantly enhances the accuracy and stability of fish tracking. Trained and tested on the dataset, the model achieves an IDF1 score of 90.3% and a MOTA accuracy of 94.3%. Experimental results show that the proposed FMRFT model effectively addresses the challenges of high similarity and mutual occlusion in fish populations, enabling accurate tracking in factory farming environments.

cs.CV

PNR: Physics-informed Neural Representation for high-resolution LFM reconstruction

Light field microscopy (LFM) has been widely utilized in various fields for its capability to efficiently capture high-resolution 3D scenes. Despite the rapid advancements in neural representations, there are few methods specifically tailored for microscopic scenes. Existing approaches often do not adequately address issues such as the loss of high-frequency information due to defocus and sample aberration, resulting in suboptimal performance. In addition, existing methods, including RLD, INR, and supervised U-Net, face challenges such as sensitivity to initial estimates, reliance on extensive labeled data, and low computational efficiency, all of which significantly diminish the practicality in complex biological scenarios. This paper introduces PNR (Physics-informed Neural Representation), a method for high-resolution LFM reconstruction that significantly enhances performance. Our method incorporates an unsupervised and explicit feature representation approach, resulting in a 6.1 dB improvement in PSNR than RLD. Additionally, our method employs a frequency-based training loss, enabling better recovery of high-frequency details, which leads to a reduction in LPIPS by at least half compared to SOTA methods (1.762 V.S. 3.646 of DINER). Moreover, PNR integrates a physics-informed aberration correction strategy that optimizes Zernike polynomial parameters during optimization, thereby reducing the information loss caused by aberrations and improving spatial resolution. These advancements make PNR a promising solution for long-term high-resolution biological imaging applications. Our code and dataset will be made publicly available.

eess.IV

OPAL: Occlusion Pattern Aware Loss for Unsupervised Light Field Disparity Estimation

Light field disparity estimation is an essential task in computer vision with various applications. Although supervised learning-based methods have achieved both higher accuracy and efficiency than traditional optimization-based methods, the dependency on ground-truth disparity for training limits the overall generalization performance not to say for real-world scenarios where the ground-truth disparity is hard to capture. In this paper, we argue that unsupervised methods can achieve comparable accuracy, but, more importantly, much higher generalization capacity and efficiency than supervised methods. Specifically, we present the Occlusion Pattern Aware Loss, named OPAL, which successfully extracts and encodes the general occlusion patterns inherent in the light field for loss calculation. OPAL enables: i) accurate and robust estimation by effectively handling occlusions without using any ground-truth information for training and ii) much efficient performance by significantly reducing the network parameters required for accurate inference. Besides, a transformer-based network and a refinement module are proposed for achieving even more accurate results. Extensive experiments demonstrate our method not only significantly improves the accuracy compared with the SOTA unsupervised methods, but also possesses strong generalization capacity, even for real-world data, compared with supervised methods. Our code will be made publicly available.

cs.CV

Reducible and irreducible approximation of complex symmetric operators

This paper aims to study reducible and irreducible approximation in the set $\textsl{CSO}$ of all complex symmetric operators on a separable, complex Hilbert space $\mathcal H$. When ${\rm dim} \mathcal H=\infty$, it is proved that both those reducible ones and those irreducible ones are norm dense in $\textsl{CSO}$. When ${\rm dim} \mathcal H<\infty$, irreducible complex symmetric operators constitute an open, dense subset of $\textsl{CSO}$.

math.FA