SearcharxivSearch

arXiv subjects

Honglin Liu

Publications and source records attributed to Honglin Liu.

16 recordsLinked to original sources

Halo Separation-guided Underwater Multi-scale Image Restoration

Underwater images captured by Autonomous Underwater Vehicles (AUVs) are inevitably affected by artificial light sources, which often produce halos in the foreground of the camera and seriously interfere with the quality of the image. The existing underwater image enhancement methods fail to fully consider this key problem, and the robustness of processing images under artificial light scenes is poor. In practical applications, since underwater image enhancement itself is a very challenging task, the influence of artificial light sources will lead to serious degradation of image performance and affect subsequent vision tasks. In order to effectively deal with this problem, this paper designs a single halo image correction method based on an iterative structure. The network is mainly divided into two sub-networks, one is the halo layer separation sub-network which aims to separate the halo by gradient minimization, and the other is the multi-scale recovery sub-network which aims to recover the image information masked by halo. The UIEB and EUVP synthetic datasets are used for training to ensure that the network can fully learn the characteristics and laws of underwater halo images. Then a large number of halo images taken in an underwater environment with real artificial light are collected for testing. In addition, the brightness distribution characteristics of underwater halo images are analyzed and the radial gradient is introduced to constraint eliminate halo to improve the effect of underwater image restoration.

cs.CV

Conditional Representation Learning for Customized Tasks

Conventional representation learning methods learn a universal representation that primarily captures dominant semantics, which may not always align with customized downstream tasks. For instance, in animal habitat analysis, researchers prioritize scene-related features, whereas universal embeddings emphasize categorical semantics, leading to suboptimal results. As a solution, existing approaches resort to supervised fine-tuning, which however incurs high computational and annotation costs. In this paper, we propose Conditional Representation Learning (CRL), aiming to extract representations tailored to arbitrary user-specified criteria. Specifically, we reveal that the semantics of a space are determined by its basis, thereby enabling a set of descriptive words to approximate the basis for a customized feature space. Building upon this insight, given a user-specified criterion, CRL first employs a large language model (LLM) to generate descriptive texts to construct the semantic basis, then projects the image representation into this conditional feature space leveraging a vision-language model (VLM). The conditional representation better captures semantics for the specific criterion, which could be utilized for multiple customized tasks. Extensive experiments on classification and retrieval tasks demonstrate the superiority and generality of the proposed CRL. The code is available at https://github.com/XLearning-SCU/2025-NeurIPS-CRL.

cs.CV

Creation of Lunar-Like Rims in Ilmenite using Synthetic Solar Wind

Space weathering of lunar minerals, due to bombardment from solar wind (SW) particles and micrometeoroid impacts, modifies the mineralogy within tens of nanometers of the surface, i.e., the rim. Spectroscopic signatures of these modifications, observed via remote sensing, have long been used to gauge surface exposure times on the Moon. However, the relative contributions of SW and micrometeoroids in the creation of rim features are still debated, particularly for the nanometer-scale clusters known as nanophase iron (npFe0), which commonly form in ferrous minerals. We address this issue in the laboratory, using deuterium ions and low-energy electrons as a synthetic solar wind plasma to irradiate ilmenite (FeTiO3), a common lunar mineral. Characterization by high-resolution scanning transmission electron microscopy and electron energy-loss spectroscopy shows that the SW alone creates rims with all the main characteristics of lunar samples. We conclusively identify npFe0 and quantify its distribution as a function of depth and fluence, allowing us to estimate the SW exposure of Apollo soil 71501. Our results confirm that small npFe0 particles (<10 nm in diameter) form from SW irradiation. Such experiments provide microscopic details of space weathering, improving the link between surface modification processes and macroscopic remote-sensing data.

astro-ph.EP

Generalization vs. Hallucination

With fast developments in computational power and algorithms, deep learning has made breakthroughs and been applied in many fields. However, generalization remains to be a critical challenge, and the limited generalization capability severely constrains its practical applications. Hallucination issue is another unresolved conundrum haunting deep learning and large models. By leveraging a physical model of imaging through scattering media, we studied the lack of generalization to system response functions in deep learning, identified its cause, and proposed a universal solution. The research also elucidates the creation process of a hallucination in image prediction and reveals its cause, and the common relationship between generalization and hallucination is discovered and clarified. Generally speaking, it enhances the interpretability of deep learning from a physics-based perspective, and builds a universal physical framework for deep learning in various fields. It may pave a way for direct interaction between deep learning and the real world, facilitating the transition of deep learning from a demo model to a practical tool in diverse applications.

physics.optics

Cross-Dataset Generalization in Deep Learning

Deep learning has been extensively used in various fields, such as phase imaging, 3D imaging reconstruction, phase unwrapping, and laser speckle reduction, particularly for complex problems that lack analytic models. Its data-driven nature allows for implicit construction of mathematical relationships within the network through training with abundant data. However, a critical challenge in practical applications is the generalization issue, where a network trained on one dataset struggles to recognize an unknown target from a different dataset. In this study, we investigate imaging through scattering media and discover that the mathematical relationship learned by the network is an approximation dependent on the training dataset, rather than the true mapping relationship of the model. We demonstrate that enhancing the diversity of the training dataset can improve this approximation, thereby achieving generalization across different datasets, as the mapping relationship of a linear physical model is independent of inputs. This study elucidates the nature of generalization across different datasets and provides insights into the design of training datasets to ultimately address the generalization issue in various deep learning-based applications.

cs.LG

Nonconvex optimization for optimum retrieval of the transmission matrix of a multimode fiber

Transmission matrix (TM) allows light control through complex media such as multimode fibers (MMFs), gaining great attention in areas like biophotonics over the past decade. The measurement of a complex-valued TM is highly desired as it supports full modulation of the light field, yet demanding as the holographic setup is usually entailed. Efforts have been taken to retrieve a TM directly from intensity measurements with several representative phase retrieval algorithms, which still see limitations like slow or suboptimum recovery, especially under noisy environment. Here, a modified non-convex optimization approach is proposed. Through numerical evaluations, it shows that the nonconvex method offers an optimum efficiency of focusing with less running time or sampling rate. The comparative test under different signal-to-noise levels further indicates its improved robustness for TM retrieval. Experimentally, the optimum retrieval of the TM of a MMF is collectively validated by multiple groups of single-spot and multi-spot focusing demonstrations. Focus scanning on the working plane of the MMF is also conducted where our method achieves 93.6% efficiency of the gold standard holography method when the sampling rate is 8. Based on the recovered TM, image transmission through the MMF with high fidelity can be realized via another phase retrieval. Thanks to parallel operation and GPU acceleration, the nonconvex approach can retrieve an 8685$\times$1024 TM (sampling rate=8) with 42.3 s on a regular computer. In brief, the proposed method provides optimum efficiency and fast implementation for TM retrieval, which will facilitate wide applications in deep-tissue optical imaging, manipulation and treatment.

physics.optics

Roles of scattered and ballistic photons in imaging through scattering media: a deep learning-based study

Scattering of light in complex media scrambles optical wavefronts and breaks the principles of conventional imaging methods. For decades, researchers have endeavored to conquer the problem by inventing approaches such as adaptive optics, iterative wavefront shaping, and transmission matrix measurement. That said, imaging through/into thick scattering media remains challenging to date. With the rapid development of computing power, deep learning has been introduced and shown potentials to reconstruct target information through complex media or from rough surfaces. But it also fails once coming to optically thick media where ballistic photons become negligible. Here, instead of treating deep learning only as an image extraction method, whose best-selling advantage is to avoid complicate physical models, we exploit it as a tool to explore the underlying physical principles. By adjusting the weights of ballistic and scattered photons through a random phasemask, it is found that although deep learning can extract images from both scattered and ballistic light, the mechanisms are different: scattering may function as an encryption key and decryption from scattered light is key sensitive, while extraction from ballistic light is stable. Based on this finding, it is hypothesized and experimentally confirmed that the foundation of the generalization capability of trained neural networks for different diffusers can trace back to the contribution of ballistic photons, even though their weights of photon counting in detection are not that significant. Moreover, the study may pave an avenue for using deep learning as a probe in exploring the unknown physical principles in various fields.

physics.optics

Different Channels to Transmit Information in a Scattering Medium

A channel should be built to transmit information from one place to another. Imaging is 2 or higher dimensional information communication. Conventionally, an imaging channel comprises a lens and free spaces of its both sides. The transfer function of each part is known; thus, the response of a conventional imaging channel is known as well. Replacing the lens with a scattering layer, the image can still be extracted from the detection plane. That is to say, the scattering medium reconstructs the channel for imaging. Aided by deep learning, we find that different from the lens there are different channels in a scattering medium, i.e., the same scattering medium can construct different channels to match different manners of source encoding. Moreover, we found that without a valid channel the convolution law for a shift-invariant system, i.e., the output is the convolution of its point spread function (PSF) and the input object, is broken, and information cannot be transmitted onto the detection plane. In other words, valid channels are essential to transmit image information through even a shift-invariant system.

physics.optics

Speckle-based optical cryptosystem and its application for human face recognition via deep learning

Face recognition has recently become ubiquitous in many scenes for authentication or security purposes. Meanwhile, there are increasing concerns about the privacy of face images, which are sensitive biometric data that should be carefully protected. Software-based cryptosystems are widely adopted nowadays to encrypt face images, but the security level is limited by insufficient digital secret key length or computing power. Hardware-based optical cryptosystems can generate enormously longer secret keys and enable encryption at light speed, but most reported optical methods, such as double random phase encryption, are less compatible with other systems due to system complexity. In this study, a plain yet high-efficient speckle-based optical cryptosystem is proposed and implemented. A scattering ground glass is exploited to generate physical secret keys of gigabit length and encrypt face images via seemingly random optical speckles at light speed. Face images can then be decrypted from the random speckles by a well-trained decryption neural network, such that face recognition can be realized with up to 98% accuracy. The proposed cryptosystem has wide applicability, and it may open a new avenue for high-security complex information encryption and decryption by utilizing optical speckles.

cs.CR

MlTr: Multi-label Classification with Transformer

The task of multi-label image classification is to recognize all the object labels presented in an image. Though advancing for years, small objects, similar objects and objects with high conditional probability are still the main bottlenecks of previous convolutional neural network(CNN) based models, limited by convolutional kernels' representational capacity. Recent vision transformer networks utilize the self-attention mechanism to extract the feature of pixel granularity, which expresses richer local semantic information, while is insufficient for mining global spatial dependence. In this paper, we point out the three crucial problems that CNN-based methods encounter and explore the possibility of conducting specific transformer modules to settle them. We put forward a Multi-label Transformer architecture(MlTr) constructed with windows partitioning, in-window pixel attention, cross-window attention, particularly improving the performance of multi-label image classification tasks. The proposed MlTr shows state-of-the-art results on various prevalent multi-label datasets such as MS-COCO, Pascal-VOC, and NUS-WIDE with 88.5%, 95.8%, and 65.5% respectively. The code will be available soon at https://github.com/starmemda/MlTr/

cs.CV

Scattering medium: randomly packed pinhole cameras

When light travels through scattering media, speckles (spatially random distribution of fluctuated intensities) are formed due to the interference of light travelling along different optical paths, preventing the perception of structure, absolute location and dimension of a target within or on the other side of the medium. Currently, the prevailing techniques such as wavefront shaping, optical phase conjugation, scattering matrix measurement, and speckle autocorrelation imaging can only picture the target structure in the absence of prior information. Here we show that a scattering medium can be conceptualized as an assembly of randomly packed pinhole cameras, and the corresponding speckle pattern is a superposition of randomly shifted pinhole images. This provides a new perspective to bridge target, scattering medium, and speckle pattern, allowing one to localize and profile a target quantitatively from speckle patterns perceived from the other side of the scattering medium, which is impossible with all existing methods. The method also allows us to interpret some phenomena of diffusive light that are otherwise challenging to understand. For example, why the morphological appearance of speckle patterns changes with the target, why information is difficult to be extracted from thick scattering media, and what determines the capability of seeing through scattering media. In summary, the concept, whilst in its infancy, opens a new door to unveiling scattering media and information extraction from scattering media in real time.

physics.optics

Lensless Wiener-Khinchin telescope based on high-order spatial autocorrelation of thermal light

The resolution of a conventional imaging system based on first-order field correlation can be directly obtained from the optical transfer function. However, it is challenging to determine the resolution of an imaging system through random media, including imaging through scattering media and imaging through randomly inhomogeneous media, since the point-to-point correspondence between the object and the image plane in these systems cannot be established by the first-order field correlation anymore. In this paper, from the perspective of ghost imaging, we demonstrate for the first time to our knowledge that the point-to-point correspondence in these imaging systems can be quantitatively recovered from the high-order correlation of light fields, and the imaging capability, such as resolution, of such imaging schemes can thus be derived by analyzing high-order correlation of the optical transfer function. Based on this theoretical analysis, we propose a lensless Wiener-Khinchin telescope based on high-order spatial autocorrelation of thermal light, which can acquire the image of an object by a snapshot via using a spatial random phase modulator. As an incoherent imaging approach illuminated by thermal light, lensless Wiener-Khinchin telescope can be applied in many fields such as X-ray astronomical observations.

physics.optics

Sub-wavelength Coherent Imaging of a Pure-Phase Object with Thermal Light

We report, for the first time, the observation of sub-wavelength coherent image of a pure phase object with thermal light,which represents an accurate Fourier transform. We demonstrate that ghost-imaging scheme (GI) retrieves amplitude transmittance knowledge of objects rather than the transmitted intensities as the HBT-type imaging scheme does.

quant-ph

Transmission area and two-photon correlated imaging

The relationship between transmission area of an object imaged and the visibility of its image is investigated in a lensless system. We show that the changes of the visibility are quite different when the transmission area is varied by different manners. An increase of the transmission by adding the slit number leads to a decrease of the visibility. While, the change is adverse when the slit width is widened for a given distance between two slits.

quant-ph

Fourier Analysis of Ghost Imaging

Fourier analysis of ghost imaging (FAGI) is proposed in this paper to analyze the properties of ghost imaging with thermal light sources. This new theory is compatible with the general correlation theory of intensity fluctuation and could explain some amazed phenomena. Furthermore we design a series of experiments to verify the new theory and investigate the inherent properties of ghost imaging.

quant-ph

Lensless Fourier-Transform Ghost Imaging with Classical Incoherent Light

The Fourier-Transform ghost imaging of both amplitude-only and pure-phase objects was experimentally observed with classical incoherent light at Fresnel distance by a new lensless scheme. The experimental results are in good agreement with the standard Fourier-transform of the corresponding objects. This scheme provides a new route towards aberration-free diffraction-limited 3D images with classically incoherent thermal light, which have no resolution and depth-of-field limitations of lens-based tomographic systems.

quant-ph