Searcharxiv⌕ Search

arXiv subjects

Yu Guo

Publications and source records attributed to Yu Guo.

At least 73 records · Page 4Linked to original sources

SIRR-LMM: Single-image Reflection Removal via Large Multimodal Model

Glass surfaces create complex interactions of reflected and transmitted light, making single-image reflection removal (SIRR) challenging. Existing datasets suffer from limited physical realism in synthetic data or insufficient scale in real captures. We introduce a synthetic dataset generation framework that path-traces 3D glass models over real background imagery to create physically accurate reflection scenarios with varied glass properties, camera settings, and post-processing effects. To leverage the capabilities of Large Multimodal Model (LMM), we concatenate the image layers into a single composite input, apply joint captioning, and fine-tune the model using task-specific LoRA rather than full-parameter training. This enables our approach to achieve improved reflection removal and separation performance compared to state-of-the-art methods.

cs.CV↗

DSP-Reg: Domain-Sensitive Parameter Regularization for Robust Domain Generalization

Domain Generalization (DG) is a critical area that focuses on developing models capable of performing well on data from unseen distributions, which is essential for real-world applications. Existing approaches primarily concentrate on learning domain-invariant features, which assume that a model robust to variations in the source domains will generalize well to unseen target domains. However, these approaches neglect a deeper analysis at the parameter level, which makes the model hard to explicitly differentiate between parameters sensitive to domain shifts and those robust, potentially hindering its overall ability to generalize. In order to address these limitations, we first build a covariance-based parameter sensitivity analysis framework to quantify the sensitivity of each parameter in a model to domain shifts. By computing the covariance of parameter gradients across multiple source domains, we can identify parameters that are more susceptible to domain variations, which serves as our theoretical foundation. Based on this, we propose Domain-Sensitive Parameter Regularization (DSP-Reg), a principled framework that guides model optimization by a soft regularization technique that encourages the model to rely more on domain-invariant parameters while suppressing those that are domain-specific. This approach provides a more granular control over the model's learning process, leading to improved robustness and generalization to unseen domains. Extensive experiments on benchmarks, such as PACS, VLCS, OfficeHome, and DomainNet, demonstrate that DSP-Reg outperforms state-of-the-art approaches, achieving an average accuracy of 66.7\% and surpassing all baselines.

cs.LG↗

Advanced Long-term Earth System Forecasting

Reliable long-term forecasting of Earth system dynamics is fundamentally limited by instabilities in current artificial intelligence (AI) models during extended autoregressive simulations. These failures often originate from inherent spectral bias, leading to inadequate representation of critical high-frequency, small-scale processes and subsequent uncontrolled error amplification. Inspired by the nested grids in numerical models used to resolve small scales, we present TritonCast. At the core of its design is a dedicated latent dynamical core, which ensures the long-term stability of the macro-evolution at a coarse scale. An outer structure then fuses this stable trend with fine-grained local details. This design effectively mitigates the spectral bias caused by cross-scale interactions. In atmospheric science, it achieves state-of-the-art accuracy on the WeatherBench 2 benchmark while demonstrating exceptional long-term stability: executing year-long autoregressive global forecasts and completing multi-year climate simulations that span the entire available $2500$-day test period without drift. In oceanography, it extends skillful eddy forecast to $120$ days and exhibits unprecedented zero-shot cross-resolution generalization. Ablation studies reveal that this performance stems from the synergistic interplay of the architecture's core components. TritonCast thus offers a promising pathway towards a new generation of trustworthy, AI-driven simulations. This significant advance has the potential to accelerate discovery in climate and Earth system science, enabling more reliable long-term forecasting and deeper insights into complex geophysical dynamics.

cs.LG↗

Binarisation-loophole-free observation of high-dimensional quantum nonlocality

Bell inequality tests based on high-dimensional entanglement usually require measurements that can resolve multiple possible outcomes. However, the implementation of high-dimensional multi-outcome measurements is often only emulated via a collection of ``click or no-click'' measurements. This reduction of multi-outcome measurements to binary-outcome measurements opens a loophole in high-dimensional tests Bell inequalities which can be exploited by local hidden variable models [Tavakoli et al., Phys. Rev. A 111, 042433 (2025)]. Here, we close this loophole by using four-dimensional photonic path-mode entanglement and multi-outcome detection. We test both the well-known Collins-Gisin-Linden-Massar-Popescu inequality and a related Bell inequality tailored for maximally entangled states in high-dimension. We observe violations that are large enough to also rule out any quantum model based on entanglement of lower dimension, thereby demonstrating genuinely high-dimensional nonlocality free of the binarisation loophole.

quant-ph↗

LLHA-Net: A Hierarchical Attention Network for Two-View Correspondence Learning

Establishing the correct correspondence of feature points is a fundamental task in computer vision. However, the presence of numerous outliers among the feature points can significantly affect the matching results, reducing the accuracy and robustness of the process. Furthermore, a challenge arises when dealing with a large proportion of outliers: how to ensure the extraction of high-quality information while reducing errors caused by negative samples. To address these issues, in this paper, we propose a novel method called Layer-by-Layer Hierarchical Attention Network, which enhances the precision of feature point matching in computer vision by addressing the issue of outliers. Our method incorporates stage fusion, hierarchical extraction, and an attention mechanism to improve the network's representation capability by emphasizing the rich semantic information of feature points. Specifically, we introduce a layer-by-layer channel fusion module, which preserves the feature semantic information from each stage and achieves overall fusion, thereby enhancing the representation capability of the feature points. Additionally, we design a hierarchical attention module that adaptively captures and fuses global perception and structural semantic information using an attention mechanism. Finally, we propose two architectures to extract and integrate features, thereby improving the adaptability of our network. We conduct experiments on two public datasets, namely YFCC100M and SUN3D, and the results demonstrate that our proposed method outperforms several state-of-the-art techniques in both outlier removal and camera pose estimation. Source code is available at http://www.linshuyuan.com.

cs.CV↗

Measure of entanglement and the monogamy relation: a topical review

Characterizing entanglement, including quantifying and distribution of entanglement, which lies at heart of the quantum resource theory, have been investigated extensively ever since Bennett \etal proposed three seminal measures of entanglement in 1996. Up to now, there are numerous measures of entanglement that have been proposed from different point of view and plenty of monogamy relations have been explored which make the distribution of entanglement became more and more clear. While this is relatively easy in the case of pure states, it is much more intricate for the case of mixed quantum states especially with higher dimension and more particles in the system. We present here an overview of the theory along this line. We outline most of the results in this field historically and focus on the finite-dimensional systems. In particular we emphasize the point of view that (i) which yardsticks haven been applied in quantifying entanglement and its distribution, (ii) what are the substantive characteristics and interrelations of these measures and their monogamy relations mathematically by comparing, and (iii) which concepts should be improved or revised and how they were developed accordingly.

quant-ph↗

Partitewise Entanglement

It is known that $ρ^{AB}$ as a bipartite reduced state of the 3-qubit GHZ state is separable, but part $A$ and part $B$ indeed ``share tripartite entanglement'' in the GHZ state. Namely, whether a state can ``share'' more entanglement is dependent on the global system it lives in. Here we explore such kind of entanglement in any $n$-partite system with arbitrary dimensions, $n\geqslant3$, and call it partitewise entanglement (PWE) which includes pairwise entanglement (PE) proposed in [Phys. Rev. A 110, 032420(2024)] as a special case. We propose three classes of the partitewise entanglement measures which are based on the genuine entanglement measure, the minimal bipartition, and the minimal distance from the partitewise separable states, respectively. The former two methods are far-ranging since all of them are defined by the reduced function. Consequently, we establish the framework of the resource theory of the partitewise entanglement. In addition, we investigate the partitewise entanglement extensibility and give a measure of such extensibility, and from which we find that the maximal partitewise entanglement extension is its purification. At last, the relation between this extensibility and the partitewise entanglement is discussed.

quant-ph↗

Complete $k$-partite entanglement measure

The $k$-partite entanglement, which focus on at most how many particles in the global system are entangled but separable from other particles, is complementary to the $k$-entanglement that reflects how many splitted subsystems are entangled under partitions of the systems in characterizing multipartite entanglement. Very recently, the theory of the complete $k$-entanglement measure has been established in [Phys. Rev. A 110, 012405 (2024)]. Here we investigate whether we can define the complete measure of the $k$-partite entanglement. Consequently, with the same spirit as that of the complete $k$-entanglement measure, we present the axiomatic postulates that a complete $k$-partite entanglement measure should require. Furthermore, we present two classes of $k$-partite entanglement measures and show that one is complete while the other one is unified but not complete except for the case of $k=2$.

quant-ph↗

Neural Geometry Image-Based Representations with Optimal Transport (OT)

Neural representations for 3D meshes are emerging as an effective solution for compact storage and efficient processing. Existing methods often rely on neural overfitting, where a coarse mesh is stored and progressively refined through multiple decoder networks. While this can restore high-quality surfaces, it is computationally expensive due to successive decoding passes and the irregular structure of mesh data. In contrast, images have a regular structure that enables powerful super-resolution and restoration frameworks, but applying these advantages to meshes is difficult because their irregular connectivity demands complex encoder-decoder architectures. Our key insight is that a geometry image-based representation transforms irregular meshes into a regular image grid, making efficient image-based neural processing directly applicable. Building on this idea, we introduce our neural geometry image-based representation, which is decoder-free, storage-efficient, and naturally suited for neural processing. It stores a low-resolution geometry-image mipmap of the surface, from which high-quality meshes are restored in a single forward pass. To construct geometry images, we leverage Optimal Transport (OT), which resolves oversampling in flat regions and undersampling in feature-rich regions, and enables continuous levels of detail (LoD) through geometry-image mipmapping. Experimental results demonstrate state-of-the-art storage efficiency and restoration accuracy, measured by compression ratio (CR), Chamfer distance (CD), and Hausdorff distance (HD).

cs.CV↗

Inverse Rendering for High-Genus Surface Meshes from Multi-View Images

We present a topology-informed inverse rendering approach for reconstructing high-genus surface meshes from multi-view images. Compared to 3D representations like voxels and point clouds, mesh-based representations are preferred as they enable the application of differential geometry theory and are optimized for modern graphics pipelines. However, existing inverse rendering methods often fail catastrophically on high-genus surfaces, leading to the loss of key topological features, and tend to oversmooth low-genus surfaces, resulting in the loss of surface details. This failure stems from their overreliance on Adam-based optimizers, which can lead to vanishing and exploding gradients. To overcome these challenges, we introduce an adaptive V-cycle remeshing scheme in conjunction with a re-parametrized Adam optimizer to enhance topological and geometric awareness. By periodically coarsening and refining the deforming mesh, our method informs mesh vertices of their current topology and geometry before optimization, mitigating gradient issues while preserving essential topological features. Additionally, we enforce topological consistency by constructing topological primitives with genus numbers that match those of ground truth using Gauss-Bonnet theorem. Experimental results demonstrate that our inverse rendering approach outperforms the current state-of-the-art method, achieving significant improvements in Chamfer Distance and Volume IoU, particularly for high-genus surfaces, while also enhancing surface details for low-genus surfaces.

cs.GR↗

Bayesian inference of the magnetic component of quark-gluon plasma

The chromo-magnetic monopoles (CMM), emergent topological excitations of non-Abelian gauge fields carrying chromo-magnetic charge, have long been postulated to play an important role in the vacuum confinement of quantum chromodynamics (QCD), the deconfinement transition at temperature $T_c\approx 160\rm MeV$, as well as the strongly coupled nature of quark-gluon plasma (QGP). While such CMMs have been found to provide solutions for challenging puzzles from heavy-ion collision measurements, they were typically introduced as model assumptions in the past. Here we show how their very existence can be determined and their abundance extracted in a data-driven way for the first time. Using the \textsc{cujet3} framework for calculations of jet energy loss and analyzing a comprehensive experimental data set for nuclear modification factor ($R_{\mathrm{AA}}$) and elliptic flow ($v_2$) of high-transverse-momentum hadrons, the fraction of CMMs in the QGP is obtained by Bayesian inference and is found to be substantial in the $1\sim 2 T_c$ region. The posterior CMM fraction is further validated by excellent agreement with additional data and is also shown to predict QGP transport properties quantitatively consistent with the state-of-the-art knowledge.

hep-ph↗

Loss investigations of high frequency lithium niobate Lamb wave resonators at ultralow temperatures

Lamb wave resonators (LWRs) operating at ultralow temperatures serve as promising acoustic platforms for implementing microwave-optical transduction and radio frequency (RF) front-ends in aerospace communications because of the exceptional electromechanical coupling (k^2) and frequency scalability. However, the properties of LWRs at cryogenic temperatures have not been well understood yet. Herein, we experimentally investigate the temperature dependence of the quality factor and resonant frequency in higher order antisymmetric LWRs down to millikelvin temperatures. The high-frequency A1 and A3 mode resonators with spurious-free responses are comprehensively designed, fabricated, and characterized. The quality factors of A1 modes gradually increase upon cryogenic cooling and shows 4 times higher than the room temperature value, while A3 mode resonators exhibit a non-monotonic temperature dependence. Our findings provide new insights into loss mechanisms of cryogenic LWRs, paving the way to strong-coupling quantum acoustodynamics and next-generation satellite wireless communications.

physics.app-ph↗

Neptune-X: Active X-to-Maritime Generation for Universal Maritime Object Detection

Maritime object detection is essential for navigation safety, surveillance, and autonomous operations, yet constrained by two key challenges: the scarcity of annotated maritime data and poor generalization across various maritime attributes (e.g., object category, viewpoint, location, and imaging environment). To address these challenges, we propose Neptune-X, a data-centric generative-selection framework that enhances training effectiveness by leveraging synthetic data generation with task-aware sample selection. From the generation perspective, we develop X-to-Maritime, a multi-modality-conditioned generative model that synthesizes diverse and realistic maritime scenes. A key component is the Bidirectional Object-Water Attention module, which captures boundary interactions between objects and their aquatic surroundings to improve visual fidelity. To further improve downstream tasking performance, we propose Attribute-correlated Active Sampling, which dynamically selects synthetic samples based on their task relevance. To support robust benchmarking, we construct the Maritime Generation Dataset, the first dataset tailored for generative maritime learning, encompassing a wide range of semantic conditions. Extensive experiments demonstrate that our approach sets a new benchmark in maritime scene synthesis, significantly improving detection accuracy, particularly in challenging and previously underrepresented settings. The code is available at https://github.com/gy65896/Neptune-X.

cs.CV↗

HSIDMamba: Exploring Bidirectional State-Space Models for Hyperspectral Denoising

Effectively modeling global context information in hyperspectral image (HSI) denoising is crucial, but prevailing methods using convolution or transformers still face localized or computational efficiency limitations. Inspired by the emerging Selective State Space Model (Mamba) with nearly linear computational complexity and efficient long-term modeling, we present a novel HSI denoising network named HSIDMamba (HSDM). HSDM is tailored to exploit the capture of potential spatial-spectral dependencies effectively and efficiently for HSI denoising. In particular, HSDM comprises multiple Hyperspectral Continuous Scan Blocks (HCSB) to strengthen spatial-spectral interactions. HCSB links forward and backward scans and enhances information from eight directions through the State Space Model (SSM), strengthening the context representation learning of HSDM and improving denoising performance more effectively. In addition, to enhance the utilization of spectral information and mitigate the degradation problem caused by long-range scanning, spectral attention mechanism. Extensive evaluations against HSI denoising benchmarks validate the superior performance of HSDM, achieving state-of-the-art performance and surpassing the efficiency of the transformer method SERT by 31%.

cs.CV↗

Highly Efficient and Broadband Optical Delay Line towards a Quantum Memory

We demonstrate a high-efficiency, free space optical delay line utilizing a nested multipass cell architecture. This design supports extended optical paths with low loss, aided by custom broadband dielectric coating that provides high reflectivity across a wide spectral bandwidth. The cell is characterized using polarization-entangled photon pairs, with signal photons routed through the delay line and idler photons used as timing reference. Quantum state tomography performed on the entangled pair reveals entanglement preservation with a fidelity of $99.6(9)\%$ following a single-transit delay of up to $687$~ns, accompanied by a photon retrieval efficiency of $95.390(5)%$. The delay is controllable and can be set between $1.8$~ns to $687$~ns in $\sim12.6$~ns increments. The longest delay and wide spectral bandwidth result in a time-bandwidth product of $3.87\times 10^7$. These results position this delay line as a strong candidate for all-optical quantum memories and synchronization modules for scalable quantum networks.

quant-ph↗

Controlling the $\mathcal{PT}$ Symmetry Breaking Threshold in Bipartite Lattice Systems with Floquet Topological Edge States

We investigate the control of the parity-time ($\mathcal{PT}$)-symmetry breaking threshold in a periodically driven one-dimensional dimerized lattice with spatially symmetric gain and loss defects. We elucidate the contrasting roles played by Floquet topological edge states in determining the $\mathcal{PT}$ symmetry breaking threshold within the high- and low-frequency driving regimes. In the high-frequency regime, the participation of topological edge states in $\mathcal{PT}$ symmetry breaking is contingent upon the position of the $\mathcal{PT}$-symmetric defect pairs, whereas in the low-frequency regime, their participation is unconditional and independent of the defect pairs placement, resulting in a universal zero threshold. We establish a direct link between the symmetry-breaking threshold and how the spatial profile of the Floquet topological edge states evolves over one driving period. We further demonstrate that lattices with an odd number of sites exhibit unique threshold patterns, in contrast to even-sized systems. Moreover, applying co-frequency periodic driving to the defect pairs, which preserves time-reversal symmetry, can significantly enhance the $\mathcal{PT}$ symmetry-breaking threshold.

quant-ph↗

Complete Genuine Multipartite Entanglement Monotone

A complete characterization and quantification of entanglement, particularly the multipartite entanglement, remains an unfinished long-term goal in quantum information theory. As long as the multipartite system is concerned, the relation between the entanglement contained in different partitions or different subsystems need to take into account. The complete multipartite entanglement measure and the complete monogamy relation is a framework that just deals with such a issue. In this paper, we put forward conditions to justify whether the multipartite entanglement monotone (MEM) and genuine multipartite entanglement monotone (GMEM) are complete, completely monogamous, and tightly complete monogamous according to the feature of the reduced function. Especially, with the assumption that the maximal reduced function is nonincreasing on average under LOCC, we proposed a class of complete MEMs and a class of complete GMEMs via the maximal reduced function for the first time. By comparison, it is shown that, for the tripartite case, this class of GMEMs is better than the one defined from the minimal bipartite entanglement in literature under the framework of complete MEM and complete monogamy relation. In addition, the relation between monogamy, complete monogamy, and the tightly complete monogamy are revealed in light of different kinds of MEMs and GMEMs.

quant-ph↗

UltraVSR: Achieving Ultra-Realistic Video Super-Resolution with Efficient One-Step Diffusion Space

Diffusion models have shown great potential in generating realistic image detail. However, adapting these models to video super-resolution (VSR) remains challenging due to their inherent stochasticity and lack of temporal modeling. Previous methods have attempted to mitigate this issue by incorporating motion information and temporal layers. However, unreliable motion estimation from low-resolution videos and costly multiple sampling steps with deep temporal layers limit them to short sequences. In this paper, we propose UltraVSR, a novel framework that enables ultra-realistic and temporally-coherent VSR through an efficient one-step diffusion space. A central component of UltraVSR is the Degradation-aware Reconstruction Scheduling (DRS), which estimates a degradation factor from the low-resolution input and transforms the iterative denoising process into a single-step reconstruction from low-resolution to high-resolution videos. To ensure temporal consistency, we propose a lightweight Recurrent Temporal Shift (RTS) module, including an RTS-convolution unit and an RTS-attention unit. By partially shifting feature components along the temporal dimension, it enables effective propagation, fusion, and alignment across frames without explicit temporal layers. The RTS module is integrated into a pretrained text-to-image diffusion model and is further enhanced through Spatio-temporal Joint Distillation (SJD), which improves temporally coherence while preserving realistic details. Additionally, we introduce a Temporally Asynchronous Inference (TAI) strategy to capture long-range temporal dependencies under limited memory constraints. Extensive experiments show that UltraVSR achieves state-of-the-art performance, both qualitatively and quantitatively, in a single sampling step. Code is available at https://github.com/yongliuy/UltraVSR.

cs.CV↗