Searcharxiv⌕ Search

arXiv subjects

Ying Fu

Publications and source records attributed to Ying Fu.

At least 109 records · Page 6Linked to original sources

GTAV-NightRain: Photometric Realistic Large-scale Dataset for Night-time Rain Streak Removal

Rain is transparent, which reflects and refracts light in the scene to the camera. In outdoor vision, rain, especially rain streaks degrade visibility and therefore need to be removed. In existing rain streak removal datasets, although density, scale, direction and intensity have been considered, transparency is not fully taken into account. This problem is particularly serious in night scenes, where the appearance of rain largely depends on the interaction with scene illuminations and changes drastically on different positions within the image. This is problematic, because unrealistic dataset causes serious domain bias. In this paper, we propose GTAV-NightRain dataset, which is a large-scale synthetic night-time rain streak removal dataset. Unlike existing datasets, by using 3D computer graphic platform (namely GTA V), we are allowed to infer the three dimensional interaction between rain and illuminations, which insures the photometric realness. Current release of the dataset contains 12,860 HD rainy images and 1,286 corresponding HD ground truth images in diversified night scenes. A systematic benchmark and analysis are provided along with the dataset to inspire further research.

cs.CV↗

Mutual Contrastive Low-rank Learning to Disentangle Whole Slide Image Representations for Glioma Grading

Whole slide images (WSI) provide valuable phenotypic information for histological assessment and malignancy grading of tumors. The WSI-based grading promises to provide rapid diagnostic support and facilitate digital health. Currently, the most commonly used WSIs are derived from formalin-fixed paraffin-embedded (FFPE) and Frozen section. The majority of automatic tumor grading models are developed based on FFPE sections, which could be affected by the artifacts introduced by tissue processing. The frozen section exists problems such as low quality that might influence training within single modality as well. To overcome this problem in a single modal training and achieve better multi-modal and discriminative representation disentanglement in brain tumor, we propose a mutual contrastive low-rank learning (MCL) scheme to integrate FFPE and frozen sections for glioma grading. We first design a mutual learning scheme to jointly optimize the model training based on FFPE and frozen sections. In this proposed scheme, we design a normalized modality contrastive loss (NMC-loss), which could promote to disentangle multi-modality complementary representation of FFPE and frozen sections from the same patient. To reduce intra-class variance, and increase inter-class margin at intra- and inter-patient levels, we conduct a low-rank (LR) loss. Our experiments show that the proposed scheme achieves better performance than the model trained based on each single modality or mixed modalities and even improves the feature extraction in classical attention-based multiple instances learning methods (MIL). The combination of NMC-loss and low-rank loss outperforms other typical contrastive loss functions.

eess.IV↗

Deep Plug-and-Play Prior for Hyperspectral Image Restoration

Deep-learning-based hyperspectral image (HSI) restoration methods have gained great popularity for their remarkable performance but often demand expensive network retraining whenever the specifics of task changes. In this paper, we propose to restore HSIs in a unified approach with an effective plug-and-play method, which can jointly retain the flexibility of optimization-based methods and utilize the powerful representation capability of deep neural networks. Specifically, we first develop a new deep HSI denoiser leveraging gated recurrent convolution units, short- and long-term skip connections, and an augmented noise level map to better exploit the abundant spatio-spectral information within HSIs. It, therefore, leads to the state-of-the-art performance on HSI denoising under both Gaussian and complex noise settings. Then, the proposed denoiser is inserted into the plug-and-play framework as a powerful implicit HSI prior to tackle various HSI restoration tasks. Through extensive experiments on HSI super-resolution, compressed sensing, and inpainting, we demonstrate that our approach often achieves superior performance, which is competitive with or even better than the state-of-the-art on each task, via a single model without any task-specific training.

eess.IV↗

End-to-End Video Text Spotting with Transformer

Recent video text spotting methods usually require the three-staged pipeline, i.e., detecting text in individual images, recognizing localized text, tracking text streams with post-processing to generate final results. These methods typically follow the tracking-by-match paradigm and develop sophisticated pipelines. In this paper, rooted in Transformer sequence modeling, we propose a simple, but effective end-to-end video text DEtection, Tracking, and Recognition framework (TransDETR). TransDETR mainly includes two advantages: 1) Different from the explicit match paradigm in the adjacent frame, TransDETR tracks and recognizes each text implicitly by the different query termed text query over long-range temporal sequence (more than 7 frames). 2) TransDETR is the first end-to-end trainable video text spotting framework, which simultaneously addresses the three sub-tasks (e.g., text detection, tracking, recognition). Extensive experiments in four video text datasets (i.e.,ICDAR2013 Video, ICDAR2015 Video, Minetto, and YouTube Video Text) are conducted to demonstrate that TransDETR achieves state-of-the-art performance with up to around 8.0% improvements on video text spotting tasks. The code of TransDETR can be found at https://github.com/weijiawu/TransDETR.

cs.CV↗

ProbNVS: Fast Novel View Synthesis with Learned Probability-Guided Sampling

Existing state-of-the-art novel view synthesis methods rely on either fairly accurate 3D geometry estimation or sampling of the entire space for neural volumetric rendering, which limit the overall efficiency. In order to improve the rendering efficiency by reducing sampling points without sacrificing rendering quality, we propose to build a novel view synthesis framework based on learned MVS priors that enables general, fast and photo-realistic view synthesis simultaneously. Specifically, fewer but important points are sampled under the guidance of depth probability distributions extracted from the learned MVS architecture. Based on the learned probability-guided sampling, a neural volume rendering module is elaborately devised to fully aggregate source view information as well as the learned scene structures to synthesize photorealistic target view images. Finally, the rendering results in uncertain, occluded and unreferenced regions can be further improved by incorporating a confidence-aware refinement module. Experiments show that our method achieves 15 to 40 times faster rendering compared to state-of-the-art baselines, with strong generalization capacity and comparable high-quality novel view synthesis performance.

cs.CV↗

Estimating Fine-Grained Noise Model via Contrastive Learning

Image denoising has achieved unprecedented progress as great efforts have been made to exploit effective deep denoisers. To improve the denoising performance in realworld, two typical solutions are used in recent trends: devising better noise models for the synthesis of more realistic training data, and estimating noise level function to guide non-blind denoisers. In this work, we combine both noise modeling and estimation, and propose an innovative noise model estimation and noise synthesis pipeline for realistic noisy image generation. Specifically, our model learns a noise estimation model with fine-grained statistical noise model in a contrastive manner. Then, we use the estimated noise parameters to model camera-specific noise distribution, and synthesize realistic noisy training data. The most striking thing for our work is that by calibrating noise models of several sensors, our model can be extended to predict other cameras. In other words, we can estimate cameraspecific noise models for unknown sensors with only testing images, without laborious calibration frames or paired noisy/clean data. The proposed pipeline endows deep denoisers with competitive performances with state-of-the-art real noise modeling methods.

eess.IV↗

Layer-controlled Ferromagnetism in Atomically Thin CrSiTe$_3$ Flakes

The research on two-dimensional (2D) van der Waals (vdW) ferromagnets has promoted the development of ultrahigh-density and nanoscale data storage. However, intrinsic ferromagnetism in layered magnets is always subject to many factors, such as stacking orders, interlayer couplings, and the number of layers. Here, we report a magnetic transition from soft to hard ferromagnetic behaviors as the thickness of CrSiTe$_3$ flakes decreases down to several nanometers. Phenomenally, in contrast to the negligible hysteresis loop in the bulk counterparts, atomically thin CrSiTe$_3$ shows a rectangular loop with finite magnetization and coercivity as thickness decreases down to ~8 nm, indicative of a single-domain and out-of-plane ferromagnetic order. We find that the stray field is weakened with decreasing thickness, which suppresses the formation of the domain wall. In addition, thickness-dependent ferromagnetic properties also reveal a crossover from 3 dimensional to 2 dimensional Ising ferromagnets at a ~7 nm thickness of CrSiTe$_3$, accompanied by a drop of the Curie temperature from 33 K for bulk to ~17 K for 4 nm sample.

cond-mat.mtrl-sci↗

Dynamic Proximal Unrolling Network for Compressive Imaging

Compressive imaging aims to recover a latent image from under-sampled measurements, suffering from a serious ill-posed inverse problem. Recently, deep neural networks have been applied to this problem with superior results, owing to the learned advanced image priors. These approaches, however, require training separate models for different imaging modalities and sampling ratios, leading to overfitting to specific settings. In this paper, a dynamic proximal unrolling network (dubbed DPUNet) was proposed, which can handle a variety of measurement matrices via one single model without retraining. Specifically, DPUNet can exploit both the embedded observation model via gradient descent and imposed image priors by learned dynamic proximal operators, achieving joint reconstruction. A key component of DPUNet is a dynamic proximal mapping module, whose parameters can be dynamically adjusted at the inference stage and make it adapt to different imaging settings. Experimental results demonstrate that the proposed DPUNet can effectively handle multiple compressive imaging modalities under varying sampling ratios and noise levels via only one trained model, and outperform the state-of-the-art approaches.

eess.IV↗

TFPnP: Tuning-free Plug-and-Play Proximal Algorithm with Applications to Inverse Imaging Problems

Plug-and-Play (PnP) is a non-convex optimization framework that combines proximal algorithms, for example, the alternating direction method of multipliers (ADMM), with advanced denoising priors. Over the past few years, great empirical success has been obtained by PnP algorithms, especially for the ones that integrate deep learning-based denoisers. However, a key challenge of PnP approaches is the need for manual parameter tweaking as it is essential to obtain high-quality results across the high discrepancy in imaging conditions and varying scene content. In this work, we present a class of tuning-free PnP proximal algorithms that can determine parameters such as denoising strength, termination time, and other optimization-specific parameters automatically. A core part of our approach is a policy network for automated parameter search which can be effectively learned via a mixture of model-free and model-based deep reinforcement learning strategies. We demonstrate, through rigorous numerical and visual experiments, that the learned policy can customize parameters to different settings, and is often more efficient and effective than existing handcrafted criteria. Moreover, we discuss several practical considerations of PnP denoisers, which together with our learned policy yield state-of-the-art results. This advanced performance is prevalent on both linear and nonlinear exemplar inverse imaging problems, and in particular shows promising results on compressed sensing MRI, sparse-view CT, single-photon imaging, and phase retrieval.

cs.CV↗

$LnCu_3(OH)_6Cl_3 (Ln = Gd, Tb, Dy)$: Heavy Lanthanides on Spin-1/2 Kagome Magnets

The spin-1/2 kagome antiferromagnets are key prototype materials for studying frustrated magnetism. Three isostructural kagome antiferromagnets LnCu$_3$(OH)$_6$Cl$_3$ (Ln = Gd, Tb, Dy) have been successfully synthesized by the hydrothermal method. LnCu$_3$(OH)$_6$Cl$_3$ adopts space group $P\overline{3}m1$ and features the layered Cu-kagome lattice with lanthanide Ln$^{3+}$ cations sitting at the center of the hexagons. Although heavy lanthanides (Ln = Gd, Tb, Dy) in LnCu$_3$(OH)$_6$Cl$_3$ provide a large effective magnetic moment and ferromagnetic-like spin correlations compared to light-lanthanides (Nd, Sm, Eu) analogues, Cu-kagome holds an antiferromagnetically ordered state at around 17 K like YCu$_3$(OH)$_6$Cl$_3$.

cond-mat.str-el↗

Dzyaloshinskii-Moriya anisotropy effect on field-induced magnon condensation in kagome antiferromagnet $α-Cu_3Mg(OH)_6Br_2$

We performed a comprehensive electron spin resonance, magnetization and heat capacity study on the field-induced magnetic phase transitions in the kagome antiferromagnet $α-Cu_3Mg(OH)_6Br_2$. With the successful preparation of single crystals, we mapped out the magnetic phase diagrams under the $c$-axis and $ab$-plane directional magnetic fields $B$. For $B\|c$, the three-dimensional (3D) magnon Bose-Einstein condensation (BEC) is evidenced by the power law scaling of the transition temperature, $T_c\propto (B_c-B)^{2/3}$. For $B\|ab$, the transition from the canted antiferromagetic (CAFM) state to the fully polarized (FP) state is a crossover rather than phase transition, and the characteristic temperature has a significant deviation from the 3D BEC scaling. The different behaviors of the field-induced magnetic transitions for $B\|c$ and $B\|ab$ could result from the Dzyaloshinkii-Moriya (DM) interaction with the DM vector along the $c$-axis, which preserves the $c$-axis directional spin rotation symmetry and breaks the spin rotation symmetry when $B\|ab$. The 3D magnon BEC scaling for $B\|c$ is immune to the off-stoichiometric disorder in our sample $α-Cu_{3.26}Mg_{0.74}(OH)_6Br_2$. Our findings have the potential to shed light on the investigations of the magnetic anisotropy and disorder effects on the field-induced magnon BEC in the quantum antiferromagnet.

cond-mat.str-el↗

Physics-based Noise Modeling for Extreme Low-light Photography

Enhancing the visibility in extreme low-light environments is a challenging task. Under nearly lightless condition, existing image denoising methods could easily break down due to significantly low SNR. In this paper, we systematically study the noise statistics in the imaging pipeline of CMOS photosensors, and formulate a comprehensive noise model that can accurately characterize the real noise structures. Our novel model considers the noise sources caused by digital camera electronics which are largely overlooked by existing methods yet have significant influence on raw measurement in the dark. It provides a way to decouple the intricate noise structure into different statistical distributions with physical interpretations. Moreover, our noise model can be used to synthesize realistic training data for learning-based low-light denoising algorithms. In this regard, although promising results have been shown recently with deep convolutional neural networks, the success heavily depends on abundant noisy clean image pairs for training, which are tremendously difficult to obtain in practice. Generalizing their trained models to images from new devices is also problematic. Extensive experiments on multiple low-light denoising datasets -- including a newly collected one in this work covering various devices -- show that a deep neural network trained with our proposed noise formation model can reach surprisingly-high accuracy. The results are on par with or sometimes even outperform training with paired real data, opening a new door to real-world extreme low-light photography.

eess.IV↗

Weakly-supervised Semantic Segmentation in Cityscape via Hyperspectral Image

High-resolution hyperspectral images (HSIs) contain the response of each pixel in different spectral bands, which can be used to effectively distinguish various objects in complex scenes. While HSI cameras have become low cost, algorithms based on it have not been well exploited. In this paper, we focus on a novel topic, weakly-supervised semantic segmentation in cityscape via HSIs. It is based on the idea that high-resolution HSIs in city scenes contain rich spectral information, which can be easily associated to semantics without manual labeling. Therefore, it enables low cost, highly reliable semantic segmentation in complex scenes. Specifically, in this paper, we theoretically analyze the HSIs and introduce a weakly-supervised HSI semantic segmentation framework, which utilizes spectral information to improve the coarse labels to a finer degree. The experimental results show that our method can obtain highly competitive labels and even have higher edge fineness than artificial fine labels in some classes. At the same time, the results also show that the refined labels can effectively improve the effect of semantic segmentation. The combination of HSIs and semantic segmentation proves that HSIs have great potential in high-level visual tasks.

cs.CV↗

LocalTrans: A Multiscale Local Transformer Network for Cross-Resolution Homography Estimation

Cross-resolution image alignment is a key problem in multiscale gigapixel photography, which requires to estimate homography matrix using images with large resolution gap. Existing deep homography methods concatenate the input images or features, neglecting the explicit formulation of correspondences between them, which leads to degraded accuracy in cross-resolution challenges. In this paper, we consider the cross-resolution homography estimation as a multimodal problem, and propose a local transformer network embedded within a multiscale structure to explicitly learn correspondences between the multimodal inputs, namely, input images with different resolutions. The proposed local transformer adopts a local attention map specifically for each position in the feature. By combining the local transformer with the multiscale structure, the network is able to capture long-short range correspondences efficiently and accurately. Experiments on both the MS-COCO dataset and the real-captured cross-resolution dataset show that the proposed network outperforms existing state-of-the-art feature-based and deep-learning-based homography estimation methods, and is able to accurately align images under $10\times$ resolution gap.

cs.CV↗

Heat Transport in Herbertsmithite: Can a Quantum Spin Liquid Survive Disorder?

Arguably the most favorable situation for spins to enter the long-sought quantum spin liquid (QSL) state is when they sit on a kagome lattice. No consensus has been reached in theory regarding the true ground state of this promising platform. The experimental efforts, relying mostly on one archetypal material ZnCu$_3$(OH)$_6$Cl$_2$, have also led to diverse possibilities. Apart from subtle interactions in the Hamiltonian, there is the additional degree of complexity associated with disorder in the real material ZnCu$_3$(OH)$_6$Cl$_2$ that haunts most experimental probes. Here we resort to heat transport measurement, a cleaner probe in which instead of contributing directly, the disorder only impacts the signal from the kagome spins. For ZnCu$_3$(OH)$_6$Cl$_2$ and a related QSL candidate Cu$_3$Zn(OH)$_6$FBr, we observed no contribution by any spin excitation nor any field-induced change to the thermal conductivity. These results impose different constraints on various scenarios about the ground state of these two kagome compounds: while a gapped QSL, or certain quantum paramagnetic state other than a QSL, is compatible with our results, a gapless QSL must be dramatically modified by the disorder so that gapless spin excitations are localized.

cond-mat.str-el↗

Disentangled Face Attribute Editing via Instance-Aware Latent Space Search

Recent works have shown that a rich set of semantic directions exist in the latent space of Generative Adversarial Networks (GANs), which enables various facial attribute editing applications. However, existing methods may suffer poor attribute variation disentanglement, leading to unwanted change of other attributes when altering the desired one. The semantic directions used by existing methods are at attribute level, which are difficult to model complex attribute correlations, especially in the presence of attribute distribution bias in GAN's training set. In this paper, we propose a novel framework (IALS) that performs Instance-Aware Latent-Space Search to find semantic directions for disentangled attribute editing. The instance information is injected by leveraging the supervision from a set of attribute classifiers evaluated on the input images. We further propose a Disentanglement-Transformation (DT) metric to quantify the attribute transformation and disentanglement efficacy and find the optimal control factor between attribute-level and instance-specific directions based on it. Experimental results on both GAN-generated and real-world images collectively show that our method outperforms state-of-the-art methods proposed recently by a wide margin. Code is available at https://github.com/yxuhan/IALS.

cs.CV↗

Pressure-enhanced ferromagnetism in layered CrSiTe3 flakes

The research on van der Waals (vdW) layered ferromagnets have promoted the development of nanoscale spintronics and applications. However, low-temperature ferromagnetic properties of these materials greatly hinder their applications. Here, we report pressure-enhanced ferromagnetic behaviours in layered CrSiTe3 flakes revealed by high-pressure magnetic circular dichroism (MCD) measurement. At ambient pressure, CrSiTe3 undergoes a paramagnetic-to-ferromagnetic phase transition at 32.8 K, with a negligible hysteresis loop, indicating a soft ferromagnetic behaviour. Under 4.6 GPa pressure, the soft ferromagnet changes into hard one, signalled by a rectangular hysteretic loop with remnant magnetization at zero field. Interestingly, with further increasing pressure, the coercive field (H_c) dramatically increases from 0.02 T at 4.6 GPa to 0.17 T at 7.8 GPa, and the Curie temperature (T_c^h: the temperature for closing the hysteresis loop) also increases from ~36 K at 4.6 GPa to ~138 K at 7.8 GPa. The influences of pressure on exchange interactions are further investigated by density functional theory calculations, which reveal that the in-plane nearest-neighbor exchange interaction and magneto-crystalline anisotropy increase simultaneously as pressure increases, leading to increased H_c and T_c^h in experiments. The effective interaction between magnetic couplings and external pressure offers new opportunities for both searching room-temperature layered ferromagnets and designing pressure-sensitive magnetic functional devices.

cond-mat.mtrl-sci↗

Multiple Symmetry-Protected Dirac Nodal Lines in A Quasi-One-Dimensional Semimetal

Nodal-line semimetals (NLSMs) contains Dirac/Weyl type band-crossing nodes extending into shapes of line, loop and chain in the reciprocal space, leading to novel band topology and transport responses. Robust NLSMs against spin-orbit coupling typically occur in three-dimensional materials with more symmetry operations to protect the line nodes of band crossing, while the possibilities in lower-dimensional materials are rarely discussed. Here we demonstrate robust NLSM phase in a quasi-one-dimensional nonmagnetic semimetal TaNiTe5. Combining angle-resolved photoemission spectroscopy measurements and first-principles calculations, we reveal how reduced dimension can interact with nonsymmorphic symmetry and result into multiple Dirac-type nodal lines with four-fold degeneracy. Our findings suggest rich physics and application in (quasi-)one-dimensional topological materials and call for further investigation on the interplay between the quantum confinement and nontrivial band topology.

cond-mat.other↗