SearcharxivSearch

arXiv subjects

Yi Yu

Publications and source records attributed to Yi Yu.

At least 127 records · Page 7Linked to original sources

Dynamic Tuning of Single-Photon Emission in Monolayer WSe2 via Localized Strain Engineering

Two-dimensional (2D) materials have emerged as promising candidates for next-generation integrated single-photon emitters (SPEs). However, significant variability in the emission energies of 2D SPEs presents a major challenge in producing identical single photons from different SPEs, which may become crucial for various quantum applications including quantum information processing. Although various approaches to dynamically tuning the emission energies of 2D SPEs have been developed to address the issue, the practical solution to matching multiple individual SPEs in a single 2D flake is still scarce. In this work, we demonstrate a precise emission energy tuning of individual SPEs in a WSe2 monolayer. Our approach utilizes localized strain fields near individual SPEs, which we control independently by adjusting the physical volume of an SU-8-based stressor layer via focused laser annealing. This technique allows continuous emission energy tuning of up to 15 meV while maintaining the qualities of SPEs. Additionally, we showcase the precise spectral alignment of three distinct SPEs in a single WSe2 monolayer to the same wavelength. The tunability of 2D SPEs represents a solid step towards the on-chip integrated photonics with 2D materials for quantum technologies.

cond-mat.mes-hall

Backdoor Attacks against No-Reference Image Quality Assessment Models via a Scalable Trigger

No-Reference Image Quality Assessment (NR-IQA), responsible for assessing the quality of a single input image without using any reference, plays a critical role in evaluating and optimizing computer vision systems, e.g., low-light enhancement. Recent research indicates that NR-IQA models are susceptible to adversarial attacks, which can significantly alter predicted scores with visually imperceptible perturbations. Despite revealing vulnerabilities, these attack methods have limitations, including high computational demands, untargeted manipulation, limited practical utility in white-box scenarios, and reduced effectiveness in black-box scenarios. To address these challenges, we shift our focus to another significant threat and present a novel poisoning-based backdoor attack against NR-IQA (BAIQA), allowing the attacker to manipulate the IQA model's output to any desired target value by simply adjusting a scaling coefficient $α$ for the trigger. We propose to inject the trigger in the discrete cosine transform (DCT) domain to improve the local invariance of the trigger for countering trigger diminishment in NR-IQA models due to widely adopted data augmentations. Furthermore, the universal adversarial perturbations (UAP) in the DCT space are designed as the trigger, to increase IQA model susceptibility to manipulation and improve attack effectiveness. In addition to the heuristic method for poison-label BAIQA (P-BAIQA), we explore the design of clean-label BAIQA (C-BAIQA), focusing on $α$ sampling and image data refinement, driven by theoretical insights we reveal. Extensive experiments on diverse datasets and various NR-IQA models demonstrate the effectiveness of our attacks. Code can be found at https://github.com/yuyi-sd/BAIQA.

cs.CV

Wholly-WOOD: Wholly Leveraging Diversified-quality Labels for Weakly-supervised Oriented Object Detection

Accurately estimating the orientation of visual objects with compact rotated bounding boxes (RBoxes) has become a prominent demand, which challenges existing object detection paradigms that only use horizontal bounding boxes (HBoxes). To equip the detectors with orientation awareness, supervised regression/classification modules have been introduced at the high cost of rotation annotation. Meanwhile, some existing datasets with oriented objects are already annotated with horizontal boxes or even single points. It becomes attractive yet remains open for effectively utilizing weaker single point and horizontal annotations to train an oriented object detector (OOD). We develop Wholly-WOOD, a weakly-supervised OOD framework, capable of wholly leveraging various labeling forms (Points, HBoxes, RBoxes, and their combination) in a unified fashion. By only using HBox for training, our Wholly-WOOD achieves performance very close to that of the RBox-trained counterpart on remote sensing and other areas, significantly reducing the tedious efforts on labor-intensive annotation for oriented objects. The source codes are available at https://github.com/VisionXLab/whollywood (PyTorch-based) and https://github.com/VisionXLab/whollywood-jittor (Jittor-based).

cs.CV

Point2RBox-v2: Rethinking Point-supervised Oriented Object Detection with Spatial Layout Among Instances

With the rapidly increasing demand for oriented object detection (OOD), recent research involving weakly-supervised detectors for learning OOD from point annotations has gained great attention. In this paper, we rethink this challenging task setting with the layout among instances and present Point2RBox-v2. At the core are three principles: 1) Gaussian overlap loss. It learns an upper bound for each instance by treating objects as 2D Gaussian distributions and minimizing their overlap. 2) Voronoi watershed loss. It learns a lower bound for each instance through watershed on Voronoi tessellation. 3) Consistency loss. It learns the size/rotation variation between two output sets with respect to an input image and its augmented view. Supplemented by a few devised techniques, e.g. edge loss and copy-paste, the detector is further enhanced. To our best knowledge, Point2RBox-v2 is the first approach to explore the spatial layout among instances for learning point-supervised OOD. Our solution is elegant and lightweight, yet it is expected to give a competitive performance especially in densely packed scenes: 62.61%/86.15%/34.71% on DOTA/HRSC/FAIR1M. Code is available at https://github.com/VisionXLab/point2rbox-v2.

cs.CV

A Simple Aerial Detection Baseline of Multimodal Language Models

The multimodal language models (MLMs) based on generative pre-trained Transformer are considered powerful candidates for unifying various domains and tasks. MLMs developed for remote sensing (RS) have demonstrated outstanding performance in multiple tasks, such as visual question answering and visual grounding. In addition to visual grounding that detects specific objects corresponded to given instruction, aerial detection, which detects all objects of multiple categories, is also a valuable and challenging task for RS foundation models. However, aerial detection has not been explored by existing RS MLMs because the autoregressive prediction mechanism of MLMs differs significantly from the detection outputs. In this paper, we present a simple baseline for applying MLMs to aerial detection for the first time, named LMMRotate. Specifically, we first introduce a normalization method to transform detection outputs into textual outputs to be compatible with the MLM framework. Then, we propose a evaluation method, which ensures a fair comparison between MLMs and conventional object detection models. We construct the baseline by fine-tuning open-source general-purpose MLMs and achieve impressive detection performance comparable to conventional detector. We hope that this baseline will serve as a reference for future MLM development, enabling more comprehensive capabilities for understanding RS images. Code is available at https://github.com/Li-Qingyun/mllm-mmrotate.

cs.CV

Transfer Learning for Nonparametric Contextual Dynamic Pricing

Dynamic pricing strategies are crucial for firms to maximize revenue by adjusting prices based on market conditions and customer characteristics. However, designing optimal pricing strategies becomes challenging when historical data are limited, as is often the case when launching new products or entering new markets. One promising approach to overcome this limitation is to leverage information from related products or markets to inform the focal pricing decisions. In this paper, we explore transfer learning for nonparametric contextual dynamic pricing under a covariate shift model, where the marginal distributions of covariates differ between source and target domains while the reward functions remain the same. We propose a novel Transfer Learning for Dynamic Pricing (TLDP) algorithm that can effectively leverage pre-collected data from a source domain to enhance pricing decisions in the target domain. The regret upper bound of TLDP is established under a simple Lipschitz condition on the reward function. To establish the optimality of TLDP, we further derive a matching minimax lower bound, which includes the target-only scenario as a special case and is presented for the first time in the literature. Extensive numerical experiments validate our approach, demonstrating its superiority over existing methods and highlighting its practical utility in real-world applications.

cs.LG

Quasi-Monte Carlo finite element approximation of the Navier-Stokes equations with initial data modeled by log-normal random fields

In this paper, we analyze the numerical approximation of the Navier-Stokes problem over a bounded polygonal domain in $\mathbb{R}^2$, where the initial condition is modeled by a log-normal random field. This problem usually arises in the area of uncertainty quantification. We aim to compute the expectation value of linear functionals of the solution to the Navier-Stokes equations and perform a rigorous error analysis for the problem. In particular, our method includes the finite element, fully-discrete discretizations, truncated Karhunen-Loéve expansion for the realizations of the initial condition, and lattice-based quasi-Monte Carlo (QMC) method to estimate the expected values over the parameter space. Our QMC analysis is based on randomly-shifted lattice rules for the integration over the domain in high-dimensional space, which guarantees the error decays with $\mathcal{O}(N^{-1+δ})$, where $N$ is the number of sampling points, $δ>0$ is an arbitrary small number, and the constant in the decay estimate is independent of the dimension of integration.

math.NA

Enhancement and speed-up of carrier dynamics in a dielectric nanocavity with deep sub-wavelength confinement

The emergence of dielectric bowtie cavities enable optical confinement with ultrahigh quality factor and ultra-small optical mode volumes with perspectives for enhanced light-matter interaction. Experimental work has so far emphasized the realization of these nanocavities. Here, we experimentally investigate the ultrafast dynamics of a topology-optimized dielectric (silicon) bowtie nanocavity, with device dimensions down to 12 nm, that localizes light to a mode volume deep below the so-called diffraction limit given by the half-wavelength cubed. This strong spatial light concentration is shown to significantly enhance the carrier generation rate through two-photon absorption, as well as reducing the time it takes for the carriers to recover. A diffusion time below 1 ps is achieved for the bowtie cavity, which is more than an order of magnitude smaller than for a conventional microcavity. Additionally, parametric effects due to coherent interactions between pump and probe signals are also enhanced in the bowtie cavity, leading to an improved extinction ratio. These results demonstrate important fundamental advantages of dielectric bowtie cavities compared to conventional point-defect cavities, laying a foundation for novel low-power and ultrafast optical devices, including switches and modulators.

physics.optics

A nanolaser with extreme dielectric confinement

The interaction between light and matter can be enhanced by spatially concentrating the light field to boost the photon energy density and increasing the photon dwell time to prolong energy transfer between light and matter. Traditionally, strong spatial light localization has been achieved using plasmonics, which, despite its effectiveness, entails ohmic losses. Recent advances in nanostructured dielectrics offer an avenue for achieving strong light confinement without metallic losses. However, previous studies primarily focused on minimizing the optical mode volume without adequately addressing light-matter interactions. Here, we develop a nanolaser that simultaneously localizes the electromagnetic field and excited carriers within the same region of a dielectric nanobridge. This extreme dielectric confinement of both light and matter achieves a mode volume below the diffraction limit and a subwavelength carrier volume without the introduction of lateral quantum confinement, enabling continuous-wave lasing at room-temperature. Moreover, we observe a strong correlation between the mode field and carrier distribution, and unexpectedly, the enhanced mode field localization automatically leads to more pronounced carrier localization, promoting self-alignment of light and matter, which significantly reduces the laser threshold. We quantify the intensified light-matter interaction with a newly proposed interaction volume, which generalizes the concept of mode volume to a broad class of active media. Our work lays the ground for developing ultra-efficient optoelectronic devices by greatly enhancing light-matter interactions through advanced material nanostructuring.

physics.optics

Robust and Transferable Backdoor Attacks Against Deep Image Compression With Selective Frequency Prior

Recent advancements in deep learning-based compression techniques have surpassed traditional methods. However, deep neural networks remain vulnerable to backdoor attacks, where pre-defined triggers induce malicious behaviors. This paper introduces a novel frequency-based trigger injection model for launching backdoor attacks with multiple triggers on learned image compression models. Inspired by the widely used DCT in compression codecs, triggers are embedded in the DCT domain. We design attack objectives tailored to diverse scenarios, including: 1) degrading compression quality in terms of bit-rate and reconstruction accuracy; 2) targeting task-driven measures like face recognition and semantic segmentation. To improve training efficiency, we propose a dynamic loss function that balances loss terms with fewer hyper-parameters, optimizing attack objectives effectively. For advanced scenarios, we evaluate the attack's resistance to defensive preprocessing and propose a two-stage training schedule with robust frequency selection to enhance resilience. To improve cross-model and cross-domain transferability for downstream tasks, we adjust the classification boundary in the attack loss during training. Experiments show that our trigger injection models, combined with minor modifications to encoder parameters, successfully inject multiple backdoors and their triggers into a single compression model, demonstrating strong performance and versatility. (*Due to the notification of arXiv "The Abstract field cannot be longer than 1,920 characters", the appeared Abstract is shortened. For the full Abstract, please download the Article.)

cs.CV

A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning

Audio-visual correlation learning aims to capture and understand natural phenomena between audio and visual data. The rapid growth of Deep Learning propelled the development of proposals that process audio-visual data and can be observed in the number of proposals in the past years. Thus encouraging the development of a comprehensive survey. Besides analyzing the models used in this context, we also discuss some tasks of definition and paradigm applied in AI multimedia. In addition, we investigate objective functions frequently used and discuss how audio-visual data is exploited in the optimization process, i.e., the different methodologies for representing knowledge in the audio-visual domain. In fact, we focus on how human-understandable mechanisms, i.e., structured knowledge that reflects comprehensible knowledge, can guide the learning process. Most importantly, we provide a summarization of the recent progress of Audio-Visual Correlation Learning (AVCL) and discuss the future research directions.

cs.MM

Deterministic formation of carbon-functionalized quantum emitters in hexagonal boron nitride

Forming single-photon emitters (SPEs) in insulating hexagonal boron nitride (hBN) has sparked wide interests in the quantum photonics. Despite significant progress, it remains challenging to deterministically create SPEs at precise locations with a specific type of element for creating defects. In this study, we present a straightforward approach to generate site-deterministic carbon-functionalized quantum emitters in hBN by harnessing ultrasonic nanoindentation. The obtained SPEs are high-quality and can be scaled up to large arrays in a single fabrication step. Comprehensive experimental analyses reveal that the insertion of carbon atoms into the hBN lattice is the source of the robust quantum emission. Complementary theoretical studies suggest possible candidates for the structural origin of the defects based on our experimental results. This rapid and scalable nanoindentation method provides a new way to create SPE arrays with specific types of atoms, enabling the comprehensive investigation of the origins and mechanics of SPE formations in two-dimensional (2D) materials and beyond.

physics.app-ph

PointOBB-v2: Towards Simpler, Faster, and Stronger Single Point Supervised Oriented Object Detection

Single point supervised oriented object detection has gained attention and made initial progress within the community. Diverse from those approaches relying on one-shot samples or powerful pretrained models (e.g. SAM), PointOBB has shown promise due to its prior-free feature. In this paper, we propose PointOBB-v2, a simpler, faster, and stronger method to generate pseudo rotated boxes from points without relying on any other prior. Specifically, we first generate a Class Probability Map (CPM) by training the network with non-uniform positive and negative sampling. We show that the CPM is able to learn the approximate object regions and their contours. Then, Principal Component Analysis (PCA) is applied to accurately estimate the orientation and the boundary of objects. By further incorporating a separation mechanism, we resolve the confusion caused by the overlapping on the CPM, enabling its operation in high-density scenarios. Extensive comparisons demonstrate that our method achieves a training speed 15.58x faster and an accuracy improvement of 11.60%/25.15%/21.19% on the DOTA-v1.0/v1.5/v2.0 datasets compared to the previous state-of-the-art, PointOBB. This significantly advances the cutting edge of single point supervised oriented detection in the modular track.

cs.CV

Single-Image Shadow Removal Using Deep Learning: A Comprehensive Survey

Shadow removal aims at restoring the image content within shadow regions, pursuing a uniform distribution of illumination that is consistent between shadow and non-shadow regions. {Comparing to other image restoration tasks, there are two unique challenges in shadow removal:} 1) The patterns of shadows are arbitrary, varied, and often have highly complex trace structures, making ``trace-less'' image recovery difficult. 2) The degradation caused by shadows is spatially non-uniform, resulting in inconsistencies in illumination and color between shadow and non-shadow areas. Recent developments in this field are primarily driven by deep learning-based solutions, employing a variety of learning strategies, network architectures, loss functions, and training data. Nevertheless, a thorough and insightful review of deep learning-based shadow removal techniques is still lacking. In this paper, we are the first to provide a comprehensive survey to cover various aspects ranging from technical details to applications. We highlight the major advancements in deep learning-based single-image shadow removal methods, thoroughly review previous research across various categories, and provide insights into the historical progression of these developments. Additionally, we summarize performance comparisons both quantitatively and qualitatively. Beyond the technical aspects of shadow removal methods, we also explore potential future directions for this field.

cs.CV

A Survey for Large Language Models in Biomedicine

Recent breakthroughs in large language models (LLMs) offer unprecedented natural language understanding and generation capabilities. However, existing surveys on LLMs in biomedicine often focus on specific applications or model architectures, lacking a comprehensive analysis that integrates the latest advancements across various biomedical domains. This review, based on an analysis of 484 publications sourced from databases including PubMed, Web of Science, and arXiv, provides an in-depth examination of the current landscape, applications, challenges, and prospects of LLMs in biomedicine, distinguishing itself by focusing on the practical implications of these models in real-world biomedical contexts. Firstly, we explore the capabilities of LLMs in zero-shot learning across a broad spectrum of biomedical tasks, including diagnostic assistance, drug discovery, and personalized medicine, among others, with insights drawn from 137 key studies. Then, we discuss adaptation strategies of LLMs, including fine-tuning methods for both uni-modal and multi-modal LLMs to enhance their performance in specialized biomedical contexts where zero-shot fails to achieve, such as medical question answering and efficient processing of biomedical literature. Finally, we discuss the challenges that LLMs face in the biomedicine domain including data privacy concerns, limited model interpretability, issues with dataset quality, and ethics due to the sensitive nature of biomedical data, the need for highly reliable model outputs, and the ethical implications of deploying AI in healthcare. To address these challenges, we also identify future research directions of LLM in biomedicine including federated learning methods to preserve data privacy and integrating explainable AI methodologies to enhance the transparency of LLMs.

cs.CL

Longitudinal double spin asymmetry of $π^{\pm}$-tagged jet, $Λ$, $\overlineΛ$, and $K_S^0$ in polarized $p+p$ collisions at $\sqrt{s}=200$ $\rm{GeV}$ at STAR

Understanding the origin of the proton spin is one of the most fundamental and challenging questions in QCD. Much progress has been made since the first surprising result by the EMC experiment in the late 1980s. However, the helicity distributions of strange quarks and anti-quarks inside the proton are still not well constrained by the experimental data. Measurement of the longitudinal double spin asymmetry $A_{LL}$ of the inclusive jets tagged with a $π^{\pm}/π^-$ carrying high jet momentum fraction, $z$, in $p+p$ collisions can provide further constraints on the gluon helicity distribution in the proton. In addition, the $A_{LL}$ of $Λ$, $\overlineΛ$ and $K_S^0$ in the longitudinally polarized $p+p$ collisions may shed light on the strange quark and anti-quark helicity distributions. In this contribution, we report the preliminary results on the $A_{LL}$ measurements of inclusive jets tagged with a high-$z$ $π^{\pm}$, and the $Λ$, $\overlineΛ$ and $K_S^0$. We utilize the longitudinally polarized $p+p$ collisions at $\sqrt{s}=200$ $\rm{GeV}$ collected by the STAR experiment with an integrated luminosity of about 52 $\rm{pb^{-1}}$.

hep-ex

Towards Physical World Backdoor Attacks against Skeleton Action Recognition

Skeleton Action Recognition (SAR) has attracted significant interest for its efficient representation of the human skeletal structure. Despite its advancements, recent studies have raised security concerns in SAR models, particularly their vulnerability to adversarial attacks. However, such strategies are limited to digital scenarios and ineffective in physical attacks, limiting their real-world applicability. To investigate the vulnerabilities of SAR in the physical world, we introduce the Physical Skeleton Backdoor Attacks (PSBA), the first exploration of physical backdoor attacks against SAR. Considering the practicalities of physical execution, we introduce a novel trigger implantation method that integrates infrequent and imperceivable actions as triggers into the original skeleton data. By incorporating a minimal amount of this manipulated data into the training set, PSBA enables the system misclassify any skeleton sequences into the target class when the trigger action is present. We examine the resilience of PSBA in both poisoned and clean-label scenarios, demonstrating its efficacy across a range of datasets, poisoning ratios, and model architectures. Additionally, we introduce a trigger-enhancing strategy to strengthen attack performance in the clean label setting. The robustness of PSBA is tested against three distinct backdoor defenses, and the stealthiness of PSBA is evaluated using two quantitative metrics. Furthermore, by employing a Kinect V2 camera, we compile a dataset of human actions from the real world to mimic physical attack situations, with our findings confirming the effectiveness of our proposed attacks. Our project website can be found at https://qichenzheng.github.io/psba-website.

cs.CR

Unlearnable Examples Detection via Iterative Filtering

Deep neural networks are proven to be vulnerable to data poisoning attacks. Recently, a specific type of data poisoning attack known as availability attacks has led to the failure of data utilization for model learning by adding imperceptible perturbations to images. Consequently, it is quite beneficial and challenging to detect poisoned samples, also known as Unlearnable Examples (UEs), from a mixed dataset. In response, we propose an Iterative Filtering approach for UEs identification. This method leverages the distinction between the inherent semantic mapping rules and shortcuts, without the need for any additional information. We verify that when training a classifier on a mixed dataset containing both UEs and clean data, the model tends to quickly adapt to the UEs compared to the clean data. Due to the accuracy gaps between training with clean/poisoned samples, we employ a model to misclassify clean samples while correctly identifying the poisoned ones. The incorporation of additional classes and iterative refinement enhances the model's ability to differentiate between clean and poisoned samples. Extensive experiments demonstrate the superiority of our method over state-of-the-art detection approaches across various attacks, datasets, and poison ratios, significantly reducing the Half Total Error Rate (HTER) compared to existing methods.

cs.CR