SearcharxivSearch

arXiv subjects

Fei Kong

Publications and source records attributed to Fei Kong.

At least 19 recordsLinked to original sources

Removing inhomogeneous broadening in cross-relaxation spectra via hole burning

Quantum relaxometry based on nitrogen-vacancy (NV) centers in diamond is an easy-to-use technique that detects magnetic noise by measuring the longitudinal relaxation of NV centers. It is favored in chemical and biological applications due to the robustness against imperfect spin control. By tuning the energy level, NV relaxometry can detect magnetic noise generated by other spins and obtain magnetic resonance spectra through cross relaxation. However, the inhomogeneous broadening of NV centers greatly limits the spectral resolution of cross-relaxation spectra. Here we demonstrate a hole-burning technique to remove the inhomogeneous broadening. We utilize a weak pump field to deplete specific NV centers during the relaxation measurement, which contribute a reverse signal in the cross-relaxation spectra with a much narrower linewidth. Our method retains the convenience and sensitivity of NV relaxometry, opening up new avenues for high-resolution magnetic resonance spectroscopy in complex scenarios.

quant-ph

Filtration and gradation of quantum affine vertex algebras

In this paper, we present a unified construction of quantum affine vertex algebras associated to double Yangians, untwisted quantum affinization algebras, and twisted quantum affinization algebras. We show that these $\hbar$-adic quantum vertex algebras contain a dense quantum vertex subalgebra equipped with a increasing filtration. Moreover, we prove that the associated graded algebras with respect to these filtrations are isomorphic to a $\mathbb Z$-graded dense quantum vertex subalgebra of the quantum affine vertex algebras associated to double Yangians.

math.QA

Quantum affine vertex algebra at root of unity

Let $\mathfrak g$ be a finite simple Lie algebra, and let $r$ denote the ratio of the square length of long roots to that of short roots. Let $\wp>2r$ be an integer and $\zeta$ a primitive $\wp$-th root of unity. Denote by $\mathcal U_\zeta(\widehat{\mathfrak g})$ the Lusztig big quantum affine algebra at root of unity defined by divided powers. In this paper, we establish a current algebra presentation of $\mathcal U_\zeta(\widehat{\mathfrak g})$. Based on this presentation, we construct a $\mathbb Z_\wp$-module quantum vertex algebras $V_{\wp,\tau}^\ell(\mathfrak g)$ for each integer $\ell$. Moreover, we establish a fully faithful functor from the category of smooth weighted $\mathcal U_\zeta(\widehat{\mathfrak g})$-modules of level $\ell$ to the category of $(\mathbb Z_\wp,\chi_\phi)$-equivariant $\phi$-coordinated quasi-modules of $V_{\wp,\tau}^\ell(\mathfrak g)$, where $\chi_\phi:\mathbb Z_\wp\to\mathbb C^\times$ is the group homomorphism defined by $s\mapsto \zeta^s$. We also determine the image of this functor. The structure $V_{\wp,\tau}^\ell(\mathfrak g)$ is substantially different from that of affine vertex algebras. We realize $V_{\wp,\tau}^\ell(\mathfrak g)$ as a deformation of a simpler quantum vertex algebra $V_{\wp,\varepsilon}^\ell(\mathfrak g)$ by using vertex bialgebras, and decompose $V_{\wp,\varepsilon}^\ell(\mathfrak g)$ into a Heisenberg vertex algebra and a more interesting quantum vertex algebra determined by a quiver.

math.QA

Double-Layered Silica-Engineered Fluorescent Nanodiamonds for Catalytic Generation and Quantum Sensing of Active Radicals

Fluorescent nanodiamonds (FNDs) hosting nitrogen-vacancy (NV) centers have attracted considerable attention for quantum sensing applications, particularly owing to notable advancements achieved in the field of weak magnetic signal detection in recent years. Here, we report a practical quantum-sensing platform for the controlled production and real-time monitoring of ultra-short-lived reactive free radicals using a double-layered silica modification strategy. An inner dense silica layer preserves the intrinsic properties of NV centers, while an outer porous silica layer facilitates efficient adsorption and stabilization of hydroxyl radicals and their precursor reactants. By doping this mesoporous shell with gadolinium (III) catalysts, we achieve sustained, light-free generation of hydroxyl radicals via catalytic water splitting, eliminating reliance on external precursors. The mechanism underlying this efficient radical generation is discussed in detail. The radical production is monitored in real time and in situ through spin-dependent T1 relaxometry of the NV centers, demonstrating stable and tunable radical fluxes, with concentration tunable across a continuous range from approximately 100 mM to molar levels by adjusting the catalyst condition. This study extends the technical application of nanodiamonds from relaxation sensing to the controlled synthesis of reactive free radicals, thereby providing robust experimental evidence to support the advancement of quantum sensing systems in intelligent manufacturing.

quant-ph

Quantum relaxometry for detecting biomolecular interactions with single NV centers

The investigation of biomolecular interactions at the single-molecule level has emerged as a pivotal research area in life science, particularly through optical, mechanical, and electrochemical approaches. Spins existing widely in biological systems, offer a unique degree of freedom for detecting such interactions. However, most previous studies have been largely confined to ensemble-level detection in the spin degree. Here, we developed a molecular interaction analysis method approaching single-molecule level based on relaxometry using the quantum sensor, nitrogen-vacancy (NV) center in diamond. Experiments utilized an optimized diamond surface functionalized with a polyethylenimine nanogel layer, achieving $\sim$10 nm average protein distance and mitigating interfacial steric hindrance. Then we measured the strong interaction between streptavidin and spin-labeled biotin complexes, as well as the weak interaction between bovine serum albumin and biotin complexes, at both the micrometer scale and nanoscale. For the micrometer-scale measurements using ensemble NV centers, we re-examined the often-neglected fast relaxation component and proposed a relaxation rate evaluation method, substantially enhancing the measurement sensitivity. Furthermore, we achieved nanoscale detection approaching single-molecule level using single NV centers. This methodology holds promise for applications in molecular screening, identification and kinetic studies at the single-molecule level, offering critical insights into molecular function and activity mechanisms.

quant-ph

Parallel accelerated electron paramagnetic resonance spectroscopy using diamond sensors

The nitrogen-vacancy (NV) center can serve as a magnetic sensor for electron paramagnetic resonance (EPR) measurements. Benefiting from its atomic size, the diamond chip can integrate a tremendous amount of NV centers to improve the magnetic-field sensitivity. However, EPR spectroscopy using NV ensembles is less efficient due to inhomogeneities in both sensors and targets. Spectral line broadening induced by ensemble averaging is even detrimental to spectroscopy. Here we show a kind of cross-relaxation EPR spectroscopy at zero field, where the sensor is tuned by an amplitude-modulated control field to match the target. The modulation makes detection robust to the sensor's inhomogeneity, while zero-field EPR is naturally robust to the target's inhomogeneity. We demonstrate an efficient EPR measurement on an ensemble of roughly 30000 NV centers. Our method shows the ability to not only acquire unambiguous EPR spectra of free radicals, but also monitor their spectroscopic dynamics in real time.

quant-ph

Generalized Denoising Diffusion Codebook Models (gDDCM): Tokenizing images using a pre-trained diffusion model

Denoising diffusion models have emerged as a dominant paradigm in image generation. Discretizing image data into tokens is a critical step for effectively integrating images with Transformer and other architectures. Although the Denoising Diffusion Codebook Models (DDCM) pioneered the use of pre-trained diffusion models for image tokenization, it strictly relies on the traditional discrete-time DDPM architecture. Consequently, it fails to adapt to modern continuous-time variants-such as Flow Matching and Consistency Models-and suffers from inefficient sampling in high-noise regions. To address these limitations, this paper proposes the Generalized Denoising Diffusion Codebook Models (gDDCM). We establish a unified theoretical framework and introduce a generic "De-noise and Back-trace" sampling strategy. By integrating a deterministic ODE denoising step with a residual-aligned noise injection step, our method resolves the challenge of adaptation. Furthermore, we introduce a backtracking parameter $p$ and significantly enhance tokenization ability. Extensive experiments on CIFAR10 and LSUN Bedroom datasets demonstrate that gDDCM achieves comprehensive compatibility with mainstream diffusion variants and significantly outperforms DDCM in terms of reconstruction quality and perceptual fidelity.

cs.CV

Double Yangians and quantum vertex algebras, I

For any symmetrizable generalized Cartan matrix $A$, we introduce an algebra $\widehat{\mathcal{DY}}(A)$, which is essentially the centrally extended double Yangian when $A$ is of finite type, and we give a new field (current) presentation of $\widehat{\mathcal{DY}}(A)$. Among the main results, for any $\ell\in \mathbb C$ we construct a universal vacuum $\widehat{\mathcal{DY}}(A)$-module $\mathcal{V}_A(\ell)$ of level $\ell$, prove that there exists a natural $\hbar$-adic weak quantum vertex algebra structure on $\mathcal{V}_A(\ell)$, and give an isomorphism between the category of restricted $\widehat{\mathcal{DY}}(A)$-modules of level $\ell$ and the category of $\mathcal{V}_A(\ell)$-modules.

math.QA

Double Yangians and lattice quantum vertex algebras

For any simply-laced GCM $A$, a $\mathbb C[[\hbar]]$-algebra $\widehat{\mathcal{DY}}(A)$ was introduced in [KL1], where it was proved that the universal vacuum $\widehat{\mathcal{DY}}(A)$-module ${\mathcal{V}}_A(\ell)$ for any fixed level $\ell$ is naturally an $\hbar$-adic weak quantum vertex algebra. Let $L$ be the root lattice of $\mathfrak g(A)$. As the main results of this paper, we construct an $\hbar$-adic quantum vertex algebra $V_L[[\hbar]]^{\eta}$ as a formal deformation of the lattice vertex algebra $V_L$ and show that every $V_L[[\hbar]]^{\eta}$-module is naturally a restricted $\widehat{\mathcal{DY}}(A)$-module of level one. For $A$ of finite type, we obtain a realization of $V_L[[\hbar]]^{\eta}$ as a quotient of the $\hbar$-adic weak quantum vertex algebra ${\mathcal{V}}_A(1)$, giving a characterization of $V_L[[\hbar]]^{\eta}$-modules as restricted $\widehat{\mathcal{DY}}(A)$-modules of level one.

math.QA

PhaseNAS: Language-Model Driven Architecture Search with Dynamic Phase Adaptation

Neural Architecture Search (NAS) is challenged by the trade-off between search space exploration and efficiency, especially for complex tasks. While recent LLM-based NAS methods have shown promise, they often suffer from static search strategies and ambiguous architecture representations. We propose PhaseNAS, an LLM-based NAS framework with dynamic phase transitions guided by real-time score thresholds and a structured architecture template language for consistent code generation. On the NAS-Bench-Macro benchmark, PhaseNAS consistently discovers architectures with higher accuracy and better rank. For image classification (CIFAR-10/100), PhaseNAS reduces search time by up to 86% while maintaining or improving accuracy. In object detection, it automatically produces YOLOv8 variants with higher mAP and lower resource cost. These results demonstrate that PhaseNAS enables efficient, adaptive, and generalizable NAS across diverse vision tasks.

cs.LG

LRR-Bench: Left, Right or Rotate? Vision-Language models Still Struggle With Spatial Understanding Tasks

Real-world applications, such as autonomous driving and humanoid robot manipulation, require precise spatial perception. However, it remains underexplored how Vision-Language Models (VLMs) recognize spatial relationships and perceive spatial movement. In this work, we introduce a spatial evaluation pipeline and construct a corresponding benchmark. Specifically, we categorize spatial understanding into two main types: absolute spatial understanding, which involves querying the absolute spatial position (e.g., left, right) of an object within an image, and 3D spatial understanding, which includes movement and rotation. Notably, our dataset is entirely synthetic, enabling the generation of test samples at a low cost while also preventing dataset contamination. We conduct experiments on multiple state-of-the-art VLMs and observe that there is significant room for improvement in their spatial understanding abilities. Explicitly, in our experiments, humans achieve near-perfect performance on all tasks, whereas current VLMs attain human-level performance only on the two simplest tasks. For the remaining tasks, the performance of VLMs is distinctly lower than that of humans. In fact, the best-performing Vision-Language Models even achieve near-zero scores on multiple tasks. The dataset and code are available on https://github.com/kong13661/LRR-Bench.

cs.CV

TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination Via Latent Truthful-Guided Pre-Intervention

Object Hallucination (OH) has been acknowledged as one of the major trustworthy challenges in Large Vision-Language Models (LVLMs). Recent advancements in Large Language Models (LLMs) indicate that internal states, such as hidden states, encode the "overall truthfulness" of generated responses. However, it remains under-explored how internal states in LVLMs function and whether they could serve as "per-token" hallucination indicators, which is essential for mitigating OH. In this paper, we first conduct an in-depth exploration of LVLM internal states with OH issues and discover that (1) LVLM internal states are high-specificity per-token indicators of hallucination behaviors. Moreover, (2) different LVLMs encode universal patterns of hallucinations in common latent subspaces, indicating that there exist "generic truthful directions" shared by various LVLMs. Based on these discoveries, we propose Truthful-Guided Pre-Intervention (TruthPrInt) that first learns the truthful direction of LVLM decoding and then applies truthful-guided inference-time intervention during LVLM decoding. We further propose TruthPrInt to enhance both cross-LVLM and cross-data hallucination detection transferability by constructing and aligning hallucination latent subspaces. We evaluate TruthPrInt in extensive experimental settings, including in-domain and out-of-domain scenarios, over popular LVLMs and OH benchmarks. Experimental results indicate that TruthPrInt significantly outperforms state-of-the-art methods. Codes will be available at https://github.com/jinhaoduan/TruthPrInt.

cs.CV

Quantum vertex algebra associated to quantum toroidal $\mathfrak{gl}_N$

In this paper, we associate the quantum toroidal algebra $\mathcal{E}_N$ of type $\mathfrak{gl}_N$ with quantum vertex algebra through equivariant $ϕ$-coordinated quasi modules. More precisely, for every $\ell\in \mathbb{C}$, by deforming the universal affine vertex algebra of $\mathfrak{sl}_\infty$, we construct an $\hbar$-adic quantum $\Z$-vertex algebra $V_{\widehat{\mathfrak{sl}}_{\infty},\hbar}(\ell,0)$. Then we prove that the category of restricted $\mathcal{E}_N$-modules of level $\ell$ is canonically isomorphic to that of equivariant $ϕ$-coordinated quasi $V_{\widehat{\mathfrak{sl}}_{\infty},\hbar}(\ell,0)$-modules.

math.QA

Twisted tensor products of quantum affine vertex algebras and coproducts

Let $\mathfrak g$ be a symmetrizable Kac-Moody Lie algebra, and let $V_{\hat{\mathfrak g},\hbar}^\ell$, $L_{\hat{\mathfrak g},\hbar}^\ell$ be the quantum affine vertex algebras constructed in [11]. For any complex numbers $\ell$ and $\ell'$, we present an $\hbar$-adic quantum vertex algebra homomorphism $Δ$ from $V_{\hat{\mathfrak g},\hbar}^{\ell+\ell'}$ to the twisted tensor product $\hbar$-adic quantum vertex algebra $V_{\hat{\mathfrak g},\hbar}^\ell\widehat\otimes V_{\hat{\mathfrak g},\hbar}^{\ell'}$. In addition, if both $\ell$ and $\ell'$ are positive integers, we show that $Δ$ induces an $\hbar$-adic quantum vertex algebra homomorphism from $L_{\hat{\mathfrak g},\hbar}^{\ell+\ell'}$ to the twisted tensor product $\hbar$-adic quantum vertex algebra $L_{\hat{\mathfrak g},\hbar}^\ell\widehat\otimes L_{\hat{\mathfrak g},\hbar}^{\ell'}$. Moreover, we prove the coassociativity of $Δ$.

math.QA

Representations of quantum lattice vertex algebras

Let $Q$ be a non-degenerated even lattice, let $V_Q$ be the lattice vertex algebra associated to $Q$, and let $V_Q^\eta$ be a quantum lattice vertex algebra. In this paper, we prove the equivalence between the category $V_Q$-modules and the category of $V_Q^\eta$-modules. As a consequence, we show that every $V_Q^\eta$-module is completely reducible, and the set of simple $V_Q^\eta$-modules are in one-to-one correspondence with the set of cosets of $Q$ in its dual lattice.

math.QA

Federated attention consistent learning models for prostate cancer diagnosis and Gleason grading

Artificial intelligence (AI) holds significant promise in transforming medical imaging, enhancing diagnostics, and refining treatment strategies. However, the reliance on extensive multicenter datasets for training AI models poses challenges due to privacy concerns. Federated learning provides a solution by facilitating collaborative model training across multiple centers without sharing raw data. This study introduces a federated attention-consistent learning (FACL) framework to address challenges associated with large-scale pathological images and data heterogeneity. FACL enhances model generalization by maximizing attention consistency between local clients and the server model. To ensure privacy and validate robustness, we incorporated differential privacy by introducing noise during parameter transfer. We assessed the effectiveness of FACL in cancer diagnosis and Gleason grading tasks using 19,461 whole-slide images of prostate cancer from multiple centers. In the diagnosis task, FACL achieved an area under the curve (AUC) of 0.9718, outperforming seven centers with an average AUC of 0.9499 when categories are relatively balanced. For the Gleason grading task, FACL attained a Kappa score of 0.8463, surpassing the average Kappa score of 0.7379 from six centers. In conclusion, FACL offers a robust, accurate, and cost-effective AI training model for prostate cancer pathology while maintaining effective data safeguards.

cs.CV

ACT-Diffusion: Efficient Adversarial Consistency Training for One-step Diffusion Models

Though diffusion models excel in image generation, their step-by-step denoising leads to slow generation speeds. Consistency training addresses this issue with single-step sampling but often produces lower-quality generations and requires high training costs. In this paper, we show that optimizing consistency training loss minimizes the Wasserstein distance between target and generated distributions. As timestep increases, the upper bound accumulates previous consistency training losses. Therefore, larger batch sizes are needed to reduce both current and accumulated losses. We propose Adversarial Consistency Training (ACT), which directly minimizes the Jensen-Shannon (JS) divergence between distributions at each timestep using a discriminator. Theoretically, ACT enhances generation quality, and convergence. By incorporating a discriminator into the consistency training framework, our method achieves improved FID scores on CIFAR10 and ImageNet 64$\times$64 and LSUN Cat 256$\times$256 datasets, retains zero-shot image inpainting capabilities, and uses less than $1/6$ of the original batch size and fewer than $1/2$ of the model parameters and training steps compared to the baseline method, this leads to a substantial reduction in resource consumption. Our code is available:https://github.com/kong13661/ACT

cs.CV

An Efficient Membership Inference Attack for the Diffusion Model by Proximal Initialization

Recently, diffusion models have achieved remarkable success in generating tasks, including image and audio generation. However, like other generative models, diffusion models are prone to privacy issues. In this paper, we propose an efficient query-based membership inference attack (MIA), namely Proximal Initialization Attack (PIA), which utilizes groundtruth trajectory obtained by $ε$ initialized in $t=0$ and predicted point to infer memberships. Experimental results indicate that the proposed method can achieve competitive performance with only two queries on both discrete-time and continuous-time diffusion models. Moreover, previous works on the privacy of diffusion models have focused on vision tasks without considering audio tasks. Therefore, we also explore the robustness of diffusion models to MIA in the text-to-speech (TTS) task, which is an audio generation task. To the best of our knowledge, this work is the first to study the robustness of diffusion models to MIA in the TTS task. Experimental results indicate that models with mel-spectrogram (image-like) output are vulnerable to MIA, while models with audio output are relatively robust to MIA. {Code is available at \url{https://github.com/kong13661/PIA}}.

cs.SD