SearcharxivSearch

arXiv subjects

Long Tian

Publications and source records attributed to Long Tian.

At least 19 recordsLinked to original sources

ConceptADapt: Concept-guided Adaptive Feature Reconstruction with Dynamic Attention for Few-Shot Industrial Anomaly Detection

Few-shot industrial anomaly detection (FS-IAD) focuses on detecting and localizing visual defects in industrial inspection during the cold-start phase, where only a limited number of normal training samples are available per category. Recent advances in this field predominantly leverage visual features from foundation-model and have achieved promising performance. Despite the strong representational power of foundation-model features, the model generalization remains fragile due to the extreme scarcity of normal training data.To address this pivotal issue, we propose ConceptADapt, a concept-guided adaptive feature reconstruction model with dynamic attention. Specifically, our model pre-learns a set of fixed normal concepts from the limited support features and leverages them to mine relationships with query features, thereby recalibrating their statistics for improved anomaly detection at test time. To mitigate the prevalent feature shortcut problem, which is particularly severe under low-data regimes, we further develop a dynamic attention mechanism integrated with sparse autoencoders to learn robust normal concepts during training. Moreover, to enable fast adaptation during inference, our model remains lightweight by incorporating LoRA into the attention module, which introduces only minimal updating parameters.Extensive experiments on three widely adopted FS-IAD benchmarks, including MVTec-AD, VisA, and MPDD, demonstrate that our model consistently outperforms state-of-the-art (SOTA) approaches across both detection and localization tasks, achieving significant improvements under various shot settings.

cs.CV

$d$-spacing distributions as a probe of nematoelastic response in iron-based superconductors

Electronic nematicity in iron-based superconductors (FeSCs) couples bilinearly to orthorhombic strain, allowing nematic correlations to appear in the lattice response. Here we use neutron Larmor diffraction to measure the temperature-dependent distribution of relative $d$ spacings in electron-doped Ba(Fe$_{1-x}$Co$_x$)$_2$As$_2$, hole-doped Ba$_{0.83}$K$_{0.17}$Fe$_2$As$_2$, FeSe, and Fe$_{1.07}$Te. In Ba(Fe$_{1-x}$Co$_x$)$_2$As$_2$ crystals without intentionally applied uniaxial stress, the in-plane distribution width, $\varepsilon_{\rm FWHM}$, increases on cooling in the tetragonal phase and can be described phenomenologically by a Curie--Weiss-like form. The fitted scale $T^*$ decreases with Co doping and evolves similarly to the nematic phase diagram inferred from elastoresistance, although the two experiments probe different response functions. Related broadening in Ba$_{0.83}$K$_{0.17}$Fe$_2$As$_2$ and FeSe supports extending this interpretation beyond electron-doped BaFe$_2$As$_2$. By contrast, Fe$_{1.07}$Te shows no extended Curie--Weiss-like regime without applied stress, whereas uniaxial pressure produces a strongly anisotropic broadening that can contain contributions from both the field-biased lattice response and inhomogeneous loading. A mean-field model with bilinear nematoelastic coupling and spatially varying symmetry-breaking stress explains the Curie--Weiss-like broadening in terms of the renormalized orthorhombic compliance. Neutron Larmor diffraction therefore provides a bulk-sensitive probe of nematic-related lattice broadening that complements electronic and elastic measurements.

cond-mat.supr-con

One-Step Diffusion with Inverse Residual Fields for Unsupervised Industrial Anomaly Detection

Diffusion models have achieved outstanding performance in unsupervised industrial anomaly detection (uIAD) by learning a manifold of normal data under the common assumption that off-manifold anomalies are harder to generate, resulting in larger reconstruction errors in data space or lower probability densities in the tractable latent space. However, their iterative denoising and noising nature leads to slow inference. In this paper, we propose OSD-IRF, a novel one-step diffusion with inverse residual fields, to address this limitation for uIAD task. We first train a deep diffusion probabilistic model (DDPM) on normal data without any conditioning. Then, for a test sample, we predict its inverse residual fields (IRF) based on the noise estimated by the well-trained parametric noise function of the DDPM. Finally, uIAD is performed by evaluating the probability density of the IRF under a Gaussian distribution and comparing it with a threshold. Our key observation is that anomalies become distinguishable in this IRF space, a finding that has seldom been reported in prior works. Moreover, OSD-IRF requires only single step diffusion for uIAD, thanks to the property that IRF holds for any neighboring time step in the denoising process. Extensive experiments on three widely used uIAD benchmarks show that our model achieves SOTA or competitive performance across six metrics, along with roughly a 2X inference speedup without distillation.

cs.CV

Enhancing few-shot time series forecasting with LLM-guided diffusion

Time series forecasting in specialized domains is often constrained by limited data availability, where conventional models typically require large-scale datasets to effectively capture underlying temporal dynamics. To tackle this few-shot challenge, we propose LTSM-DIFF (Large-scale Temporal Sequential Memory with Diffusion), a novel learning framework that integrates the expressive power of large language models with the generative capability of diffusion models. Specifically, the LTSM module is fine-tuned and employed as a temporal memory mechanism, extracting rich sequential representations even under data-scarce conditions. These representations are then utilized as conditional guidance for a joint probability diffusion process, enabling refined modeling of complex temporal patterns. This design allows knowledge transfer from the language domain to time series tasks, substantially enhancing both generalization and robustness. Extensive experiments across diverse benchmarks demonstrate that LTSM-DIFF consistently achieves state-of-the-art performance in data-rich scenarios, while also delivering significant improvements in few-shot forecasting. Our work establishes a new paradigm for time series analysis under data scarcity.

cs.LG

Continuous variable quantum communication with 40 pairs of entangled sideband

Constructing large-scale quantum resources is an important foundation for further improving the efficiency and scalability of quantum communication. Here, we present an efficient extraction and stable control scheme of 40 pairs of entangled sideband modes from the squeezed light by specially designing optical parametric oscillator. Utilizing the low-loss optical frequency comb control technology and the local cross-correlation algorithm, we model and manage the efficient separation process of the entangled sidebands modes facilitated by the optical filtering cavities, a maximum entanglement level of 6.5 dB is achieved. The feasibility of large-capacity quantum dense coding based on these entangled sideband modes is proved experimentally, which is of great significance for optimizing the utilization of quantum resources, thereby contributing to the advancement of large-capacity quantum communication networks and enabling the realization of more secure and efficient quantum communication systems.

quant-ph

Reservoir-engineered squeezed lasing through the parametric coupling

We report the first experimental demonstration of squeezed lasing in a reservoir-engineered optical parametric oscillator (OPO). The OPO provides a basis of squeezed states and parametric amplification in lasing emission, whose vacuum reservoir is coupled to a squeezed vacuum generated by a second OPO. With a precisely controlled squeezing angle and strong squeezing injection, the parametric interaction in the first OPO is exponentially enhanced. It successfully circumvents the decoherence in the system, and eliminates the undesired noise of spontaneous photon emission in the OPO. As a result, the amplified parametric process simultaneously reserves the coherence and quantum properties in the first OPO, and yields a -6.1 dB squeezed laser in optical domain with a narrow linewidth and high brightness. Our work sheds light on potential applications of squeezed lasing in quantum metrology and quantum optics.

quant-ph

Quantum-enhanced laser phase noise filter

Quantum noise is the fundamental limit of laser phase noise filter. We cannot realize the effective quantum-enhanced phase noise suppression through simply utilizing amplitude noise suppression scheme. Here, we present the first experimental demonstration of a quantum-enhanced laser phase noise filter, achieved by employing a noise ellipse rotation phase noise readout technique combined with an excess amplitude noise suppression scheme. We address the primary limitations in the extracting of laser phase noise, and make the quantum enhancement via squeezed vacuum injection feasible. A maximum of 5 dB quantum-enhanced phase noise suppression is realized across the Fourier frequencies from 5 kHz to 60 kHz. The demonstration unlocks the application of squeezed vacuum state in laser phase noise suppression.

quant-ph

FastRef:Fast Prototype Refinement for Few-Shot Industrial Anomaly Detection

Few-shot industrial anomaly detection (FS-IAD) presents a critical challenge for practical automated inspection systems operating in data-scarce environments. While existing approaches predominantly focus on deriving prototypes from limited normal samples, they typically neglect to systematically incorporate query image statistics to enhance prototype representativeness. To address this issue, we propose FastRef, a novel and efficient prototype refinement framework for FS-IAD. Our method operates through an iterative two-stage process: (1) characteristic transfer from query features to prototypes via an optimizable transformation matrix, and (2) anomaly suppression through prototype alignment. The characteristic transfer is achieved through linear reconstruction of query features from prototypes, while the anomaly suppression addresses a key observation in FS-IAD that unlike conventional IAD with abundant normal prototypes, the limited-sample setting makes anomaly reconstruction more probable. Therefore, we employ optimal transport (OT) for non-Gaussian sampled features to measure and minimize the gap between prototypes and their refined counterparts for anomaly suppression. For comprehensive evaluation, we integrate FastRef with three competitive prototype-based FS-IAD methods: PatchCore, FastRecon, WinCLIP, and AnomalyDINO. Extensive experiments across four benchmark datasets of MVTec, ViSA, MPDD and RealIAD demonstrate both the effectiveness and computational efficiency of our approach under 1/2/4-shots.

cs.CV

Meta-SurDiff: Classification Diffusion Model Optimized by Meta Learning is Reliable for Online Surgical Phase Recognition

Online surgical phase recognition has drawn great attention most recently due to its potential downstream applications closely related to human life and health. Despite deep models have made significant advances in capturing the discriminative long-term dependency of surgical videos to achieve improved recognition, they rarely account for exploring and modeling the uncertainty in surgical videos, which should be crucial for reliable online surgical phase recognition. We categorize the sources of uncertainty into two types, frame ambiguity in videos and unbalanced distribution among surgical phases, which are inevitable in surgical videos. To address this pivot issue, we introduce a meta-learning-optimized classification diffusion model (Meta-SurDiff), to take full advantage of the deep generative model and meta-learning in achieving precise frame-level distribution estimation for reliable online surgical phase recognition. For coarse recognition caused by ambiguous video frames, we employ a classification diffusion model to assess the confidence of recognition results at a finer-grained frame-level instance. For coarse recognition caused by unbalanced phase distribution, we use a meta-learning based objective to learn the diffusion model, thus enhancing the robustness of classification boundaries for different surgical phases.We establish effectiveness of Meta-SurDiff in online surgical phase recognition through extensive experiments on five widely used datasets using more than four practical metrics. The datasets include Cholec80, AutoLaparo, M2Cai16, OphNet, and NurViD, where OphNet comes from ophthalmic surgeries, NurViD is the daily care dataset, while the others come from laparoscopic surgeries. We will release the code upon acceptance.

cs.CV

Low-Rank Adaptation of Pre-Trained Stable Diffusion for Rigid-Body Target ISAR Imaging

Traditional range-instantaneous Doppler (RID) methods for rigid-body target imaging often suffer from low resolution due to the limitations of time-frequency analysis (TFA). To address this challenge, our primary focus is on obtaining high resolution time-frequency representations (TFRs) from their low resolution counterparts. Recognizing that the curve features of TFRs are a specific type of texture feature, we argue that pre trained generative models such as Stable Diffusion (SD) are well suited for enhancing TFRs, thanks to their powerful capability in capturing texture representations. Building on this insight, we propose a novel inverse synthetic aperture radar (ISAR) imaging method for rigid-body targets, leveraging the low-rank adaptation (LoRA) of a pre-trained SD model. Our approach adopts the basic structure and pre-trained parameters of SD Turbo while incorporating additional linear operations for LoRA and adversarial training to achieve super-resolution and noise suppression. Then we integrate LoRA-SD into the RID-based ISAR imaging, enabling sharply focused and denoised imaging with super-resolution capabilities. We evaluate our method using both simulated and real radar data. The experimental results demonstrate the superiority of our approach in frequency es timation and ISAR imaging compared to traditional methods. Notably, the generalization capability is verified by training on simulated radar data and testing on measured radar data.

cs.CV

A Spatial-temporal Deep Probabilistic Diffusion Model for Reliable Hail Nowcasting with Radar Echo Extrapolation

Hail nowcasting is a considerable contributor to meteorological disasters and there is a great need to mitigate its socioeconomic effects through precise forecast that has high resolution, long lead times and local details with large landscapes. Existing medium-range weather forecasting methods primarily rely on changes in upper air currents and cloud layers to predict precipitation events, such as heavy rainfall, which are unsuitable for hail nowcasting since it is mainly caused by low-altitude local strong convection associated with terrains. Additionally, radar captures the status of low cloud layers, such as water vapor, droplets, and ice crystals, providing rich signals suitable for hail nowcasting. To this end, we introduce a Spatial-Temporal gEnerAtive Model called SteamCast for hail nowcasting with radar echo extrapolation, it is a deep probabilistic diffusion model based on spatial-temporal representations including radar echoes as well as their position/time embeddings, which we trained on historical reanalysis archive from Yan'an Meteorological Bureau in China, where the crop yield like apple suffers greatly from hail damage. Considering the short-term nature of hail, SteamCast provides 30-minute nowcasts at 6-minute intervals for a single radar reflectivity variable, across 9 different vertical angles, on a latitude-longitude grid with approximately 1 km * 1 km resolution per pixel in Yan'an City, China. By successfully fusing the spatial-temporal features of radar echoes, SteamCast delivers competitive, and in some cases superior, results compared to other deep learning-based models such as PredRNN and VMRNN.

cs.LG

Laser intensity noise suppression for space-borne gravitational wave mission

Laser intensity noise is a main limitation of measurement and sensing mission represented by gravitational wave detection. We develop a noise decomposition model and design the core elements of the feedback loop independently based on the analysis results. We construct a fiber amplifier system with ultra-low intensity noise in the 0.1 mHz-1 Hz frequency band by the employment of an optoelectronic feedback loop that is specially designed. The study provides experimental basis and technologies for precise measurement and sensing system at ultra-low frequency.

physics.optics

Polarization-Analyzed Small-Angle Neutron Scattering with an $\textit{in-situ}$ $^{3}$He neutron spin filter at the China Spallation Neutron Source

Polarization-analyzed small-angle neutron scattering (PASANS) is an advanced technique that enables the selective investigation of magnetic scattering phenomena in magnetic materials and distinguishes coherent scattering obscured by incoherent backgrounds, making it particularly valuable for cutting-edge research. The successful implementation of PASANS in China was achieved for the first time at the newly commissioned Very Small Angle Neutron Scattering (VSANS) instrument at the China Spallation Neutron Source (CSNS). This technique employs a combination of a double-V cavity supermirror polarizer and a radio frequency (RF) neutron spin flipper to manipulate the polarization of the incident neutrons. The scattered neutron polarization is stably analyzed by a specially designed $\textit{in-situ}$ optical pumping $^{3}$He neutron spin filter, which covers a spatially symmetric scattering angle coverage of about 4.8 $^{\circ}$. A comprehensive PASANS data reduction method, aimed at pulsed neutron beams, has been established and validated with a silver behenate powder sample, indicating a maximum momentum transfer coverage of approximately 0.25 {\AA} $^{-1}$.

physics.ins-det

Generation of squeezed vacuum state in the millihertz frequency band

The detection of gravitational waves has ushered in a new era of observing the universe. Quantum resource advantages offer significant enhancements to the sensitivity of gravitational wave observatories. While squeezed states for ground-based gravitational wave detection have received marked attention, the generation of squeezed states suitable for mid-to-low-frequency detection has remained unexplored. To address the gap in squeezed state optical fields at ultra-low frequencies, we report on the first direct observation of a squeezed vacuum field until Fourier frequency of 4 millihertz with the quantum noise reduction of up to 8 dB, by the employment of a multiple noise suppression scheme. Our work provides quantum resources for future gravitational wave observatories, facilitating the development of quantum precision measurement.

physics.optics

Cradle: Empowering Foundation Agents Towards General Computer Control

Despite the success in specific scenarios, existing foundation agents still struggle to generalize across various virtual scenarios, mainly due to the dramatically different encapsulations of environments with manually designed observation and action spaces. To handle this issue, we propose the General Computer Control (GCC) setting to restrict foundation agents to interact with software through the most unified and standardized interface, i.e., using screenshots as input and keyboard and mouse actions as output. We introduce Cradle, a modular and flexible LMM-powered framework, as a preliminary attempt towards GCC. Enhanced by six key modules, Cradle can understand input screenshots and output executable code for low-level keyboard and mouse control after high-level planning, so that Cradle can interact with any software and complete long-horizon complex tasks without relying on any built-in APIs. Experimental results show that Cradle exhibits remarkable generalizability and impressive performance across four previously unexplored commercial video games, five software applications, and a comprehensive benchmark, OSWorld. Cradle is the first to enable foundation agents to follow the main storyline and complete 40-minute-long real missions in the complex AAA game Red Dead Redemption 2 (RDR2). Cradle can also create a city of a thousand people in Cities: Skylines, farm and harvest parsnips in Stardew Valley, and trade and bargain with a maximal weekly total profit of 87% in Dealer's Life 2. Cradle can not only operate daily software, like Chrome, Outlook, and Feishu, but also edit images and videos using Meitu and CapCut. Cradle greatly extends the reach of foundation agents by enabling the easy conversion of any software, especially complex games, into benchmarks to evaluate agents' various abilities and facilitate further data collection, thus paving the way for generalist agents.

cs.AI

Measure upper bounds of nodal sets of solutions to Dirichlet problem of Schr\"{o}dinger equations

In this paper, we focus on estimating measure upper bounds of nodal sets of solutions to the following boundary value problem \begin{equation*} \left\{ \begin{array}{lll} \Delta u+Vu=0\quad \mbox{in}\ \Omega,\\[2mm] u=0\quad \mbox{on}\ \partial\Omega, \end{array}\right. \end{equation*} where $V\in W^{1,\infty}(\Omega)$ is a potential function, and $\Omega \subset \mathbb{R}^n$ ($n \geq 2$) is a bounded domain whose boundary is of class $C^{1,\alpha}$ for any $0<\alpha<1$. By developing a delicate dividing iteration procedure, we show that upper bound of the $(n-1)$-dimensional Hausdorff measure of the nodal set of $u$ in $\Omega$ is $$C\Big(1+\log\left(\|\nabla V\|_{L^{\infty}(\Omega)}+1\right)\Big)\cdot\left(\|V\|_{L^{\infty}(\Omega)}^{\frac{1}{2}}+\|\nabla V\|_{L^{\infty}(\Omega)}^{\frac{1}{2}}+1\right),$$ provided $V$ is analytic, here $C$ is a positive constant depending only on $n$ and $\Omega$. In particular, if $\|\nabla V\|_{L^{\infty}(\Omega)}$ is small, the upper bound for the measure of the nodal set of $u$ is $C\left(\|V\|^{\frac{1}{2}}_{L^{\infty}(\Omega)}+1\right)$, which is sharp in the sense of a famous conjecture of Yau.

math.AP

Hierarchical Vector Quantized Transformer for Multi-class Unsupervised Anomaly Detection

Unsupervised image Anomaly Detection (UAD) aims to learn robust and discriminative representations of normal samples. While separate solutions per class endow expensive computation and limited generalizability, this paper focuses on building a unified framework for multiple classes. Under such a challenging setting, popular reconstruction-based networks with continuous latent representation assumption always suffer from the "identical shortcut" issue, where both normal and abnormal samples can be well recovered and difficult to distinguish. To address this pivotal issue, we propose a hierarchical vector quantized prototype-oriented Transformer under a probabilistic framework. First, instead of learning the continuous representations, we preserve the typical normal patterns as discrete iconic prototypes, and confirm the importance of Vector Quantization in preventing the model from falling into the shortcut. The vector quantized iconic prototype is integrated into the Transformer for reconstruction, such that the abnormal data point is flipped to a normal data point.Second, we investigate an exquisite hierarchical framework to relieve the codebook collapse issue and replenish frail normal patterns. Third, a prototype-oriented optimal transport method is proposed to better regulate the prototypes and hierarchically evaluate the abnormal score. By evaluating on MVTec-AD and VisA datasets, our model surpasses the state-of-the-art alternatives and possesses good interpretability. The code is available at https://github.com/RuiyingLu/HVQ-Trans.

cs.CV

Quantitative unique continuation property for solutions to a bi-Laplacian equation with a potential

In this paper, we focus on the quantitative unique continuation property of solutions to \begin{equation*} \Delta^2u=Vu, \end{equation*} where $V\in W^{1,\infty}$. We show that the maximal vanishing order of the solutions is not large than \begin{equation} C\left(\|V\|^{\frac{1}{4}}_{L^{\infty}}+\|\nabla V\|_{L^{\infty}}+1\right). \end{equation} Our key argument is to lift the original equation to that with a positive potential, then decompose the resulted fourth-order equation into a special system of two second-order equations. Based on the special system, we define a variant frequency function with weights and derive its almost monotonicity to establishing some doubling inequalities with explicit dependence on the Sobolev norm of the potential function.

math.AP