SearcharxivSearch

arXiv subjects

Yifan Tang

Publications and source records attributed to Yifan Tang.

At least 19 recordsLinked to original sources

Witness expansion: A unified framework for analytical and measurable mixed-state resource detection

Quantum information science aims to harness different kinds of quantum resources to accomplish specific information-processing tasks. These resources also play an increasingly important role in addressing fundamental questions concerning quantum phases and dynamics. Therefore, developing powerful and practical methods for identifying and detecting quantum resources is of great significance, with applications ranging from benchmarking quantum devices to understanding the fundamental structure of quantum theory. In this work, we propose witness expansion, a unified framework for constructing nonlinear criteria for detecting quantum resources that are associated with a well-defined group of free unitaries. These criteria apply to both pure and mixed quantum states and are based on polynomial functions of the target state, which can be estimated experimentally using multiple copies of the state and evaluated analytically in certain physical models. We show how several well-known resource-detection quantities naturally emerge from our framework, including the $l_2$ norm of coherence, partial-transpose moments for entanglement, stabilizer entropy for nonstabilizerness (quantum magic), and fermionic antiflatness for fermionic non-Gaussianity. Beyond recovering these existing structures, our framework also yields new criteria for detecting qubit and qudit magic states, substantially enhancing witness-based detection capabilities. In addition, it gives, to the best of our knowledge, the first analytical criterion for detecting mixed-state fermionic non-Gaussianity with respect to the convex hull of pure fermionic Gaussian states that remains nontrivial for arbitrary numbers of qubits, demonstrating the broad applicability and conceptual unifying power of the framework.

quant-ph

Topological Signatures of Grokking

We study the grokking phenomenon through the lens of topology. Using persistent homology on point clouds derived from the embedding matrices of a range of models trained on modular arithmetic with varying primes, we identify a clear and consistent topological signature of grokking: a sharp increase in both the maximum and total persistence of first homology ($H_1$). Persistence diagrams reveal the emergence of a dominant long-lived topological feature together with increasingly structured secondary features, reflecting the underlying cyclic structure of the task. Compared to existing spectral and geometric diagnostics -- specifically, Fourier analysis and local intrinsic dimension -- persistent homology provides a unified geometric and topological characterization of representation learning, capturing both local and global multi-scale structure. Ablations across data regimes and control settings show that these topological transitions are tied to generalization rather than memorization. Our results suggest that persistent homology offers a principled and interpretable framework for analyzing how neural networks internalize latent structure during training.

cs.LG

Measurement-induced non-commutativity in adaptive fermionic linear optics

Fermionic linear optics (FLO) with Gaussian resources is efficiently classically simulable. We show that this is no longer the case for such quantum circuits for fermions with internal degrees of freedom, equipped with mid-circuit number monitoring and classical feedforward. In our architecture, the measurement record routes the selected blocks into a fixed-order Bell-fusion pairing geometry. On the level of classical description, this implies realizing a situation in which the permutation sum no longer collapses to a single determinant or Pfaffian. Each post-selected branch expands as a signed sum of path-ordered products of typically non-commuting dressed blocks, and branch amplitudes are matrix elements of the resulting non-commutative trace polynomials. Numerically, we observe Porter-Thomas statistics as the output distribution and a rapid growth of the minimal order-respecting matrix product operator bond dimension. These results thus establish mid-circuit measurement-induced non-commutativity as a route to sampling hardness for noninteracting fermions under reasonable complexity assumptions, without introducing coherent two-body interactions into the FLO evolution.

quant-ph

TabPFN for Zero-shot Parametric Engineering Design Generation

Deep generative models for engineering design often require substantial computational cost, large training datasets, and extensive retraining when design requirements or datasets change, limiting their applicability in real-world engineering design workflow. In this work, we propose a zero-shot generation framework for parametric engineering design based on TabPFN, enabling conditional design generation using only a limited number of reference samples and without any task-specific model training or fine-tuning. The proposed method generates design parameters sequentially conditioned on target performance indicators, providing a flexible alternative to conventional generative models. The effectiveness of the proposed approach is evaluated on three engineering design datasets, i.e., ship hull design, BlendedNet aircraft, and UIUC airfoil. Experimental results demonstrate that the proposed method achieves competitive diversity across highly structured parametric design spaces, remains robust to variations in sampling, resolution and parameter dimensionality of geometry generation, and achieves a low performance error (e.g., less than 2% in generated ship hull designs' performance). Compared with diffusion-based generative models, the proposed framework significantly reduces computational overhead and data requirements while preserving reliable generation performance. These results highlight the potential of zero-shot, data-efficient generation as a practical and efficient tool for engineering design, enabling rapid deployment, flexible adaptation to new design settings, and ease of integration into real-world engineering workflows.

cs.LG

RePaint-Enhanced Conditional Diffusion Model for Parametric Engineering Designs under Performance and Parameter Constraints

This paper presents a RePaint-enhanced framework that integrates a pre-trained performance-guided denoising diffusion probabilistic model (DDPM) for performance- and parameter-constraint engineering design generation. The proposed method enables the generation of missing design components based on a partial reference design while satisfying performance constraints, without retraining the underlying model. By applying mask-based resampling during inference process, RePaint allows efficient and controllable repainting of partial designs under both performance and parameter constraints, which is not supported by conventional DDPM-base methods. The framework is evaluated on two representative design problems, parametric ship hull design and airfoil design, demonstrating its ability to generate novel designs with expected performance based on a partial reference design. Results show that the method achieves accuracy comparable to or better than pre-trained models while enabling controlled novelty through fixing partial designs. Overall, the proposed approach provides an efficient, training-free solution for parameter-constraint-aware generative design in engineering applications.

cs.LG

Adaptation and Fine-tuning with TabPFN for Travelling Salesman Problem

Tabular Prior-Data Fitted Network (TabPFN) is a foundation model designed for small to medium-sized tabular data, which has attracted much attention recently. This paper investigates the application of TabPFN in Combinatorial Optimization (CO) problems. The aim is to lessen challenges in time and data-intensive training requirements often observed in using traditional methods including exact and heuristic algorithms, Machine Learning (ML)-based models, to solve CO problems. Proposing possibly the first ever application of TabPFN for such a purpose, we adapt and fine-tune the TabPFN model to solve the Travelling Salesman Problem (TSP), one of the most well-known CO problems. Specifically, we adopt the node-based approach and the node-predicting adaptation strategy to construct the entire TSP route. Our evaluation with varying instance sizes confirms that TabPFN requires minimal training, adapts to TSP using a single sample, performs better generalization across varying TSP instance sizes, and reduces performance degradation. Furthermore, the training process with adaptation and fine-tuning is completed within minutes. The methodology leads to strong solution quality even without post-processing and achieves performance comparable to other models with post-processing refinement. Our findings suggest that the TabPFN model is a promising approach to solve structured and CO problems efficiently under training resource constraints and rapid deployment requirements.

cs.LG

VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model

Vision-Language-Action (VLA) models typically bridge the gap between perceptual and action spaces by pre-training a large-scale Vision-Language Model (VLM) on robotic data. While this approach greatly enhances performance, it also incurs significant training costs. In this paper, we investigate how to effectively bridge vision-language (VL) representations to action (A). We introduce VLA-Adapter, a novel paradigm designed to reduce the reliance of VLA models on large-scale VLMs and extensive pre-training. To this end, we first systematically analyze the effectiveness of various VL conditions and present key findings on which conditions are essential for bridging perception and action spaces. Based on these insights, we propose a lightweight Policy module with Bridge Attention, which autonomously injects the optimal condition into the action space. In this way, our method achieves high performance using only a 0.5B-parameter backbone, without any robotic data pre-training. Extensive experiments on both simulated and real-world robotic benchmarks demonstrate that VLA-Adapter not only achieves state-of-the-art level performance, but also offers the fast inference speed reported to date. Furthermore, thanks to the proposed advanced bridging paradigm, VLA-Adapter enables the training of a powerful VLA model in just 8 hours on a single consumer-grade GPU, greatly lowering the barrier to deploying the VLA model. Project page: https://vla-adapter.github.io/.

cs.RO

Optimal randomized measurements for a family of non-linear quantum properties

Quantum learning encounters fundamental challenges when estimating non-linear properties, owing to the inherent linearity of quantum mechanics. Although recent advances in single-copy randomized measurement protocols have achieved optimal sample complexity for specific tasks like state purity estimation, generalizing these protocols to estimate broader classes of non-linear properties without sacrificing optimality remains an open problem. In this work, we introduce the observable-driven randomized measurement (ORM) protocol enabling the estimation of ${\rm Tr}(O\rho^2)$ for an arbitrary observable $O$ -- an essential quantity in quantum computing and many-body physics. We establish an upper bound for ORM's sample complexity and show its optimality for observables with a large trace-norm, including Pauli and local observables, closing a gap in the literature. For these observables, ORM admits an efficient implementation with Clifford circuits. Numerical experiments validate that ORM requires substantially fewer state samples to achieve the same precision compared to classical shadows. Additionally, we introduce a braiding randomized measurement protocol for multiple low-rank non-linear observables, reducing circuit complexities in practical applications.

quant-ph

Learning Attribute-aware Representations for Few-shot Scene Text Segmentation

Supervised scene text segmentation has achieved notable progress in recent years. However, its development is largely constrained by the scarcity of high-quality datasets and the high cost of pixel-level annotations. To address this limitation, we explore few-shot learning for text segmentation and propose TSAL, an attribute-aware few-shot framework that leverages a pre-trained CLIP model to learn transferable text attributes for segmentation. Our framework comprises two complementary branches: I) a Visual-Guided Branch that extracts semantic and textural features for foreground text and background regions, respectively, and II) an Adaptive Prompt-Guided Branch that employs learnable prompt templates to capture diverse text attributes with minimal data dependence. To effectively align textual attributes with visual representations, we further introduce an Adaptive Feature Alignment~(AFA) module, which aligns learnable attribute tokens with visual features and prompt prototypes, enabling the model to capture both general and distinctive textual characteristics. As a result, TSAL can accurately segment text regions using only a few annotated samples. Extensive experiments demonstrate that our method achieves state-of-the-art performance across several public text segmentation benchmarks under few-shot settings and exhibits strong generalization to text-related tasks.

cs.CV

Deep quantum Monte Carlo approach for polaritonic chemistry

Recent years have witnessed a surge of experimental and theoretical interest in controlling the properties of matter, such as its chemical reactivity, by confining it in optical cavities, where the enhancement of the light-matter coupling strength leads to the creation of hybrid light-matter states known as polaritons. However, ab initio calculations that account for the quantum nature of both the electromagnetic field and matter are challenging and have only started to be developed in recent years. We introduce a deep learning variational quantum Monte Carlo method to solve the electronic and photonic Schr\"odinger equation of molecules trapped in optical cavities. We extend typical electronic neural network wavefunction ansatzes to describe joint fermionic and bosonic systems, i.e. electron-photon systems, in a quantum Monte Carlo framework. We apply our method to hydrogen molecules in a cavity, computing both ground and excited states. We assess their energy, dipole moment, charge density shift due to the cavity, the state of the photonic field, and the entanglement developed between the electrons and photons. When possible, we compare our results with more conventional quantum chemistry methods proposed in the literature, finding good qualitative agreement, thus extending the range of scientific problems that can be tackled using machine learning techniques.

physics.chem-ph

Capturing Lifecycle System Degradation in Digital Twin Model Updating

Digital twin (DT) has emerged as a powerful tool to facilitate monitoring, control, and other decision-making tasks in real-world engineering systems. Online update methods have been proposed to update DT models. Considering the degradation behavior in the system lifecycle, these methods fail to enable DT models to predict the system responses affected by the system degradation over time. To alleviate this problem, degradation models of measurable parameters have been integrated into DT construction. However, identifying the degradation parameters relies on prior knowledge of the system and expensive experiments. To mitigate those limitations, this paper proposes a lifelong update method for DT models to capture the effects of system degradation on system responses without any prior knowledge and expensive offline experiments on the system. The core idea in the work is to represent the system degradation during the lifecycle as the dynamic changes of DT configurations (i.e., model parameters with a fixed model structure) at all degradation stages. During the lifelong update process, an Autoencoder is adopted to reconstruct the model parameters of all hidden layers simultaneously, so that the latent features taking into account the dependencies among hidden layers are obtained for each degradation stage. The dynamic behavior of latent features among successive degradation stages is then captured by a long short-term memory model, which enables prediction of the latent feature at any unseen stage. Based on the predicted latent features, the model configuration at future degradation stage is reconstructed to determine the new DT model, which predicts the system responses affected by the degradation at the same stage. The test results on two engineering datasets demonstrate that the proposed update method could capture effects of system degradation on system responses during the lifecycle.

cs.CE

Hybrid Metaheuristic Vehicle Routing Problem for Security Dispatch Operations

This paper investigates the optimization of the Vehicle Routing Problem for Security Dispatch (VRPSD). VRPSD focuses on security and patrolling applications which involve challenging constraints including precise timing and strict time windows. We propose three algorithms based on different metaheuristics, which are Adaptive Large Neighborhood Search (ALNS), Tabu Search (TS), and Threshold Accepting (TA). The first algorithm combines single-phase ALNS with TA, the second employs a multiphase ALNS with TA, and the third integrates multiphase ALNS, TS, and TA. Experiments are conducted on an instance comprising 251 customer requests. The results demonstrate that the third algorithm, the hybrid multiphase ALNS-TS-TA algorithm, delivers the best performance. This approach simultaneously leverages the large-area search capabilities of ALNS for exploration and effectively escapes local optima when the multiphase ALNS is coupled with TS and TA. Furthermore, in our experiments, the hybrid multiphase ALNS-TS-TA algorithm is the only one that shows potential for improving results with increased computation time across all attempts.

cs.AI

Toward Copyright Integrity and Verifiability via Multi-Bit Watermarking for Intelligent Transportation Systems

Intelligent transportation systems (ITS) use advanced technologies such as artificial intelligence to significantly improve traffic flow management efficiency, and promote the intelligent development of the transportation industry. However, if the data in ITS is attacked, such as tampering or forgery, it will endanger public safety and cause social losses. Therefore, this paper proposes a watermarking that can verify the integrity of copyright in response to the needs of ITS, termed ITSmark. ITSmark focuses on functions such as extracting watermarks, verifying permission, and tracing tampered locations. The scheme uses the copyright information to build the multi-bit space and divides this space into multiple segments. These segments will be assigned to tokens. Thus, the next token is determined by its segment which contains the copyright. In this way, the obtained data contains the custom watermark. To ensure the authorization, key parameters are encrypted during copyright embedding to obtain cipher data. Only by possessing the correct cipher data and private key, can the user entirely extract the watermark. Experiments show that ITSmark surpasses baseline performances in data quality, extraction accuracy, and unforgeability. It also shows unique capabilities of permission verification and tampered location tracing, which ensures the security of extraction and the reliability of copyright verification. Furthermore, ITSmark can also customize the watermark embedding position and proportion according to user needs, making embedding more flexible.

cs.CR

U-GIFT: Uncertainty-Guided Firewall for Toxic Speech in Few-Shot Scenario

With the widespread use of social media, user-generated content has surged on online platforms. When such content includes hateful, abusive, offensive, or cyberbullying behavior, it is classified as toxic speech, posing a significant threat to the online ecosystem's integrity and safety. While manual content moderation is still prevalent, the overwhelming volume of content and the psychological strain on human moderators underscore the need for automated toxic speech detection. Previously proposed detection methods often rely on large annotated datasets; however, acquiring such datasets is both costly and challenging in practice. To address this issue, we propose an uncertainty-guided firewall for toxic speech in few-shot scenarios, U-GIFT, that utilizes self-training to enhance detection performance even when labeled data is limited. Specifically, U-GIFT combines active learning with Bayesian Neural Networks (BNNs) to automatically identify high-quality samples from unlabeled data, prioritizing the selection of pseudo-labels with higher confidence for training based on uncertainty estimates derived from model predictions. Extensive experiments demonstrate that U-GIFT significantly outperforms competitive baselines in few-shot detection scenarios. In the 5-shot setting, it achieves a 14.92\% performance improvement over the basic model. Importantly, U-GIFT is user-friendly and adaptable to various pre-trained language models (PLMs). It also exhibits robust performance in scenarios with sample imbalance and cross-domain settings, while showcasing strong generalization across various language applications. We believe that U-GIFT provides an efficient solution for few-shot toxic speech detection, offering substantial support for automated content moderation in cyberspace, thereby acting as a firewall to promote advancements in cybersecurity.

cs.SD

State-of-the-art Advances of Deep-learning Linguistic Steganalysis Research

With the evolution of generative linguistic steganography techniques, conventional steganalysis falls short in robustly quantifying the alterations induced by steganography, thereby complicating detection. Consequently, the research paradigm has pivoted towards deep-learning-based linguistic steganalysis. This study offers a comprehensive review of existing contributions and evaluates prevailing developmental trajectories. Specifically, we first provided a formalized exposition of the general formulas for linguistic steganalysis, while comparing the differences between this field and the domain of text classification. Subsequently, we classified the existing work into two levels based on vector space mapping and feature extraction models, thereby comparing the research motivations, model advantages, and other details. A comparative analysis of the experiments is conducted to assess the performances. Finally, the challenges faced by this field are discussed, and several directions for future development and key issues that urgently need to be addressed are proposed.

cs.CL

Experimental measurement and a physical interpretation of quantum shadow enumerators

Throughout its history, the theory of quantum error correction has heavily benefited from translating classical concepts into the quantum setting. In particular, classical notions of weight enumerators, which relate to the performance of an error-correcting code, and MacWilliams' identity, which helps to compute enumerators, have been generalized to the quantum case. In this work, we establish a distinct relationship between the theoretical machinery of quantum weight enumerators and a seemingly unrelated physics experiment: we prove that Rains' quantum shadow enumerators - a powerful mathematical tool - arise as probabilities of observing fixed numbers of triplets in a Bell sampling experiment. This insight allows us to develop here a rigorous framework for the direct measurement of quantum weight enumerators, thus enabling experimental and theoretical studies of the entanglement structure of any quantum error-correcting code or state under investigation. On top of that, we derive concrete sample complexity bounds and physically-motivated robustness guarantees against unavoidable experimental imperfections. Finally, we experimentally demonstrate the possibility of directly measuring weight enumerators on a trapped-ion quantum computer. Our experimental findings are in good agreement with theoretical predictions and illuminate how entanglement theory and quantum error correction can cross-fertilize each other once Bell sampling experiments are combined with the theoretical machinery of quantum weight enumerators.

quant-ph

Mobile Robot Oriented Large-Scale Indoor Dataset for Dynamic Scene Understanding

Most existing robotic datasets capture static scene data and thus are limited in evaluating robots' dynamic performance. To address this, we present a mobile robot oriented large-scale indoor dataset, denoted as THUD (Tsinghua University Dynamic) robotic dataset, for training and evaluating their dynamic scene understanding algorithms. Specifically, the THUD dataset construction is first detailed, including organization, acquisition, and annotation methods. It comprises both real-world and synthetic data, collected with a real robot platform and a physical simulation platform, respectively. Our current dataset includes 13 larges-scale dynamic scenarios, 90K image frames, 20M 2D/3D bounding boxes of static and dynamic objects, camera poses, and IMU. The dataset is still continuously expanding. Then, the performance of mainstream indoor scene understanding tasks, e.g. 3D object detection, semantic segmentation, and robot relocalization, is evaluated on our THUD dataset. These experiments reveal serious challenges for some robot scene understanding tasks in dynamic scenes. By sharing this dataset, we aim to foster and iterate new mobile robot algorithms quickly for robot actual working dynamic environment, i.e. complex crowded dynamic scenes.

cs.RO

Linguistic Steganalysis via LLMs: Two Modes for Efficient Detection of Strongly Concealed Stego

To detect stego (steganographic text) in complex scenarios, linguistic steganalysis (LS) with various motivations has been proposed and achieved excellent performance. However, with the development of generative steganography, some stegos have strong concealment, especially after the emergence of LLMs-based steganography, the existing LS has low detection or cannot detect them. We designed a novel LS with two modes called LSGC. In the generation mode, we created an LS-task "description" and used the generation ability of LLM to explain whether texts to be detected are stegos. On this basis, we rethought the principle of LS and LLMs, and proposed the classification mode. In this mode, LSGC deleted the LS-task "description" and used the "causalLM" LLMs to extract steganographic features. The LS features can be extracted by only one pass of the model, and a linear layer with initialization weights is added to obtain the classification probability. Experiments on strongly concealed stegos show that LSGC significantly improves detection and reaches SOTA performance. Additionally, LSGC in classification mode greatly reduces training time while maintaining high performance.

cs.CL