Searcharxiv⌕ Search

arXiv subjects

Yi Fan

Publications and source records attributed to Yi Fan.

At least 37 records · Page 2Linked to original sources

FraudFox: Adaptable Fraud Detection in the Real World

The proposed method (FraudFox) provides solutions to adversarial attacks in a resource constrained environment. We focus on questions like the following: How suspicious is `Smith', trying to buy \$500 shoes, on Monday 3am? How to merge the risk scores, from a handful of risk-assessment modules (`oracles') in an adversarial environment? More importantly, given historical data (orders, prices, and what-happened afterwards), and business goals/restrictions, which transactions, like the `Smith' transaction above, which ones should we `pass', versus send to human investigators? The business restrictions could be: `at most $x$ investigations are feasible', or `at most \$$y$ lost due to fraud'. These are the two research problems we focus on, in this work. One approach to address the first problem (`oracle-weighting'), is by using Extended Kalman Filters with dynamic importance weights, to automatically and continuously update our weights for each 'oracle'. For the second problem, we show how to derive an optimal decision surface, and how to compute the Pareto optimal set, to allow what-if questions. An important consideration is adaptation: Fraudsters will change their behavior, according to our past decisions; thus, we need to adapt accordingly. The resulting system, \method, is scalable, adaptable to changing fraudster behavior, effective, and already in \textbf{production} at Amazon. FraudFox augments a fraud prevention sub-system and has led to significant performance gains.

cs.CR↗

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration

Preference-based reinforcement learning (PbRL) can help avoid sophisticated reward designs and align better with human intentions, showing great promise in various real-world applications. However, obtaining human feedback for preferences can be expensive and time-consuming, which forms a strong barrier for PbRL. In this work, we address the problem of low query efficiency in offline PbRL, pinpointing two primary reasons: inefficient exploration and overoptimization of learned reward functions. In response to these challenges, we propose a novel algorithm, \textbf{O}ffline \textbf{P}b\textbf{R}L via \textbf{I}n-\textbf{D}ataset \textbf{E}xploration (OPRIDE), designed to enhance the query efficiency of offline PbRL. OPRIDE consists of two key features: a principled exploration strategy that maximizes the informativeness of the queries and a discount scheduling mechanism aimed at mitigating overoptimization of the learned reward functions. Through empirical evaluations, we demonstrate that OPRIDE significantly outperforms prior methods, achieving strong performance with notably fewer queries. Moreover, we provide theoretical guarantees of the algorithm's efficiency. Experimental results across various locomotion, manipulation, and navigation tasks underscore the efficacy and versatility of our approach.

cs.LG↗

Machine learning modularity

Based on a transformer based sequence-to-sequence architecture combined with a dynamic batching algorithm, this work introduces a machine learning framework for automatically simplifying complex expressions involving multiple elliptic Gamma functions, including the $q$-$θ$ function and the elliptic Gamma function. The model learns to apply algebraic identities, particularly the SL$(2,\mathbb{Z})$ and SL$(3,\mathbb{Z})$ modular transformations, to reduce heavily scrambled expressions to their canonical forms. Experimental results show that the model achieves over 99\% accuracy on in-distribution tests and maintains robust performance (exceeding 90\% accuracy) under significant extrapolation, such as with deeper scrambling depths. This demonstrates that the model has internalized the underlying algebraic rules of modular transformations rather than merely memorizing training patterns. Our work presents the first successful application of machine learning to perform symbolic simplification using modular identities, offering a new automated tool for computations with special functions in quantum field theory and the string theory.

hep-th↗

Scalable parallel simulation of quantum circuits on CPU and GPU systems

Quantum computing enables parallelism through superposition and entanglement and offers advantages over classical computing architectures. However, due to the limitations of current quantum hardware in the noisy intermediate-scale quantum (NISQ) era, classical simulation remains a critical tool for developing quantum algorithms. In this research, we present a comprehensive parallelization solution for the Q$^2$Chemistry software package, delivering significant performance improvements for the full-amplitude simulator on both CPU and GPU platforms. By incorporating batch-buffered overlap processing, dependency-aware gate contraction and staggered multi-gate parallelism, our optimizations significantly enhance the simulation speed compared to unoptimized baselines, demonstrating the effectiveness of hybrid-level parallelism in HPC systems. Benchmark results show that Q$^2$Chemistry consistently outperforms current state-of-the-art open-source simulators across various circuit types. These benchmarks highlight the capability of Q$^2$Chemistry to effectively handle large-scale quantum simulations with high efficiency and high portability.

quant-ph↗

The Three-Dimensional Velocity Field of Kinesin-Driven Microtubules in Torroidal Channels

We study two regimes of flow in multiple three-dimensional toroidal channels by tracking the fluorescent spherical particles in the kinesin-driven microtubule systems: ``chaotic'' flow and ``coherent'' flow. In the smallest aspect ratio torus, where the channel height $h$ is a quarter of the width $w$, the active system shows zero mean velocity, small-scale isotropy in fluctuation and no persistent flow structure. In other tori with higher aspect ratios $h/w$ close to 1, we find faster coherent flows along the azimuthal direction and increasing fluctuation strengths with growing confinement geometries. Regardless of flow regimes, the flow profiles at $r-z$ cross-section and $r-θ$ plane are symmetric. The ``coherent'' profiles show two criteria: ``Poiseuille-like'' profiles, which have the peak velocities near the centers of channels; a ``peak-separated'' profile, which has four peak velocities near a certain distance to four confining surfaces. These flow profiles, after scaled by the local isotropic fluctuation strength, reveal universal three-dimensional flow structures among the ``Poiseuille-like'' criterion and the same level of scaled peak velocity at the ``peak-separated'' one. These results illustrate scalable flow structures in this kinesin-driven microtubule active system.

physics.flu-dyn↗

StealthRank: LLM Ranking Manipulation via Stealthy Prompt Optimization

The integration of large language models (LLMs) into information retrieval systems introduces new attack surfaces, particularly for adversarial ranking manipulations. We present $\textbf{StealthRank}$, a novel adversarial attack method that manipulates LLM-driven ranking systems while maintaining textual fluency and stealth. Unlike existing methods that often introduce detectable anomalies, StealthRank employs an energy-based optimization framework combined with Langevin dynamics to generate StealthRank Prompts (SRPs)-adversarial text sequences embedded within item or document descriptions that subtly yet effectively influence LLM ranking mechanisms. We evaluate StealthRank across multiple LLMs, demonstrating its ability to covertly boost the ranking of target items while avoiding explicit manipulation traces. Our results show that StealthRank consistently outperforms state-of-the-art adversarial ranking baselines in both effectiveness and stealth, highlighting critical vulnerabilities in LLM-driven ranking systems. Our code is publicly available at $\href{https://github.com/Tangyiming205069/controllable-seo}{here}$.

cs.IR↗

SCORE: Syntactic Code Representations for Static Script Malware Detection

As businesses increasingly adopt cloud technologies, they also need to be aware of new security challenges, such as server-side script attacks, to ensure the integrity of their systems and data. These scripts can steal data, compromise credentials, and disrupt operations. Unlike executables with standardized formats (e.g., ELF, PE), scripts are plaintext files with diverse syntax, making them harder to detect using traditional methods. As a result, more sophisticated approaches are needed to protect cloud infrastructures from these evolving threats. In this paper, we propose novel feature extraction and deep learning (DL)-based approaches for static script malware detection, targeting server-side threats. We extract features from plain-text code using two techniques: syntactic code highlighting (SCH) and abstract syntax tree (AST) construction. SCH leverages complex regexes to parse syntactic elements of code, such as keywords, variable names, etc. ASTs generate a hierarchical representation of a program's syntactic structure. We then propose a sequential and a graph-based model that exploits these feature representations to detect script malware. We evaluate our approach on more than 400K server-side scripts in Bash, Python and Perl. We use a balanced dataset of 90K scripts for training, validation, and testing, with the remaining from 400K reserved for further analysis. Experiments show that our method achieves a true positive rate (TPR) up to 81% higher than leading signature-based antivirus solutions, while maintaining a low false positive rate (FPR) of 0.17%. Moreover, our approach outperforms various neural network-based detectors, demonstrating its effectiveness in learning code maliciousness for accurate detection of script malware.

cs.CR↗

Differentiable matrix product states for simulating variational quantum computational chemistry

Quantum Computing is believed to be the ultimate solution for quantum chemistry problems. Before the advent of large-scale, fully fault-tolerant quantum computers, the variational quantum eigensolver~(VQE) is a promising heuristic quantum algorithm to solve real world quantum chemistry problems on near-term noisy quantum computers. Here we propose a highly parallelizable classical simulator for VQE based on the matrix product state representation of quantum state, which significantly extend the simulation range of the existing simulators. Our simulator seamlessly integrates the quantum circuit evolution into the classical auto-differentiation framework, thus the gradients could be computed efficiently similar to the classical deep neural network, with a scaling that is independent of the number of variational parameters. As applications, we use our simulator to study commonly used small molecules such as HF, HCl, LiH and H$_2$O, as well as larger molecules CO$_2$, BeH$_2$ and H$_4$ with up to $40$ qubits. The favorable scaling of our simulator against the number of qubits and the number of parameters could make it an ideal testing ground for near-term quantum algorithms and a perfect benchmarking baseline for oncoming large scale VQE experiments on noisy quantum computers.

quant-ph↗

NNQS-Transformer: an Efficient and Scalable Neural Network Quantum States Approach for Ab initio Quantum Chemistry

Neural network quantum state (NNQS) has emerged as a promising candidate for quantum many-body problems, but its practical applications are often hindered by the high cost of sampling and local energy calculation. We develop a high-performance NNQS method for \textit{ab initio} electronic structure calculations. The major innovations include: (1) A transformer based architecture as the quantum wave function ansatz; (2) A data-centric parallelization scheme for the variational Monte Carlo (VMC) algorithm which preserves data locality and well adapts for different computing architectures; (3) A parallel batch sampling strategy which reduces the sampling cost and achieves good load balance; (4) A parallel local energy evaluation scheme which is both memory and computationally efficient; (5) Study of real chemical systems demonstrates both the superior accuracy of our method compared to state-of-the-art and the strong and weak scalability for large molecular systems with up to $120$ spin orbitals.

quant-ph↗

Circuit-Depth Reduction of Unitary-Coupled-Cluster Ansatz by Energy Sorting

Quantum computation represents a revolutionary approach for solving problems in quantum chemistry. However, due to the limited quantum resources in the current noisy intermediate-scale quantum (NISQ) devices, quantum algorithms for large chemical systems remains a major task. In this work, we demonstrate that the circuit depth of the unitary coupled cluster (UCC) and UCC-based ansatzes in the algorithm of variational quantum eigensolver can be significantly reduced by an energy-sorting strategy. Specifically, subsets of excitation operators are first pre-screened from the operator pool according to its contribution to the total energy. The quantum circuit ansatz is then iteratively constructed until the convergence of the final energy to a typical accuracy. For demonstration, this method has been successfully applied to molecular and periodic systems. Particularly, a reduction of 50\%$\sim$98\% in the number of operators is observed while retaining the accuracy of the origin UCCSD operator pools. This method can be straightforwardly extended to general parametric variational ansatzes.

quant-ph↗

Towards practical and massively parallel quantum computing emulation for quantum chemistry

Quantum computing is moving beyond its early stage and seeking for commercial applications in chemical and biomedical sciences. In the current noisy intermediate-scale quantum computing era, quantum resource is too scarce to support these explorations. Therefore, it is valuable to emulate quantum computing on classical computers for developing quantum algorithms and validating quantum hardware. However, existing simulators mostly suffer from the memory bottleneck so developing the approaches for large-scale quantum chemistry calculations remains challenging. Here we demonstrate a high-performance and massively parallel variational quantum eigensolver (VQE) simulator based on matrix product states, combined with embedding theory for solving large-scale quantum computing emulation for quantum chemistry on HPC platforms. We apply this method to study the torsional barrier of ethane and the quantification of the protein-ligand interactions. Our largest simulation reaches $1000$ qubits, and a performance of $216.9$ PFLOPS is achieved on a new Sunway supercomputer, which sets the state-of-the-art for quantum computing emulation for quantum chemistry

quant-ph↗

Quantum Neural Network Inspired Hardware Adaptable Ansatz for Efficient Quantum Simulation of Chemical Systems

The variational quantum eigensolver is a promising way to solve the Schrödinger equation on a noisy intermediate-scale quantum (NISQ) computer, while its success relies on a well-designed wavefunction ansatz. Compared to physically motivated ansatzes, hardware heuristic ansatzes usually lead to a shallower circuit, but it may still be too deep for an NISQ device. Inspired by the quantum neural network, we propose a new hardware heuristic ansatz where the circuit depth can be significantly reduced by introducing ancilla qubits, which makes a practical simulation of a chemical reaction with more than 20 atoms feasible on a currently available quantum computer. More importantly, the expressibility of this new ansatz can be improved by increasing either the depth or the width of the circuit, which makes it adaptable to different hardware environments. These results open a new avenue to develop practical applications of quantum computation in the NISQ era.

quant-ph↗

Quantum circuit matrix product state ansatz for large-scale simulations of molecules

As in the density matrix renormalization group (DMRG) method, approximating many-body wave function of electrons using a matrix product state (MPS) is a promising way to solve electronic structure problems. The expressibility of an MPS is determined by the size of the matrices or in other words the bond dimension, which unfortunately should be very large in many cases. In this study, we propose to calculate the ground state energies of molecular systems by variationally optimizing quantum circuit MPS (QCMPS) with a relatively small number of qubits. It is demonstrated that with carefully chosen circuit structure and orbital localization scheme, QCMPS can reach a similar accuracy as that achieved in DMRG with an exponentially large bond dimension. QCMPS simulation of a linear molecule with 50 orbitals can reach the chemical accuracy using only 6 qubits at a moderate circuit depth. These results suggest that QCMPS is a promising wave function ansatz in the variational quantum eigensolver algorithm for molecular systems.

quant-ph↗

A real neural network state for quantum chemistry

The restricted Boltzmann machine (RBM) has been successfully applied to solve the many-electron Schr$\ddot{\text{o}}$dinger equation. In this work we propose a single-layer fully connected neural network adapted from RBM and apply it to study ab initio quantum chemistry problems. Our contribution is two-fold: 1) our neural network only uses real numbers to represent the real electronic wave function, while we obtain comparable precision to RBM for various prototypical molecules; 2) we show that the knowledge of the Hartree-Fock reference state can be used to systematically accelerate the convergence of the variational Monte Carlo algorithm as well as to increase the precision of the final energy.

quant-ph↗

Data-Driven Prediction and Evaluation on Future Impact of Energy Transition Policies in Smart Regions

To meet widely recognised carbon neutrality targets, over the last decade metropolitan regions around the world have implemented policies to promote the generation and use of sustainable energy. Nevertheless, there is an availability gap in formulating and evaluating these policies in a timely manner, since sustainable energy capacity and generation are dynamically determined by various factors along dimensions based on local economic prosperity and societal green ambitions. We develop a novel data-driven platform to predict and evaluate energy transition policies by applying an artificial neural network and a technology diffusion model. Using Singapore, London, and California as case studies of metropolitan regions at distinctive stages of energy transition, we show that in addition to forecasting renewable energy generation and capacity, the platform is particularly powerful in formulating future policy scenarios. We recommend global application of the proposed methodology to future sustainable energy transition in smart regions.

econ.GN↗

The applications of EPR steering in quantum teleportation for two- or three-qubit system

EPR steering is an important quantum resource in quantum information and computation. In this paper, its applications in quantum teleportation are investigated. First of all, the upper bound of the average teleportation fidelity based on the EPR steering is derived. When the receiver can only perform the identity or the Pauli rotation operations, the X-type states which violate the three-setting linear steering inequality could be used for teleportation. In the end, the steering observables and the average teleportation fidelities of the two-qubit reduced states for three-qubit pure states maintain the same ordering. The complementary relations between the steerable observables and the average teleportation fidelities for three-qubit pure states are also established.

quant-ph↗

An efficient dosimetry method with a Faraday cup for small animal, small-field proton irradiation under conventional and ultra-high dose rates

Introduction: We developed and evaluated a method for dose calibration and monitoring under conventional and ultra-high dose rates for small animal experiments with small-field proton beams using a Faraday cup. Methods: We determined a relationship between dose and optical density (OD) of EBT-XD Gafchromic film using scanned 10x10 cm2 proton pencil beams delivered at clinical dose rates; the dose was measured with an Advanced Markus chamber. On a small animal proton irradiation platform, double-scattered pencil beams with 5 or 8 mm diameter brass collimation at conventional and ultra-high dose rates were delivered to the EBT-XD films. The proton fluence charges were collected by a Faraday cup placed downstream from the film. The average of the irradiated film ODs was related to the Faraday cup charges. A conversion from the Faraday cup charge to the average dose of the small-field proton beam was then obtained. Results: The relationship between the small-field average profile dose and Faraday cup charge was established for 10 and 15 Gy mice FLASH experiments. The film OD was found to be independent of dose rate. At small-animal treatments, the Faraday cup readings were conveniently used to QA and monitor the delivered dose and dose rates to the mice under conventional and ultra-high dose rates. Conclusion: The dose calibration and monitoring method with Faraday cup for small animal proton FLASH experiments is time-efficient and cost-effective and can be used for irradiations of various small field sizes. The same approach can also be adopted for clinical proton dosimetry for small-field irradiations.

physics.med-ph↗

Divide-and-conquer variational quantum algorithms for large-scale electronic structure simulations

Exploring the potential application of quantum computers in material design and drug discovery has attracted a lot of interest in the age of quantum computing. However, the quantum resource requirement for solving practical electronic structure problems are far beyond the capacity of near-term quantum devices. In this work, we integrate the divide-and-conquer (DC) approaches into the variational quantum eigensolver (VQE) for large-scale quantum computational chemistry simulations. Two popular divide-and-conquer schemes, including many-body expansion~(MBE) fragmentation theory and density matrix embedding theory~(DMET), are employed to divide complicated problems into many small parts that are easy to implement on near-term quantum computers. Pilot applications of these methods to systems consisting of tens of atoms are performed with adaptive VQE algorithms. This work should encourage further studies of using the philosophy of DC to solve electronic structure problems on quantum computers.

quant-ph↗