SearcharxivSearch

arXiv subjects

Jiahao Cao

Publications and source records attributed to Jiahao Cao.

10 recordsLinked to original sources

Functional BART with Shape Priors: A Bayesian Tree Approach to Constrained Functional Regression

Motivated by the remarkable success of Bayesian additive regression trees (BART) in regression modelling, we propose a novel nonparametric Bayesian method, termed Functional BART (FBART), tailored specifically for function-on-scalar regression. FBART leverages spline-based representations for functional responses coupled with a flexible tree-based partitioning structure, effectively capturing complex and heterogeneous relationships between response curves and scalar predictors. To facilitate efficient posterior inference, we develop a customized Bayesian backfitting algorithm. Additionally, we extend FBART by introducing shape constraints (e.g., monotonicity or convexity) on the response curves, enabling enhanced estimation and prediction when prior shape information is available. The use of shape priors ensures that posterior samples respect the specified functional constraints. Under mild regularity conditions, we establish posterior convergence rates for both FBART and its shape-constrained variant, demonstrating rate adaptivity to unknown smoothness. Extensive simulation studies and analyses of two real datasets illustrate the superior estimation accuracy and predictive performance of our proposed methods compared to existing state-of-the-art alternatives.

stat.ME

Bayesian ACCESS for Understanding Latent Epidemic Trajectories from Publicly Released Suppressed Data: Application to U.S. Opioid-related Overdose Mortality

Publicly released health statistics play a central role in characterizing temporal trends and identifying structural changes in population health. However, disclosure limitation through suppression of small cell counts, as implemented in systems such as the Centers for Disease Control and Prevention Wide-ranging ONline Data for Epidemiologic Research (CDC WONDER), produces partially observed count data that complicate statistical inference. These challenges are particularly acute for rare outcomes and subgroup analyses, where suppression is widespread and varies across geographic regions, demographic populations, and time. We propose Bayesian ACCESS (Autoregressive Change-point and Clustering Estimation for Suppressed Count Series), a Bayesian hierarchical framework for inference on latent epidemic trajectories and their structural changes from disclosure-limited health statistics. The proposed model directly represents suppressed count data through a suppression-aware observation model, jointly infers multiple temporal change points and latent trajectories, and borrows information across related geographic and demographic populations through Bayesian nonparametric clustering while preserving meaningful heterogeneity. We apply Bayesian ACCESS to opioid-related overdose mortality data from CDC WONDER for U.S. states from 1999 to 2024. The analysis identifies distinct subgroup-specific epidemic trajectories and structural changes that would be difficult to characterize using publicly released health statistics without explicitly accounting for data suppression.

stat.ME

TRACE: Timely Retrieval and Alignment for Cybersecurity Knowledge Graph Construction and Expansion

The rapid evolution of cyber threats has highlighted significant gaps in security knowledge integration. Cybersecurity Knowledge Graphs (CKGs) relying on structured data inherently exhibit hysteresis, as the timely incorporation of rapidly evolving unstructured data remains limited, potentially leading to the omission of critical insights for risk analysis. To address these limitations, we introduce TRACE, a framework designed to integrate structured and unstructured cybersecurity data sources. TRACE integrates knowledge from 24 structured databases and 3 categories of unstructured data, including APT reports, papers, and repair notices. Leveraging Large Language Models (LLMs), TRACE facilitates efficient entity extraction and alignment, enabling continuous updates to the CKG. Evaluations demonstrate that TRACE achieves a 1.8x increase in node coverage compared to existing CKGs. TRACE attains the precision of 86.08%, the recall of 76.92%, and the F1 score of 81.24% in entity extraction, surpassing the best-known LLM-based baselines by 7.8%. Furthermore, our entity alignment methods effectively harmonize entities with existing knowledge structures, enhancing the integrity and utility of the CKG. With TRACE, threat hunters and attack analysts gain real-time, holistic insights into vulnerabilities, attack methods, and defense technologies.

cs.CR

String Breaking and Glueball Dynamics in $2+1$D Quantum Link Electrodynamics

At the heart of quark confinement and hadronization, the physics of flux strings has recently become a focal point in the field of quantum simulation of high-energy physics (HEP). Despite considerable progress, a detailed understanding of the behavior of flux strings in quantum simulation-relevant lattice formulations of gauge theories has remained limited to the lowest truncations of the gauge field, which are severely limited in their ability to draw conclusions about the quantum field theory limit. Here, we employ tensor network simulations to investigate the behavior of flux strings in a quantum link formulation of $2+1$D quantum electrodynamics (QED) with a spin-$1$ representation of the gauge field. We first map out the ground-state phase diagram of this model in the presence of two spatially separated static charges, revealing distinct microscopic processes responsible for string breaking, including a two-stage breaking mechanism not possible in the spin-$\frac{1}{2}$ formulation. Starting in different initial product state string configurations, we then explore far-from-equilibrium quench dynamics across various parameter regimes, demonstrating genuine $2+1$D real-time string breaking and glueball-like bound state formation, with the latter not possible in the spin-$\frac{1}{2}$ formulation. In and out of equilibrium, we consider different values and placements of the static charges. Finally, we provide efficient qudit circuits for a quantum simulation experiment in which our results can be observed in state-of-the-art ion-trap setups. Our findings lay the groundwork for quantum simulations of flux strings towards the quantum field theory limit.

hep-lat

Joint Estimation of a Two-Phase Spin Rotation beyond Classical Limit

Quantum metrology employs entanglement to enhance measurement precision. The focus and progress so far have primarily centered on estimating a single parameter. In diverse application scenarios, the estimation of more than one single parameter is often required. Joint estimation of multiple parameters can benefit from additional advantages for further enhanced precision. Here we report quantum-enhanced measurement of simultaneous spin rotations around two orthogonal axes, making use of spin-nematic squeezing in an atomic Bose-Einstein condensate. Aided by the $F=2$ atomic ground hyperfine manifold coupled to the nematic-squeezed $F=1$ states as an auxiliary field through a sequence of microwave (MW) pulses, simultaneous measurement of multiple spin-1 observables is demonstrated, reaching an enhancement of 3.3 to 6.3 decibels (dB) beyond the classical limit over a wide range of rotation angles. Our work realizes the first enhanced multi-parameter estimation using entangled massive particles as a probe. The techniques developed and the protocols implemented also highlight the application of two-mode squeezed vacuum states in quantum-enhanced sensing of noncommuting spin rotations simultaneously.

quant-ph

How do the professional players select their shot locations? An analysis of Field Goal Attempts via Bayesian Additive Regression Trees

Basketball analytics has significantly advanced our understanding of the game, with shot selection emerging as a critical factor in both individual and team performance. With the advent of player tracking technologies, a wealth of granular data on shot attempts has become available, enabling a deeper analysis of shooting behavior. However, modeling shot selection presents unique challenges due to the spatial and contextual complexities influencing shooting decisions. This paper introduces a novel approach to the analysis of basketball shot data, focusing on the spatial distribution of shot attempts, also known as intensity surfaces. We model these intensity surfaces using a Functional Bayesian Additive Regression Trees (FBART) framework, which allows for flexible, nonparametric regression, and uncertainty quantification while addressing the nonlinearity and nonstationarity inherent in shot selection patterns to provide a more accurate representation of the factors driving player performance; we further propose the Adaptive Functional Bayesian Additive Regression Trees (AFBART) model, which builds on FBART by introducing adaptive basis functions for improved computational efficiency and model fit. AFBART is particularly well suited for the analysis of two-dimensional shot intensity surfaces and provides a robust tool for uncovering latent patterns in shooting behavior. Through simulation studies and real-world applications to NBA player data, we demonstrate the effectiveness of the model in quantifying shooting tendencies, improving performance predictions, and informing strategic decisions for coaches, players, and team managers. This work represents a significant step forward in the statistical modeling of basketball shot selection and its applications in optimizing game strategies.

stat.AP

What Influences the Field Goal Attempts of Professional Players? Analysis of Basketball Shot Charts via Log Gaussian Cox Processes with Spatially Varying Coefficients

Basketball shot charts provide valuable information regarding local patterns of in-game performance to coaches, players, sports analysts, and statisticians. The spatial patterns of where shots were attempted and whether the shots were successful suggest options for offensive and defensive strategies as well as historical summaries of performance against particular teams and players. The data represent a marked spatio-temporal point process with locations representing locations of attempted shots and an associated mark representing the shot's outcome (made/missed). Here, we develop a Bayesian log Gaussian Cox process model allowing joint analysis of the spatial pattern of locations and outcomes of shots across multiple games. We build a hierarchical model for the log intensity function using Gaussian processes, and allow spatially varying effects for various game-specific covariates. We aim to model the spatial relative risk under different covariate values. For inference via posterior simulation, we design a Markov chain Monte Carlo (MCMC) algorithm based on a kernel convolution approach. We illustrate the proposed method using extensive simulation studies. A case study analyzing the shot data of NBA legends Stephen Curry, LeBron James, and Michael Jordan highlights the effectiveness of our approach in real-world scenarios and provides practical insights into optimizing shooting strategies by examining how different playing conditions, game locations, and opposing team strengths impact shooting efficiency.

stat.ME

Enhancing Pre-Trained Language Models for Vulnerability Detection via Semantic-Preserving Data Augmentation

With the rapid development and widespread use of advanced network systems, software vulnerabilities pose a significant threat to secure communications and networking. Learning-based vulnerability detection systems, particularly those leveraging pre-trained language models, have demonstrated significant potential in promptly identifying vulnerabilities in communication networks and reducing the risk of exploitation. However, the shortage of accurately labeled vulnerability datasets hinders further progress in this field. Failing to represent real-world vulnerability data variety and preserve vulnerability semantics, existing augmentation approaches provide limited or even counterproductive contributions to model training. In this paper, we propose a data augmentation technique aimed at enhancing the performance of pre-trained language models for vulnerability detection. Given the vulnerability dataset, our method performs natural semantic-preserving program transformation to generate a large volume of new samples with enriched data diversity and variety. By incorporating our augmented dataset in fine-tuning a series of representative code pre-trained models (i.e., CodeBERT, GraphCodeBERT, UnixCoder, and PDBERT), up to 10.1% increase in accuracy and 23.6% increase in F1 can be achieved in the vulnerability detection task. Comparison results also show that our proposed method can substantially outperform other prominent vulnerability augmentation approaches.

cs.CR

Integrating Coarse Granularity Part-level Features with Supervised Global-level Features for Person Re-identification

Holistic person re-identification (Re-ID) and partial person re-identification have achieved great progress respectively in recent years. However, scenarios in reality often include both holistic and partial pedestrian images, which makes single holistic or partial person Re-ID hard to work. In this paper, we propose a robust coarse granularity part-level person Re-ID network (CGPN), which not only extracts robust regional level body features, but also integrates supervised global features for both holistic and partial person images. CGPN gains two-fold benefit toward higher accuracy for person Re-ID. On one hand, CGPN learns to extract effective body part features for both holistic and partial person images. On the other hand, compared with extracting global features directly by backbone network, CGPN learns to extract more accurate global features with a supervision strategy. The single model trained on three Re-ID datasets including Market-1501, DukeMTMC-reID and CUHK03 achieves state-of-the-art performances and outperforms any existing approaches. Especially on CUHK03, which is the most challenging dataset for person Re-ID, in single query mode, we obtain a top result of Rank-1/mAP=87.1\%/83.6\% with this method without re-ranking, outperforming the current best method by +7.0\%/+6.7\%.

cs.CV

When the Differences in Frequency Domain are Compensated: Understanding and Defeating Modulated Replay Attacks on Automatic Speech Recognition

Automatic speech recognition (ASR) systems have been widely deployed in modern smart devices to provide convenient and diverse voice-controlled services. Since ASR systems are vulnerable to audio replay attacks that can spoof and mislead ASR systems, a number of defense systems have been proposed to identify replayed audio signals based on the speakers' unique acoustic features in the frequency domain. In this paper, we uncover a new type of replay attack called modulated replay attack, which can bypass the existing frequency domain based defense systems. The basic idea is to compensate for the frequency distortion of a given electronic speaker using an inverse filter that is customized to the speaker's transform characteristics. Our experiments on real smart devices confirm the modulated replay attacks can successfully escape the existing detection mechanisms that rely on identifying suspicious features in the frequency domain. To defeat modulated replay attacks, we design and implement a countermeasure named DualGuard. We discover and formally prove that no matter how the replay audio signals could be modulated, the replay attacks will either leave ringing artifacts in the time domain or cause spectrum distortion in the frequency domain. Therefore, by jointly checking suspicious features in both frequency and time domains, DualGuard can successfully detect various replay attacks including the modulated replay attacks. We implement a prototype of DualGuard on a popular voice interactive platform, ReSpeaker Core v2. The experimental results show DualGuard can achieve 98% accuracy on detecting modulated replay attacks.

cs.CR