SearcharxivSearch

arXiv subjects

Benjamin J. Choi

Publications and source records attributed to Benjamin J. Choi.

17 recordsLinked to original sources

Bias-Corrected Machine-Learning Estimation of Chiral Condensate Cumulants: A Retrospective Lattice QCD Case Study

We present a retrospective case study of bias-corrected machine learning (ML) estimates of traces of the inverse Dirac operator, $\text{Tr}\,M^{-n}$ ($n=1,2,3,4$), using a fixed lattice QCD dataset and examining how the results depend on the relative proportions of the labeled and training sets. Two supervised learning approaches are examined: one using $\text{Tr}\,M^{-1}$ as the input feature, and the other employing gauge observables such as the plaquette and rectangle. Beyond the direct estimation of $\text{Tr}\,M^{-n}$, we further investigate two derived applications of the ML estimations: the evaluation of the cumulants of the chiral condensate within a single ensemble and that obtained through multi-ensemble reweighting across ensembles with different quark masses. Within this fixed dataset, the bias-corrected estimates show close agreement with the full-data reference under the adopted evaluation criteria, while the uncorrected estimates can exhibit amplified deviations after the nonlinear cumulant and reweighting steps. For the approach using $\text{Tr}\,M^{-1}$ as the input feature, nominal solve-count accounting suggests that the Dirac-inversion cost could be reduced to approximately $25.75\%$ of that of the conventional calculation in the present setup. This value is a cost projection rather than an end-to-end benchmark: it assumes comparable costs for successive inversions and excludes model-training and analysis overhead.

hep-lat

Latent Structure of Affective Representations in Large Language Models

The geometric structure of latent representations in large language models (LLMs) is an active area of research, driven in part by its implications for model transparency and AI safety. Existing literature has focused mainly on general geometric and topological properties of the learnt representations, but due to a lack of ground-truth latent geometry, validating the findings of such approaches is challenging. Emotion processing provides an intriguing testbed for probing representational geometry, as emotions exhibit both categorical organization and continuous affective dimensions, which are well-established in the psychology literature. Moreover, understanding such representations carries safety relevance. In this work, we investigate the latent structure of affective representations in LLMs using geometric data analysis tools. We present three main findings. First, we show that LLMs learn coherent latent representations of affective emotions that align with widely used valence--arousal models from psychology. Second, we find that these representations exhibit nonlinear geometric structure that can nonetheless be well-approximated linearly, providing empirical support for the linear representation hypothesis commonly assumed in model transparency methods. Third, we demonstrate that the learned latent representation space can be leveraged to quantify uncertainty in emotion processing tasks. Our findings suggest that LLMs acquire affective representations with geometric structure paralleling established models of human emotion, with practical implications for model interpretability and safety.

cs.LG

A Machine Learning Approach for Lattice Gauge Fixing

Gauge fixing is an essential step in lattice QCD calculations, particularly for studying gauge-dependent observables. Traditional iterative algorithms are computationally expensive and often suffer from critical slowing down and scaling bottlenecks on large lattices. We present a novel machine learning framework for lattice gauge fixing, where Wilson lines are utilized to construct gauge transformation matrices within a convolutional neural network. The model parameters are optimized via backpropagation, and we introduce a hybrid strategy that combines a neural-network-based transformation with subsequent iterative methods. Preliminary tests on SU(3) gauge theory ensembles for Coulomb gauge demonstrate the potential of this approach to improve the efficiency of lattice gauge fixing. Furthermore, we show that the model exhibits lattice size transferability, where parameters optimized on smaller lattices remain effective for larger volumes without additional training. This framework provides a scalable path toward mitigating critical slowing down in high-precision gauge fixing.

hep-lat

Machine Learning-Based Estimation of Cumulants of Chiral Condensate via Multi-Ensemble Reweighting with Deborah.jl

We investigate a bias-corrected machine learning (ML) strategy for estimating traces of the inverse Dirac operator, $\text{Tr}\, M^{-n}$ ($n=1,2,3,4$), motivated by the need for higher-order cumulants of the chiral condensate near the finite-temperature QCD critical endpoint. Our supervised regression framework is trained on Wilson-clover ensembles with the Iwasaki gauge action, and we explore two input feature scenarios: one using $\text{Tr}\, M^{-1}$ and another relying solely on gauge observables (plaquette and rectangle), enabling a fully feature-based prediction pipeline. Using $\text{Tr}\, M^{-1}$ both as a physical input to cumulant construction and as a feature for predicting higher powers, we find that even with $\sim1\%$ labeled data, the resulting susceptibility, skewness, and kurtosis remain statistically consistent with fully measured baselines, reducing computational cost to about $26\%$. In the feature-only approach, where correlations rather than explicit stochastic traces drive the predictions, bias correction plays a more pronounced role. We quantify this impact through multi ensemble reweighting across nearby quark masses. Our results demonstrate that bias-corrected ML estimates can significantly reduce measurement overhead while preserving the stability of higher-order observables relevant for locating the QCD critical endpoint. Code for this work is available at https://github.com/saintbenjamin/Deborah.jl .

hep-lat

A Statistical Mixture-of-Experts Framework for EMG Artifact Removal in EEG: Empirical Insights and a Proof-of-Concept Application

Effective control of neural interfaces is limited by poor signal quality. While neural network-based electroencephalography (EEG) denoising methods for electromyogenic (EMG) artifacts have improved in recent years, current state-of-the-art (SOTA) models perform suboptimally in settings with high noise. To address the shortcomings of current machine learning (ML)-based denoising algorithms, we present a signal filtration algorithm driven by a new mixture-of-experts (MoE) framework. Our algorithm leverages three new statistical insights into the EEG-EMG denoising problem: (1) EMG artifacts can be partitioned into quantifiable subtypes to aid downstream MoE classification, (2) local experts trained on narrower signal-to-noise ratio (SNR) ranges can achieve performance increases through specialization, and (3) correlation-based objective functions, in conjunction with rescaling algorithms, can enable faster convergence in a neural network-based denoising context. We empirically demonstrate these three insights into EMG artifact removal and use our findings to create a new downstream MoE denoising algorithm consisting of convolutional (CNN) and recurrent (RNN) neural networks. We tested all results on a major benchmark dataset (EEGdenoiseNet) collected from 67 subjects. We found that our MoE denoising model achieved competitive overall performance with SOTA ML denoising algorithms and superior lower bound performance in high noise settings. These preliminary results highlight the promise of our MoE framework for enabling advances in EMG artifact removal for EEG processing, especially in high noise settings. Further research and development will be necessary to assess our MoE framework on a wider range of real-world test cases and explore its downstream potential to unlock more effective neural interfaces.

eess.SP

Removing Neural Signal Artifacts with Autoencoder-Targeted Adversarial Transformers (AT-AT)

Electromyogenic (EMG) noise is a major contamination source in EEG data that can impede accurate analysis of brain-specific neural activity. Recent literature on EMG artifact removal has moved beyond traditional linear algorithms in favor of machine learning-based systems. However, existing deep learning-based filtration methods often have large compute footprints and prohibitively long training times. In this study, we present a new machine learning-based system for filtering EMG interference from EEG data using an autoencoder-targeted adversarial transformer (AT-AT). By leveraging the lightweight expressivity of an autoencoder to determine optimal time-series transformer application sites, our AT-AT architecture achieves a >90% model size reduction compared to published artifact removal models. The addition of adversarial training ensures that filtered signals adhere to the fundamental characteristics of EEG data. We trained AT-AT using published neural data from 67 subjects and found that the system was able to achieve comparable test performance to larger models; AT-AT posted a mean reconstructive correlation coefficient above 0.95 at an initial signal-to-noise ratio (SNR) of 2 dB and 0.70 at -7 dB SNR. Further research generalizing these results to broader sample sizes beyond these isolated test cases will be crucial; while outside the scope of this study, we also include results from a real-world deployment of AT-AT in the Appendix.

cs.LG

Geometric Machine Learning on EEG Signals

Brain-computer interfaces (BCIs) offer transformative potential, but decoding neural signals presents significant challenges. The core premise of this paper is built around demonstrating methods to elucidate the underlying low-dimensional geometric structure present in high-dimensional brainwave data in order to assist in downstream BCI-related neural classification tasks. We demonstrate two pipelines related to electroencephalography (EEG) signal processing: (1) a preliminary pipeline removing noise from individual EEG channels, and (2) a downstream manifold learning pipeline uncovering geometric structure across networks of EEG channels. We conduct preliminary validation using two EEG datasets and situate our demonstration in the context of the BCI-relevant imagined digit decoding problem. Our preliminary pipeline uses an attention-based EEG filtration network to extract clean signal from individual EEG channels. Our primary pipeline uses a fast Fourier transform, a Laplacian eigenmap, a discrete analog of Ricci flow via Ollivier's notion of Ricci curvature, and a graph convolutional network to perform dimensionality reduction on high-dimensional multi-channel EEG data in order to enable regularizable downstream classification. Our system achieves competitive performance with existing signal processing and classification benchmarks; we demonstrate a mean test correlation coefficient of >0.95 at 2 dB on semi-synthetic neural denoising and a downstream EEG-based classification accuracy of 0.97 on distinguishing digit- versus non-digit- thoughts. Results are preliminary and our geometric machine learning pipeline should be validated by more extensive follow-up studies; generalizing these results to larger inter-subject sample sizes, different hardware systems, and broader use cases will be crucial.

cs.LG

Targeted Adversarial Denoising Autoencoders (TADA) for Neural Time Series Filtration

Current machine learning (ML)-based algorithms for filtering electroencephalography (EEG) time series data face challenges related to cumbersome training times, regularization, and accurate reconstruction. To address these shortcomings, we present an ML filtration algorithm driven by a logistic covariance-targeted adversarial denoising autoencoder (TADA). We hypothesize that the expressivity of a targeted, correlation-driven convolutional autoencoder will enable effective time series filtration while minimizing compute requirements (e.g., runtime, model size). Furthermore, we expect that adversarial training with covariance rescaling will minimize signal degradation. To test this hypothesis, a TADA system prototype was trained and evaluated on the task of removing electromyographic (EMG) noise from EEG data in the EEGdenoiseNet dataset, which includes EMG and EEG data from 67 subjects. The TADA filter surpasses conventional signal filtration algorithms across quantitative metrics (Correlation Coefficient, Temporal RRMSE, Spectral RRMSE), and performs competitively against other deep learning architectures at a reduced model size of less than 400,000 trainable parameters. Further experimentation will be necessary to assess the viability of TADA on a wider range of deployment cases.

cs.LG

Machine Learning Estimation on the Trace of Inverse Dirac Operator using the Gradient Boosting Decision Tree Regression

We present our preliminary results on the machine learning estimation of $\text{Tr} \, M^{-n}$ from other observables with the gradient boosting decision tree regression, where $M$ is the Dirac operator. Ordinarily, $\text{Tr} \, M^{-n}$ is obtained by linear CG solver for stochastic sources which needs considerable computational cost. Hence, we explore the possibility of cost reduction on the trace estimation by the adoption of gradient boosting decision tree algorithm. We also discuss effects of bias and its correction.

hep-lat

Progress report on testing robustness of the Newton method in data analysis on 2-point correlation function using a MILC HISQ ensemble

We report recent progress in data analysis on the two point correlation functions which will be prerequisite to obtain semileptonic form factors for the $B_{(s)} \to D_{(s)}\ellν$ decays. We use a MILC HISQ ensemble for the measurement. We use the HISQ action for light quarks, and the Oktay-Kronfeld (OK) action for the heavy quarks ($b$ and $c$). We used a sequential Bayesian method for the data analysis. Here we test the new fitting methodology of Benjamin J.~Choi in a completely independent manner.

hep-lat

Current progress on the semileptonic form factors for $\bar{B} \to D^{\ast} \ell \barν$ decay using the Oktay-Kronfeld action

We present recent progress in calculating the semileptonic form factors $h_{A_1}(w)$ for the $\bar{B} \to D^{\ast} \ell \barν$ decays. We use the Oktay-Kronfeld (OK) action for the charm and bottom valence quarks and the HISQ action for light quarks. We adopt the Newton method combined with the scanning method to find a good initial guess for the $χ^2$ minimizer in the fitting of the 2pt correlation functions. The main advantage is that the Newton method lets us to consume all the time slices allowed by the physical positivity. We report the first, reliable, but preliminary results for $h_{A_1}(w)/ρ_{A_1}$ at zero recoil ($w=1$). Here we use a MILC HISQ ensemble ($a = 0.12$ fm, $M_π$ = 220 MeV, and $N_f = 2 + 1 + 1$ flavors).

hep-lat

Improved data analysis on two-point correlation function with sequential Bayesian method

We report our progress in data analysis on two-point correlation functions of the $B$ meson using sequential Bayesian method. The data set of measurement is obtained using the Oktay-Kronfeld (OK) action for the bottom quarks (valence quarks) and the HISQ action for the light quarks on the MILC HISQ lattices. We find that the old initial guess for the $χ^2$ minimizer in the fitting code is poor enough to slow down the analysis somewhat. In order to find a better initial guess, we adopt the Newton method. We find that the Newton method provides a natural test to check whether the $χ^2$ minimizer finds a local minimum or the global minimum, and it also reduces the number of iterations dramatically.

hep-lat

Semileptonic $B \to D^{(\ast)} \ellν$ Decay Form Factors using the Oktay-Kronfeld Action

We report recent progress in calculating semileptonic form factors for the $\bar{B} \to D^\ast \ell \barν$ and $\bar{B} \to D \ell \barν$ decays using the Oktay-Kronfeld (OK) action for bottom and charm quarks. We use the second order in heavy quark effective power counting $\mathcal{O}(λ^2)$ improved currents in this work. The HISQ action is used for the light spectator quarks. We analyzed four $2+1+1$-flavor MILC HISQ ensembles with $a\approx 0.09\,\mathrm{fm}$, $0.12\,\mathrm{fm}$ and $M_π\approx 220\,\mathrm{MeV}$, $310\,\mathrm{MeV}$: $a09m220$, $a09m310$, $a12m220$, $a12m310$. Preliminary results for $B\to D^\ast\ellν$ decays form factor $h_{A_1}(w)$ at zero recoil ($w=1$) are reported. Preliminary results for $B \to D\,\ellν$ decays form factors $h_\pm(w)$ over a kinematic range $1<w<1.3$ are reported as well.

hep-lat

Leptonic decays of $B_{(s)}$ and $D_{(s)}$ using the OK action

We present recent progress in the lattice calculation of leptonic decay constants for $B_{(s)}$ and $D_{(s)}$ mesons using the Oktay-Kronfeld (OK) action for charm and bottom valence quarks, whose masses are tuned non-perturbatively. The calculations are done on 6 HISQ ensembles generated by the MILC collaboration with $N_f=2+1+1$ flavors. We also use the HISQ action for the light spectator quarks. Results are presented for the ratios $f_{B_s}/f_B$ and $f_{D_s}/f_D$, which reflect $SU(3)$ flavor symmetry breaking, and are independent of the renormalization constants of the axial currents.

hep-lat

Update on $B\to D^\ast \ell ν$ form factor at zero-recoil using the Oktay-Kronfeld action

We present an update on the calculation of $\bar{B}\to D^\ast \ell \barν$ semileptonic form factor at zero recoil using the Oktay-Kronfeld bottom and charm quarks on $N_f=2+1+1$ flavor HISQ ensembles generated by the MILC collaboration. Preliminary results are given for two ensembles with $a\approx 0.12$ and $0.09$ fm and $M_π\approx 310$ MeV. Calculations have been done with a number of valence quark masses, and the dependence of the form factor on them is investigated on the $a\approx 0.12$ fm ensemble. The excited state is controlled by using multistate fits to the three-point correlators measured at 4--6 source-sink separations.

hep-lat

Kaon BSM B-parameters using improved staggered fermions from $N_f=2+1$ unquenched QCD

We present results for the matrix elements of the additional $ΔS=2$ operators that appear in models of physics beyond the Standard Model (BSM), expressed in terms of four BSM $B$-parameters. Combined with experimental results for $ΔM_K$ and $ε_K$, these constrain the parameters of BSM models. We use improved staggered fermions, with valence HYP-smeared quarks and $N_f=2+1$ flavors of "asqtad" sea quarks. The configurations have been generated by the MILC collaboration. The matching between lattice and continuum four-fermion operators and bilinears is done perturbatively at one-loop order. We use three lattice spacings for the continuum extrapolation: $a\approx 0.09$, $0.06$ and $0.045\;$fm. Valence light-quark masses range down to $\approx m_s^{\rm phys}/13$ while the light sea-quark masses range down to $\approx m_s^{\rm phys}/20$. Compared to our previous published work, we have added four additional lattice ensembles, leading to better controlled extrapolations in the lattice spacing and sea-quark masses. We report final results for two renormalization scales, $μ=2\;\text{GeV}$ and $3\;\text{GeV}$, and compare them to those obtained by other collaborations. Agreement is found for two of the four BSM $B$-parameters ($B_2$ and $B_3^\text{SUSY}$). The other two ($B_4$ and $B_5$) differ significantly from those obtained using RI-MOM renormalization as an intermediate scheme, but are in agreement with recent preliminary results obtained by the RBC-UKQCD collaboration using RI-SMOM intermediate schemes.

hep-lat