SearcharxivSearch

arXiv subjects

Kevin Zhou

Publications and source records attributed to Kevin Zhou.

At least 19 recordsLinked to original sources

Ultralight Axial Dark Matter

Ultralight dark matter could be an axial vector field, a simple possibility which motivates new experiments. Couplings to fermions, through an axial vector current or a dark electric dipole moment operator, lead to enhanced spin torques compared to axion dark matter. An axial vector also has an analogue of the axion-photon coupling, though its consistent realization requires a photon mass. Under this "axial-photon" coupling, the effect of a background magnetic field is suppressed, and the strongest experimental probes involve polarimetry, electric fields in superconducting cavities, and the cosmic microwave background. Since a massive axial vector is dual to a massive two-form, our results also apply to "Kalb-Ramond" dark matter motivated by string theory.

hep-ph

Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI

Multimodal clinical models are usually judged on accuracy with every modality present, but deployment removes modalities; an echocardiogram is often unavailable where an ECG is routine. Two questions then matter beyond the size of the accuracy loss: which modality was responsible, and whether the model fails loudly or silently once that modality is dropped. The distinction is per-example and modality-level, and is separate from post-hoc feature attribution (e.g. SHAP). Models are replaced often; the evaluation that answers these questions is reused. We present a model-agnostic modality-failure framework: given N modality embeddings, any mask-aware probe, and labels, it returns a per-example failure taxonomy, a per-modality complementarity matrix that attributes error to modalities, and a loud-vs-silent dropout profile separating monitorable failures from those that pass unflagged far from the decision boundary, using only deployment-observable signals. We release it as a small, unit-tested harness and validate it against planted ground truth. Across seeds it recovers that planted modality dominance and complementary subset, reports per-modality loud-vs-silent rates, and scales to a three-modality complementarity matrix; because the planted structure is known by construction, this validates recovery of per-example attribution rather than clinical performance. We then instantiate the framework on frozen EchoJEPA and HuBERT-ECG embeddings for LVEF and the EF <= 40% HFrEF gate over a paired MIMIC-IV cohort, where on the held-out test split (n = 245) dropping echo nearly doubles error. The narrow echo-to-ECG overlap that bounds cohort size is itself a deployment finding for cardiac foundation models. All of our work can be found at https://github.com/criticaldata/PRIMED-AI.

cs.AI

Suppressed Quantum Effects of Weakly Coupled Waves

Precision experiments increasingly target weakly coupled waves, including axion dark matter and gravitational radiation. Such waves are commonly described as classical fields, yet they could exist in quantum states with no classical counterpart. We exhibit two severe obstructions to detecting nonclassical effects, both independent of the mode occupancy. First, realistic detectors cannot resolve the fundamental modes of a field; instead they couple to coarse-grained "effective" modes, which often washes out nonclassical effects. Second, all nonclassical effects are suppressed by extra powers of the weak coupling, making them much harder to detect than the waves themselves. We prove this in general, and explicitly show how the suppression arises for quadrature and number statistics, entanglement, and decoherence. The suppression can in principle be overcome given suitable quantum resources, such as highly squeezed detector states, but the required parameters are far beyond current experimental capabilities. We use the axion cavity haloscope as an explicit example, although our conclusions apply to many ultralight dark matter searches, and rule out proposals to establish the quantization of gravity from observations of gravitational waves.

hep-ph

Repetition Mismatch: Why Data Mixture Experiments Don't Scale and How to Fix Them

Pre-training data mixtures are commonly tuned by running small-scale experiments and extrapolating to the target training budget. When high-quality data is scarce and must be repeated, this extrapolation frequently fails, but the source of the failure has not been isolated. We show that a primary culprit is a repetition mismatch: because high-quality datasets are small, their repetition rate changes as the training budget grows, shifting the optimal mixture in ways that small-scale proxy experiments do not anticipate. A subsampling procedure that matches the target repetition rate controls for this effect. In a two-source setting combining limited high-quality data with web crawl, a single repetition-controlled experiment using only 1/16 of the target tokens recovers a mixture within 0.10 of the optimum on Wiki-Text for a 1.17B parameter model, compared to an error of 0.85 without repetition control. Achieving comparable accuracy without repetition control requires multiple training horizons, consuming 19%, 44%, and 94% of the target token budget when using the results from two, three, and four horizons respectively. With three data sources, the larger mixture space requires more than a single experiment to constrain, but the approach remains effective: at the 757M scale, just two repetition-controlled horizons recover the optimal mixture, outperforming baselines that instead require the full two-source experiments to construct. Our results reveal that repetition dynamics, not scale alone, shape whether small-scale mixture experiments generalize. More broadly, they suggest that data repetition deserves treatment as a first-class variable in mixture optimization, rather than an inconvenient side effect of limited data.

cs.LG

Hybrid Spatiotemporal Logic for Automotive Applications: Modeling and Model-Checking

We introduce a hybrid spatiotemporal logic for automotive safety applications (HSTL), focused on highway driving. Spatiotemporal logic features specifications about vehicles throughout space and time, while hybrid logic enables precise references to individual vehicles and their historical positions. We define the semantics of HSTL and provide a baseline model-checking algorithm for it. We propose two optimized model-checking algorithms, which reduce the search space based on the reachable states and possible transitions from one state to another. All three model-checking algorithms are evaluated on a series of common driving scenarios such as safe following, safe crossings, overtaking, and platooning. An exponential performance improvement is observed for the optimized algorithms.

cs.LO

Quantum Calculations of the Cavity Shift in Electron Magnetic Moment Measurements

The measurement of the anomalous electron magnetic moment $g-2$ through quantum transitions of a single trapped electron is the most stringent test of quantum field theory. These experiments are now so precise that they must account for the effects of the cavity containing the electron. Classical calculations of this "cavity shift" must subtract the electron's divergent self-field, and thus require knowledge of the exact Green's function for the cavity's electromagnetic field. We perform the first fully quantum calculation of the cavity shift in a closed cavity, which instead involves subtracting linearly divergent cavity mode sums and integrals. Using contour integration methods, we find perfect agreement with existing classical results for both spherical and cylindrical cavities, justifying their current use. Moreover, our mode-based results can be naturally generalized to account for systematic effects, necessary to push future measurements to the next order of magnitude in precision.

hep-ph

Intrinsically Quantum Effects of Axion Dark Matter are Undetectable

Is the usual treatment of axion dark matter as a classical field reliable? We show that the answer is subtle: the axion field could well be in a quantum state that has no complete classical description, but realistic detectors cannot tell the difference. To see this, we solve a fully quantum model of axion detection using quantum optics techniques. We show that intrinsically quantum effects are washed out by mode averaging or small amounts of noise, and significantly suppressed by the weakness of the axion coupling. Our work exemplifies that there should always be a classical analog for axion dark matter effects, extends to other wave (ultralight) dark-matter candidates, and gives a general method to compute the effects of exotic dark-matter states.

hep-ph

Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark

While large language models (LLMs) with reasoning capabilities are progressing rapidly on high-school math competitions and coding, can they reason effectively through complex, open-ended challenges found in frontier physics research? And crucially, what kinds of reasoning tasks do physicists want LLMs to assist with? To address these questions, we present the CritPt (Complex Research using Integrated Thinking - Physics Test, pronounced "critical point"), the first benchmark designed to test LLMs on unpublished, research-level reasoning tasks that broadly covers modern physics research areas, including condensed matter, quantum physics, atomic, molecular & optical physics, astrophysics, high energy physics, mathematical physics, statistical physics, nuclear physics, nonlinear dynamics, fluid dynamics and biophysics. CritPt consists of 71 composite research challenges designed to simulate full-scale research projects at the entry level, which are also decomposed to 190 simpler checkpoint tasks for more fine-grained insights. All problems are newly created by 50+ active physics researchers based on their own research. Every problem is hand-curated to admit a guess-resistant and machine-verifiable answer and is evaluated by an automated grading pipeline heavily customized for advanced physics-specific output formats. We find that while current state-of-the-art LLMs show early promise on isolated checkpoints, they remain far from being able to reliably solve full research-scale challenges: the best average accuracy among base models is only 5.7%, achieved by GPT-5 (high), moderately rising to around 10% when equipped with coding tools. Through the realistic yet standardized evaluation offered by CritPt, we highlight a large disconnect between current model capabilities and realistic physics research demands, offering a foundation to guide the development of scientifically grounded AI tools.

cs.AI

Evaluating Uncertainty Quantification Methods in Argumentative Large Language Models

Research in uncertainty quantification (UQ) for large language models (LLMs) is increasingly important towards guaranteeing the reliability of this groundbreaking technology. We explore the integration of LLM UQ methods in argumentative LLMs (ArgLLMs), an explainable LLM framework for decision-making based on computational argumentation in which UQ plays a critical role. We conduct experiments to evaluate ArgLLMs' performance on claim verification tasks when using different LLM UQ methods, inherently performing an assessment of the UQ methods' effectiveness. Moreover, the experimental procedure itself is a novel way of evaluating the effectiveness of UQ methods, especially when intricate and potentially contentious statements are present. Our results demonstrate that, despite its simplicity, direct prompting is an effective UQ strategy in ArgLLMs, outperforming considerably more complex approaches.

cs.CL

A Prototype Hybrid Mode Cavity for Heterodyne Axion Detection

In the heterodyne approach to axion detection, axion dark matter induces transitions between two modes of a microwave cavity, resulting in a parametrically enhanced signal power. We describe the fabrication and characterization of a prototype normal conducting cavity specifically optimized for heterodyne detection. Corrugations on the cavity walls support linearly polarized hybrid modes which maximize the signal power while strongly suppressing noise. We demonstrate tuning mechanisms which allow one mode's frequency to be scanned across a 4 MHz range, while suppressing cross-coupling noise by at least 80 dB. A future superconducting cavity with identical geometry to our prototype would have the potential to probe orders of magnitude beyond astrophysical bounds.

physics.ins-det

Determining Spin-Dependent Light Dark Matter Rates from Neutron Scattering

The scattering and absorption rates of light dark matter with electron spin-dependent interactions depend on the target's spin response. We show how this response is encoded by the target's dynamical magnetic susceptibility, which can be measured using neutron scattering. We directly use existing neutron scattering data to compute the dark matter scattering rate in a candidate target material, finding close agreement with the previous first-principles calculation at MeV dark matter masses. Complementary experiments and measurements can extend the reach of this technique to other dark matter models and masses, and identify promising target materials for future experiments.

hep-ph

Ponderomotive Effects of Ultralight Dark Matter

I exhibit a new class of quadratic effects of ultralight dark matter. Axions, dark photons, and dilatons can exert rapidly oscillating forces, torques, and mass shifts on Standard Model particles. These effects average to zero at first order, but shift particle properties at second order, in analogy to the ponderomotive force in optics. Remarkably, these effects scale with the square of the amplitude of the dark matter field, even when the field's direct physical effects depend only on its derivatives. I calculate the resulting observables in electron $g_e - 2$ experiments using classical mechanics, recovering results previously derived using field theory. When considered properly, these particular experiments do not beat astrophysical bounds, but other precision experiments may have interesting sensitivity.

hep-ph

The surprising subtlety of electrostatic field lines

Electric fields are commonly visualized with field line diagrams, which only unambiguously specify the field's direction. We consider two simple questions. First, can one deduce if an electric field is conservative, as required e.g. in electrostatics, from its field lines alone? Second, are there conservative electric fields with straight field lines, besides the familiar textbook examples with spherical, cylindrical, or planar symmetry? We give a self-contained introduction to the differential geometry required to answer these questions, assuming only vector calculus background.

physics.class-ph

Query Learning of Advice and Nominal Automata

Learning automata by queries is a long-studied area initiated by Angluin in 1987 with the introduction of the $L^*$ algorithm to learn regular languages, with a large body of work afterwards on many different variations and generalizations of DFAs. Recently, Chase and Freitag introduced a novel approach to proving query learning bounds by computing combinatorial complexity measures for the classes in question, which they applied to the setting of DFAs to obtain qualitatively different results compared to the $L^*$ algorithm. Using this approach, we prove new query learning bounds for two generalizations of DFAs. The first setting is that of advice DFAs, which are DFAs augmented with an advice string that informs the DFA's transition behavior at each step. For advice DFAs, we give the first known upper bounds for query complexity. The second setting is that of nominal DFAs, which generalize DFAs to infinite alphabets which admit some structure via symmetries. For nominal DFAs, we make qualitative improvements over prior results.

cs.FL

Bounds for Rainbow-uncommon Graphs

We say a graph $H$ is $r$-rainbow-uncommon if the maximum number of rainbow copies of $H$ under an $r$-coloring of $E(K_n)$ is asymptotically (as $n \to \infty$) greater than what is expected from uniformly random $r$-colorings. Via explicit constructions, we show that for $H\in\{K_3,K_4, K_5\}$, $H$ is $r$-rainbow-uncommon for all $r\geq {|V(H)|\choose 2}$. We also construct colorings to show that for $t \geq 6$, $K_t$ is $r$-rainbow-uncommon for sufficiently large $r$.

math.CO

Physical Signatures of Fermion-Coupled Axion Dark Matter

In the presence of axion dark matter, fermion spins experience an "axion wind" torque and an "axioelectric" force. We investigate new experimental probes of these effects and find that magnetized analogs of multilayer dielectric haloscopes can explore orders of magnitude of new parameter space for the axion-electron coupling. We also revisit the calculation of axion absorption into in-medium excitations, showing that axioelectric absorption is screened in spin-polarized targets, and axion wind absorption can be characterized in terms of a magnetic energy loss function. Finally, our detailed theoretical treatment allows us to critically examine recent claims in the literature. We find that axioelectric corrections to electronic energy levels are smaller than previously estimated and that the purported electron electric dipole moment due to a constant axion field is entirely spurious.

hep-ph

The Benefits of Being Distributional: Small-Loss Bounds for Reinforcement Learning

While distributional reinforcement learning (DistRL) has been empirically effective, the question of when and why it is better than vanilla, non-distributional RL has remained unanswered. This paper explains the benefits of DistRL through the lens of small-loss bounds, which are instance-dependent bounds that scale with optimal achievable cost. Particularly, our bounds converge much faster than those from non-distributional approaches if the optimal cost is small. As warmup, we propose a distributional contextual bandit (DistCB) algorithm, which we show enjoys small-loss regret bounds and empirically outperforms the state-of-the-art on three real-world tasks. In online RL, we propose a DistRL algorithm that constructs confidence sets using maximum likelihood estimation. We prove that our algorithm enjoys novel small-loss PAC bounds in low-rank MDPs. As part of our analysis, we introduce the $\ell_1$ distributional eluder dimension which may be of independent interest. Then, in offline RL, we show that pessimistic DistRL enjoys small-loss PAC bounds that are novel to the offline setting and are more robust to bad single-policy coverage.

cs.LG

DiffuseExpand: Expanding dataset for 2D medical image segmentation using diffusion models

Dataset expansion can effectively alleviate the problem of data scarcity for medical image segmentation, due to privacy concerns and labeling difficulties. However, existing expansion algorithms still face great challenges due to their inability of guaranteeing the diversity of synthesized images with paired segmentation masks. In recent years, Diffusion Probabilistic Models (DPMs) have shown powerful image synthesis performance, even better than Generative Adversarial Networks. Based on this insight, we propose an approach called DiffuseExpand for expanding datasets for 2D medical image segmentation using DPM, which first samples a variety of masks from Gaussian noise to ensure the diversity, and then synthesizes images to ensure the alignment of images and masks. After that, DiffuseExpand chooses high-quality samples to further enhance the effectiveness of data expansion. Our comparison and ablation experiments on COVID-19 and CGMH Pelvis datasets demonstrate the effectiveness of DiffuseExpand. Our code is released at https://github.com/shaoshitong/DiffuseExpand.

eess.IV