SearcharxivSearch

arXiv subjects

Yufeng Du

Publications and source records attributed to Yufeng Du.

15 recordsLinked to original sources

Beyond Storage: State as a Runtime Control Problem in Parallel and Distributed Systems

Shared state increasingly shapes both performance and failure behavior in streaming, serving, retrieval, and continual-learning systems. Existing studies, however, often isolate access control, hardware-aware execution, memory management, and long-horizon updates. The review organizes this literature around three coupled dimensions: state access and scheduling, state-aware execution, and state evolution and reuse. Across these dimensions, the literature is synthesized through a common scaffold: state object, control surface, coupling path, evaluation boundary, and unresolved contract. This comparison identifies recurring anti-patterns and informs a contract-oriented blueprint and disturbance-aware evaluation agenda. Taken together, the evidence characterizes state management as a runtime control problem.

cs.DC

RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably

We identify intrinsic limitations of Rotary Positional Embeddings (RoPE) in Transformer-based long-context language models. Our theoretical analysis abstracts away from the specific content of the context and depends only on its length. We prove that as context length increases, RoPE-based attention becomes unpredictable and loses two properties that are central to its effectiveness. First, it loses its locality bias: RoPE is no more likely to favor nearer positions than substantially farther ones. Second, it loses consistency in token relevance: a key vector that receives a higher attention score than an alternative at one position may receive a lower score at another. In both cases, the probability of failure approaches 0.5, no better than random guessing. We further prove that the attention score can remain unchanged when a key token is moved to a different position, or even replaced by a different token, indicating a failure to distinguish positions or tokens. Adjusting the RoPE base trades off distinguishing positions against distinguishing tokens but cannot preserve both at the same time. Increasing the RoPE base hyperparameter, a common practice in today's long-context models, helps distinguish different tokens, but inevitably sacrifices the ability to distinguish positions. Our empirical analysis shows that multi-head, multi-layer architectures are insufficient to overcome these limitations. Our findings suggest that fundamentally new mechanisms for encoding position and token order may be needed in future Transformer long-context language models.

cs.CL

Context Length Alone Hurts LLM Performance Despite Perfect Retrieval

Large language models (LLMs) often fail to scale their performance on long-context tasks performance in line with the context lengths they support. This gap is commonly attributed to retrieval failures -- the models' inability to identify relevant information in the long inputs. Accordingly, recent efforts often focus on evaluating and improving LLMs' retrieval performance: if retrieval is perfect, a model should, in principle, perform just as well on a long input as it does on a short one -- or should it? This paper presents findings that the answer to this question may be negative. Our systematic experiments across 5 open- and closed-source LLMs on math, question answering, and coding tasks reveal that, even when models can perfectly retrieve all relevant information, their performance still degrades substantially (13.9%--85%) as input length increases but remains well within the models' claimed lengths. This failure occurs even when the irrelevant tokens are replaced with minimally distracting whitespace, and, more surprisingly, when they are all masked and the models are forced to attend only to the relevant tokens. A similar performance drop is observed when all relevant evidence is placed immediately before the question. Our findings reveal a previously-unrealized limitation: the sheer length of the input alone can hurt LLM performance, independent of retrieval quality and without any distraction. They motivate our simple, model-agnostic mitigation strategy that transforms a long-context task into a short-context one by prompting the model to recite the retrieved evidence before attempting to solve the problem. On RULER, we observe a consistent improvement of GPT-4o up to 4% on an already strong baseline.

cs.CL

Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark

While large language models (LLMs) with reasoning capabilities are progressing rapidly on high-school math competitions and coding, can they reason effectively through complex, open-ended challenges found in frontier physics research? And crucially, what kinds of reasoning tasks do physicists want LLMs to assist with? To address these questions, we present the CritPt (Complex Research using Integrated Thinking - Physics Test, pronounced "critical point"), the first benchmark designed to test LLMs on unpublished, research-level reasoning tasks that broadly covers modern physics research areas, including condensed matter, quantum physics, atomic, molecular & optical physics, astrophysics, high energy physics, mathematical physics, statistical physics, nuclear physics, nonlinear dynamics, fluid dynamics and biophysics. CritPt consists of 71 composite research challenges designed to simulate full-scale research projects at the entry level, which are also decomposed to 190 simpler checkpoint tasks for more fine-grained insights. All problems are newly created by 50+ active physics researchers based on their own research. Every problem is hand-curated to admit a guess-resistant and machine-verifiable answer and is evaluated by an automated grading pipeline heavily customized for advanced physics-specific output formats. We find that while current state-of-the-art LLMs show early promise on isolated checkpoints, they remain far from being able to reliably solve full research-scale challenges: the best average accuracy among base models is only 5.7%, achieved by GPT-5 (high), moderately rising to around 10% when equipped with coding tools. Through the realistic yet standardized evaluation offered by CritPt, we highlight a large disconnect between current model capabilities and realistic physics research demands, offering a foundation to guide the development of scientifically grounded AI tools.

cs.AI

Detecting gravitational signatures of dark matter with atom gradiometers

We study the purely gravitational signatures of dark matter from the ultralight to the ultraheavy mass range in proposed long-baseline atom gradiometers, focusing on terrestrial designs, such as AION-km and MAGIS-km, as well as space-based concepts, such as MAGIS-space, AEDGE and AEDGE+. Due to its exceptional acceleration sensitivity and depending on astrophysical backgrounds, a detector similar to AEDGE+ could detect a dark matter subcomponent which constitutes $\mathcal{O}(10\%)$ of the local dark matter energy density and is populated by compact clumps of mass between $10^6$~kg and $10^{10}$~kg ($10^{-25}~M_\odot\lesssim M \lesssim 10^{-21}~M_\odot$) in an otherwise unexplored region of dark matter model space. Furthermore, because the gravitational observable depends on the relative gravitational time delay measured by spatially separated atomic clouds, we find that atom gradiometers are parametrically more sensitive than laser interferometers, such as LIGO and LISA, to fast-oscillating spacetime perturbations sourced by energy density and pressure fluctuations of ultralight dark matter. Depending on astrophysical backgrounds, a detector akin to AEDGE+ could probe a DM overdensity of $\mathcal{O}(10)$ times the local dark matter energy density for masses $m\lesssim 10^{-17}$~eV.

hep-ph

Physics beyond the Standard Model with the DSA-2000

The upcoming Deep Synoptic Array 2000 (DSA-2000) will map the radio sky at $0.7-2$ GHz ($2.9 - 8.3 \, \mu$eV) with unprecedented sensitivity. This will enable searches for dark matter and other physics beyond the Standard Model, of which we study four cases: axions, dark photons, dark matter subhalos and neutrino masses. We forecast DSA-2000's potential to detect axions through two mechanisms in neutron star magnetospheres: photon conversion of axion dark matter and radio emission from axion clouds, developing the first analytical treatment of the latter. We also forecast DSA-2000's sensitivity to discover kinetically mixed dark photons from black hole superradiance, constrain dark matter substructure and fifth forces through pulsar timing, and improve cosmological neutrino mass inference through fast radio burst dispersion measurements. Our analysis indicates that in its planned five year run the DSA-2000 could reach sensitivity to QCD axion parameters, improve current limits on compact dark matter by an order of magnitude, and enhance cosmological weak lensing neutrino mass constraints by a factor of three.

hep-ph

Signatures of linearized gravity in atom interferometers: A simplified computational framework

We develop a general framework for calculating the leading-order, general relativistic contributions to the gravitational phase shift in single-photon atom interferometers within the context of linearized gravity. We show that the atom gradiometer observable, which only depends on the atom interferometer propagation phase, can be written in terms of three distinct contributions: the Doppler phase shift, which accounts for the tidal displacement of atoms along the baseline, the Shapiro phase shift, which accounts for the delay in the arrival time of photons at atom-light interaction points, and the Einstein phase shift, which accounts for the gravitational redshift measured by the atoms. For specific atom gradiometer configurations, we derive the signal and response functions for two physically motivated scenarios: (i) transient gravitational waves in the transverse-traceless gauge and, for the first time, in the proper detector frame, and (ii) transient massive objects sourcing weak and slow-varying Newtonian potentials. We find that the Doppler contribution of realistic Newtonian noise sources (e.g., a freight truck or a piece of space debris) at proposed atom gradiometer experiments, such as AION, MAGIS and AEDGE, can exceed the shot noise level and thus affect physics searches if not properly subtracted.

gr-qc

EAIRA: Establishing a Methodology for Evaluating AI Models as Scientific Research Assistants

Recent advancements have positioned AI, and particularly Large Language Models (LLMs), as transformative tools for scientific research, capable of addressing complex tasks that require reasoning, problem-solving, and decision-making. Their exceptional capabilities suggest their potential as scientific research assistants but also highlight the need for holistic, rigorous, and domain-specific evaluation to assess effectiveness in real-world scientific applications. This paper describes a multifaceted methodology for Evaluating AI models as scientific Research Assistants (EAIRA) developed at Argonne National Laboratory. This methodology incorporates four primary classes of evaluations. 1) Multiple Choice Questions to assess factual recall; 2) Open Response to evaluate advanced reasoning and problem-solving skills; 3) Lab-Style Experiments involving detailed analysis of capabilities as research assistants in controlled environments; and 4) Field-Style Experiments to capture researcher-LLM interactions at scale in a wide range of scientific domains and applications. These complementary methods enable a comprehensive analysis of LLM strengths and weaknesses with respect to their scientific knowledge, reasoning abilities, and adaptability. Recognizing the rapid pace of LLM advancements, we designed the methodology to evolve and adapt so as to ensure its continued relevance and applicability. This paper describes the methodology state at the end of February 2025. Although developed within a subset of scientific domains, the methodology is designed to be generalizable to a wide range of scientific domains.

cs.AI

Contrast Loss from Astrophysical Backgrounds in Space-Based Matter-Wave Interferometers

Atom and matter interferometers are precise quantum sensing experiments that can probe differential forces along separated spacetime paths. Various atom and matter interferometer experiments have been proposed to study dark matter, gravitational waves, and exotic new physics. Increasingly, these experimental concepts have proposed space-based designs to maximize interrogation times and baselines. However, decoherence and phase shifts caused by astrophysical backgrounds could largely undermine or destroy the target sensitivity of the experiments. We calculate the decoherence effects induced by solar photons, the solar wind, cosmic rays, solar neutrinos and zodiacal dust on space-based atom and matter interferometers. We find that, in future space-based atom and matter interferometers, the solar wind generically produces decoherence beyond the quantum noise limit, without proper shielding. In addition, solar photons are also an important background for matter interferometers.

quant-ph

Atom Interferometer Tests of Dark Matter

Direct detection experiments for dark matter are increasingly ruling out large parameter spaces. However, light dark matter models with particle masses $<$ GeV are still largely unconstrained. Here we examine a proposal to use atom interferometers to detect a light dark matter subcomponent at sub-GeV masses. We describe the decoherence and phase shifts caused by dark matter scattering off of one "arm" of an atom interferometer using a generalized dark matter direct detection framework. This allows us to consider multiple channels: nuclear recoils, hidden photon processes, and axion interactions. We apply this framework to several proposed atom interferometer experiments. Because atom interferometers are sensitive to extremely low momentum deposition and their coherent atoms may give them a boost in sensitivity, these experiments will be highly competitive and complementary to other direct detection methods. In particular, atom interferometers are uniquely able to probe a dark matter sub-component with $m_χ\lesssim 10~\rm{keV}$. We find that, for a mediator mass $m_ϕ=10^{-5}m_χ$, future atom interferometers could close a gap in the existing constraints on nuclear recoils down to $\barσ_n \sim 10^{-42}~\rm{cm}^2$ for $m_χ\sim 10^{-5} - 10^{-1}~\rm{MeV}$ dark matter masses.

hep-ph

SciCode: A Research Coding Benchmark Curated by Scientists

Since language models (LMs) now outperform average humans on many challenging tasks, it has become increasingly difficult to develop challenging, high-quality, and realistic evaluations. We address this issue by examining LMs' capabilities to generate code for solving real scientific research problems. Incorporating input from scientists and AI researchers in 16 diverse natural science sub-fields, including mathematics, physics, chemistry, biology, and materials science, we created a scientist-curated coding benchmark, SciCode. The problems in SciCode naturally factorize into multiple subproblems, each involving knowledge recall, reasoning, and code synthesis. In total, SciCode contains 338 subproblems decomposed from 80 challenging main problems. It offers optional descriptions specifying useful scientific background information and scientist-annotated gold-standard solutions and test cases for evaluation. Claude3.5-Sonnet, the best-performing model among those tested, can solve only 4.6% of the problems in the most realistic setting. We believe that SciCode demonstrates both contemporary LMs' progress towards becoming helpful scientific assistants and sheds light on the development and evaluation of scientific AI in the future.

cs.AI

Macroscopic Dark Matter Detection with Gravitational Wave Experiments

We study signatures of macroscopic dark matter (DM) in current and future gravitational wave (GW) experiments. Transiting DM with a mass of $\sim10^5-10^{15}$ kg that saturates the local DM density can be potentially detectable by GW detectors, depending on the baseline of the detector and the strength of the force mediating the interaction. In the context of laser interferometers, we derive the gauge invariant observable due to a transiting DM, including the Shapiro effect (gravitational time delay accumulated during the photon propagation), and adequately account for the finite photon travel time within an interferometer arm. In particular, we find that the Shapiro effect can be dominant for short-baseline interferometers such as Holometer and GQuEST. We also find that proposed experiments such as Cosmic Explorer and Einstein Telescope can constrain a fifth force between DM and baryons, at the level of strength $\sim 10^3$ times stronger than gravity for, e.g., kg mass DM with a fifth-force range of $10^6$ m.

astro-ph.CO

Quantum Gravity Background in Next-Generation Gravitational Wave Detectors

We study the effects of geontropic vacuum fluctuations in quantum gravity on next-generation terrestrial gravitational wave detectors. If the VZ effect proposed in Ref. [1], as modeled in Refs. [2, 3], appears in the upcoming GQuEST experiment, we show that it will be a large background for astrophysical gravitational wave searches in observatories like Cosmic Explorer and the Einstein Telescope.

gr-qc

Evaluating Modules in Graph Contrastive Learning

The recent emergence of contrastive learning approaches facilitates the application on graph representation learning (GRL), introducing graph contrastive learning (GCL) into the literature. These methods contrast semantically similar and dissimilar sample pairs to encode the semantics into node or graph embeddings. However, most existing works only performed \textbf{model-level} evaluation, and did not explore the combination space of modules for more comprehensive and systematic studies. For effective \textbf{module-level} evaluation, we propose a framework that decomposes GCL models into four modules: (1) a \textbf{sampler} to generate anchor, positive and negative data samples (nodes or graphs); (2) an \textbf{encoder} and a \textbf{readout} function to get sample embeddings; (3) a \textbf{discriminator} to score each sample pair (anchor-positive and anchor-negative); and (4) an \textbf{estimator} to define the loss function. Based on this framework, we conduct controlled experiments over a wide range of architectural designs and hyperparameter settings on node and graph classification tasks. Specifically, we manage to quantify the impact of a single module, investigate the interaction between modules, and compare the overall performance with current model architectures. Our key findings include a set of module-level guidelines for GCL, e.g., simple samplers from LINE and DeepWalk are strong and robust; an MLP encoder associated with Sum readout could achieve competitive performance on graph classification. Finally, we release our implementations and results as OpenGCL, a modularized toolkit that allows convenient reproduction, standard model and module evaluation, and easy extension. OpenGCL is available at \url{https://github.com/thunlp/OpenGCL}.

cs.LG

The Jet-Disk Boundary Layer in Black Hole Accretion

Magnetic fields lines are trapped in black hole event horizons by accreting plasma. If the trapped field lines are lightly loaded with plasma, then their motion is controlled by their footpoints on the horizon and thus by the spin of the black hole. In this paper, we investigate the boundary layer between lightly loaded polar field lines and a dense, equatorial accretion flow. We present an analytic model for aligned prograde and retrograde accretion systems and argue that there is significant shear across this "jet-disk boundary" at most radii for all black hole spins. Specializing to retrograde aligned accretion, where the model predicts the strongest shear, we show numerically that the jet-disk boundary is unstable. The resulting mixing layer episodically loads plasma onto trapped field lines where it is heated, forced to rotate with the hole, and permitted to escape outward into the jet. In one case we follow the mass loading in detail using Lagrangian tracer particles and find a time-averaged mass-loading rate ~ 0.01 Mdot.

astro-ph.HE