SearcharxivSearch

arXiv subjects

Rui Wei

Publications and source records attributed to Rui Wei.

12 recordsLinked to original sources

The Diminishing Returns of Early-Exit Decoding in Modern LLMs

In Large Language Model (LLM) inference, early-exit refers to stopping computation at an intermediate layer once the prediction is sufficiently confident, thereby reducing latency and cost. However, recent LLMs adopt improved pretraining recipes and architectures that reduce layer redundancy, potentially limiting early-exit opportunities. We re-evaluate layer-wise early-exit in modern LLMs and analyze how intermediate representations evolve during training. We introduce a metric to quantify a model's intrinsic suitability for early-exit and propose a benchmark for researchers to explore the potential early-exit benefits on different models and workloads. Our results show a diminishing trend in early-exit effectiveness across newer model generations. We further find that dense transformers generally offer greater early-exit potential than Mixture-of-Experts and State Space Models. In addition, larger models, particularly those with more than 20 billion parameters, and base pretrained models without specialized tuning tend to exhibit higher early-exit potential.

cs.CL

RLHFless: Serverless Computing for Efficient RLHF

Reinforcement Learning from Human Feedback (RLHF) has been widely applied to Large Language Model (LLM) post-training to align model outputs with human preferences. Recent models, such as DeepSeek-R1, have also shown RLHF's potential to improve LLM reasoning on complex tasks. In RL, inference and training co-exist, creating dynamic resource demands throughout the workflow. Compared to traditional RL, RLHF further challenges training efficiency due to expanding model sizes and resource consumption. Several RLHF frameworks aim to balance flexible abstraction and efficient execution. However, they rely on serverful infrastructures, which struggle with fine-grained resource variability. As a result, during synchronous RLHF training, idle time between or within RL components often causes overhead and resource wastage. To address these issues, we present RLHFless, the first scalable training framework for synchronous RLHF, built on serverless computing environments. RLHFless adapts to dynamic resource demands throughout the RLHF pipeline, pre-computes shared prefixes to avoid repeated computation, and uses a cost-aware actor scaling strategy that accounts for response length variation to find sweet spots with lower cost and higher speed. In addition, RLHFless assigns workloads efficiently to reduce intra-function imbalance and idle time. Experiments on both physical testbeds and a large-scale simulated cluster show that RLHFless achieves up to 1.35x speedup and 44.8% cost reduction compared to the state-of-the-art baseline.

cs.AI

High-Capacity Metasurface at Limits of Polarization and Wavelength Multiplexing

Polarization and wavelength multiplexing are the two most widely employed techniques to improve the capacity in the metasurfaces. Existing works have pushed each technique to its individual limits. For example, the polarization multiplexing channels working at a single wavelength have been significantly increased by using noise engineering. However, it is still challenging to achieve the multiplexing limits of wavelength and polarization simultaneously. Besides, such multiplexing methods suffer from computational inefficiencies, hindering their application in tasks like image recognition that require extensive training computation. In this work, we introduce a gradient-based optimization algorithm using deep neural network (DNN) to achieve the limits of both polarization and wavelength multiplexing with high computational efficiency. We experimentally demonstrate this capability, achieving a record-breaking capacity of 15 holographic images across five wavelengths and the maximum of three independent polarization channels, as well as 18 holographic images across three wavelengths and six corelated polarization channels. Moreover, leveraging the high computational efficiency of our DNN-based method, which is well-suited for processing large datasets, we implement large-scale image recognition tasks across 36 classes encoded in a record of nine multiplexed channels (three wavelengths * three polarizations), achieving 96% classification accuracy in calculations and 91.5% in experiments. This work sets a new benchmark for high-capacity multiplexing with metasurfaces and demonstrates the power of gradient-based inverse design for realizing multi-functional optical elements.

physics.optics

Metasurface-Based Full-Parameter Optical Multiplexing

Optical multiplexing is a key technique that enhances the capacity of optical systems by independently modulating various optical parameters to carry distinct information. Among these parameters, wavelength, polarization, and angle are the primary ones for multiplexing in plane waves with uniform cross-sectional distribution. While metasurfaces have recently emerged as a powerful platform for optical multiplexing, they are typically restricted to partial parameter multiplexing and exhibit a low number of multiplexing channels. In this work, we propose and experimentally demonstrate the full-parameter multiplexing of polarization, wavelength, and angle, achieving hundreds of distinct multiplexing channels,the largest reported to date. Our design utilizes a gradient-based optimization algorithm to enable high-efficiency performance and independent functionalities with minimal cross-talk among channels. This approach represents a significant advancement in metasurface design and optical multiplexing, with potential applications in complex and dynamic optical systems.

physics.optics

Glauber-based evaluations of the odd moments of the initial eccentricity relative to the even order participant planes

Monte Carlo simulations are used to compute the centrality dependence of the odd moments of the initial eccentricity $ε_{n+1}$, relative to the even order (n) participant planes $Ψ^*_n$ in Au+Au collisions. The results obtained for two models of the eccentricity -- the Glauber and the factorized Kharzeev-Levin-Nardi (fKLN) models -- indicate magnitudes which are essentially zero. They suggest that a possible correlation between the orientations of the the odd and even participant planes ($Ψ^*_{n+1}$ and $Ψ^*_n$ respectively), do not have a significant influence on the calculated eccentricities. An experimental verification test for correlations between the orientations of the the odd and even participant planes is also proposed.

nucl-ex

Initial eccentricity fluctuations and their relation to higher-order flow harmonics

Monte Carlo simulations are used to compute the centrality dependence of the participant eccentricities ($ε_{n}$) in Au+Au collisions, for the two primary models currently employed for eccentricity estimates -- the Glauber and the factorized Kharzeev-Levin-Nardi (fKLN) models. They suggest specific testable predictions for the magnitude and centrality dependence of the flow coefficients $v_n$, respectively measured relative to the event planes $Ψ_n$. They also indicate that the ratios of several of these coefficients may provide an additional constraint for distinguishing between the models. Such a constraint could be important for a more precise determination of the specific viscosity of the matter produced in heavy ion collisions.

nucl-ex

Dissecting the role of initial collision geometry for jet quenching observables in relativistic heavy ion collisions

The observation of large azimuthal anisotropy or $v_2$ for hadrons above $p_T>5$ GeV/$c$ in Au+Au collisions at $\sqrt{s_{\rm nn}}=200$ GeV has been a longstanding challenge for jet quenching models based on perturbative QCD (pQCD). Using a simple jet absorption model, we seek to clarify the situation by exploring in detail how the calculated $v_2$ varies with choices of the collision geometry as well as choices of the path length dependence and thermalization time $τ_0$ in the energy loss formula. Besides the change of eccentricity due to distortion from gluon saturation or event-by-event fluctuation, we find that the $v_2$ is also sensitive to the centrality dependence of multiplicity and the relative size between the matter profile and the jet profile. We find that the $v_2$ calculated for the naive quadratic path length dependence of energy loss, even including eccentricity fluctuation and the gluon saturation, is not enough to describe the experimental data at high $p_T$ ($\sim$ 6 GeV/$c$) in Au+Au collisions. However, it can match the full centrality dependence of $v_2$ data if higher power path length dependence of energy loss is allowed. We also find that the calculated $v_2$ is sensitive to the assumption of the early time dynamics but generally increases with $τ_0$, opposite to what one expects for elliptic flow. This study attests to the importance of confining the initial geometry, possibly by combining jet quenching $v_2$ with elliptic flow and other jet quenching observables, for proper interpretation of the experimental data.

nucl-th

Constraints on models for the initial collision geometry in ultra relativistic heavy ion collisions

Monte Carlo (MC) simulations are used to compute the centrality dependence of the collision zone eccentricities ($ε_{2,4}$), for both spherical and deformed ground state nuclei, for different model scenarios. Sizable model dependent differences are observed. They indicate that measurements of the $2^{\text{nd}}$ and $4^{\text{th}}$ order Fourier flow coefficients $v_{2,4}$, expressed as the ratio $\frac{v_4}{(v_2)^2}$, can provide robust constraints for distinguishing between different theoretical models for the initial-state eccentricity. Such constraints could remove one of the largest impediments to a more precise determination of the specific viscosity from precision $v_{2,4}$ measurements at the Relativistic Heavy Ion Collider (RHIC).

nucl-ex

Scaling patterns of the suppression of $π^0$ yields in Au+Au collisions at $\sqrt{s_{NN}}=200$ GeV: links to the transport properties of the QGP

Suppression measurements for neutral pions ($π^0$) are used to investigate the predicted path length ($L$) and transverse momentum ($p_T$) dependent jet quenching patterns of the hot QCD medium produced in Au+Au collisions at $\sqrt{s_{NN}}=200$ GeV. The observed scaling patterns show the predicted trends for jet-medium interactions dominated by radiative energy loss. They also allow simple estimates of the transport coefficient $\hat{q}$ and the ratio of viscosity to entropy density $η/s$. These estimates indicate that the short mean free path ($λ$) in the QCD medium leading to hydrodynamic-like flow with a small value of $η/s$, is also responsible for the strong suppression observed.

nucl-ex

PHENIX Measurements of Azimuthal Anisotropy for $π^0$ Production at High $p_T$ in Au+Au Collisions at $\sqrt{s_{NN}}=200$GeV

An improved measurement of $v_2$ for $π^0$ in a broad range of $p_T$ and centrality is presented. By combining $v_2$ with the $R_{AA}$, we provide new insights on jet-medium interactions. We show that current pQCD energy loss models cannot describe the suppression of the $π^0$ as a function of the angle with respect to the reaction plane. Our result could help to resolve the factor of 4 differences in the predicted transport coefficients among these models. Alternatively, it may suggest that non-perturbative effects associated with the strongly coupled QGP are important, and new theoretical developments are needed to fully understand the jet medium interactions.

nucl-ex

Away-side asymmetry of jet correlation relative to reaction plane: a sensitive probe for jet in-medium modifications

We proposed a new observable based on two particle azimuth correlation to study the away-side medium response in mid-central Au+Au collisions. We argue that a left/right asymmetry may appear at the away-side by selecting triggers separately in the left and right side of the reaction plane. A simple model estimation suggests that the magnitude of such asymmetry could reach 30% with details depends on the medium response mechanisms. This asymmetry, if observed, can help to distinguish competing theoretical models.

hep-ph

Is the quark gluon plasma produced in RHIC collisions strongly coupled?

Recent hexadecapole (v4) and elliptic (v2) flow measurements are used to constrain estimates for the degree of local equilibrium, mean free path $λ$, and the viscosity to entropy density ratio (eta/s) of the plasma produced in Au+Au collisions at RootS = 200 GeV. The eccentricity-scaled flow coefficients v2/e2 and v4/e4 indicate that the plasma achieves a degree of local equilibrium within 5 - 10% of the value expected for a fluid with eta/s equal to the conjectured lower bound of 1/4pi. Estimates for $λ$ and eta/s as a function of collision centrality and particle transverse momentum pT, points to transverse expansion dynamics compatible with a strongly coupled low viscosity plasma.

nucl-ex