SearcharxivSearch

arXiv subjects

Zhenyu Han

Publications and source records attributed to Zhenyu Han.

At least 19 recordsLinked to original sources

MindSpeed RL: Distributed Dataflow for Scalable and Efficient RL Training on Ascend NPU Cluster

Reinforcement learning (RL) is a paradigm increasingly used to align large language models. Popular RL algorithms utilize multiple workers and can be modeled as a graph, where each node is the status of a worker and each edge represents dataflow between nodes. Owing to the heavy cross-node dependencies, the RL training system usually suffers from poor cluster scalability and low memory utilization. In this article, we introduce MindSpeed RL, an effective and efficient system for large-scale RL training. Unlike existing centralized methods, MindSpeed RL organizes the essential data dependencies in RL training, i.e., sample flow and resharding flow, from a distributed view. On the one hand, a distributed transfer dock strategy, which sets controllers and warehouses on the basis of the conventional replay buffer, is designed to release the dispatch overhead in the sample flow. A practical allgather--swap strategy is presented to eliminate redundant memory usage in resharding flow. In addition, MindSpeed RL further integrates numerous parallelization strategies and acceleration techniques for systematic optimization. Compared with existing state-of-the-art systems, comprehensive experiments on the RL training of popular Qwen2.5-Dense-7B/32B, Qwen3-MoE-30B, and DeepSeek-R1-MoE-671B show that MindSpeed RL increases the throughput by 1.42 ~ 3.97 times. Finally, we open--source MindSpeed RL and perform all the experiments on a super pod of Ascend with 384 neural processing units (NPUs) to demonstrate the powerful performance and reliability of Ascend.

cs.LG

AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training

Reinforcement learning (RL) has become a pivotal technology in the post-training phase of large language models (LLMs). Traditional task-colocated RL frameworks suffer from significant scalability bottlenecks, while task-separated RL frameworks face challenges in complex dataflows and the corresponding resource idling and workload imbalance. Moreover, most existing frameworks are tightly coupled with LLM training or inference engines, making it difficult to support custom-designed engines. To address these challenges, we propose AsyncFlow, an asynchronous streaming RL framework for efficient post-training. Specifically, we introduce a distributed data storage and transfer module that provides a unified data management and fine-grained scheduling capability in a fully streamed manner. This architecture inherently facilitates automated pipeline overlapping among RL tasks and dynamic load balancing. Moreover, we propose a producer-consumer-based asynchronous workflow engineered to minimize computational idleness by strategically deferring parameter update process within staleness thresholds. Finally, the core capability of AsynFlow is architecturally decoupled from underlying training and inference engines and encapsulated by service-oriented user interfaces, offering a modular and customizable user experience. Extensive experiments demonstrate an average of 1.59 throughput improvement compared with state-of-the-art baseline. The presented architecture in this work provides actionable insights for next-generation RL training system designs.

cs.LG

Large Language Model-driven Meta-structure Discovery in Heterogeneous Information Network

Heterogeneous information networks (HIN) have gained increasing popularity in recent years for capturing complex relations between diverse types of nodes. Meta-structures are proposed as a useful tool to identify the important patterns in HINs, but hand-crafted meta-structures pose significant challenges for scaling up, drawing wide research attention towards developing automatic search algorithms. Previous efforts primarily focused on searching for meta-structures with good empirical performance, overlooking the importance of human comprehensibility and generalizability. To address this challenge, we draw inspiration from the emergent reasoning abilities of large language models (LLMs). We propose ReStruct, a meta-structure search framework that integrates LLM reasoning into the evolutionary procedure. ReStruct uses a grammar translator to encode the meta-structures into natural language sentences, and leverages the reasoning power of LLMs to evaluate their semantic feasibility. Besides, ReStruct also employs performance-oriented evolutionary operations. These two competing forces allow ReStruct to jointly optimize the semantic explainability and empirical performance of meta-structures. Furthermore, ReStruct contains a differential LLM explainer to generate and refine natural language explanations for the discovered meta-structures by reasoning through the search history. Experiments on eight representative HIN datasets demonstrate that ReStruct achieves state-of-the-art performance in both recommendation and node classification tasks. Moreover, a survey study involving 73 graduate students shows that the discovered meta-structures and generated explanations by ReStruct are substantially more comprehensible. Our code and questionnaire are available at https://github.com/LinChen-65/ReStruct.

cs.LG

Devil in the Landscapes: Inferring Epidemic Exposure Risks from Street View Imagery

Built environment supports all the daily activities and shapes our health. Leveraging informative street view imagery, previous research has established the profound correlation between the built environment and chronic, non-communicable diseases; however, predicting the exposure risk of infectious diseases remains largely unexplored. The person-to-person contacts and interactions contribute to the complexity of infectious disease, which is inherently different from non-communicable diseases. Besides, the complex relationships between street view imagery and epidemic exposure also hinder accurate predictions. To address these problems, we construct a regional mobility graph informed by the gravity model, based on which we propose a transmission-aware graph convolutional network (GCN) to capture disease transmission patterns arising from human mobility. Experiments show that the proposed model significantly outperforms baseline models by 8.54% in weighted F1, shedding light on a low-cost, scalable approach to assess epidemic exposure risks from street view imagery.

cs.CV

How enlightened self-interest guided global vaccine sharing benefits all: a modelling study

Background: Despite the consensus that vaccines play an important role in combating the global spread of infectious diseases, vaccine inequity is still rampant with deep-seated mentality of self-priority. This study aims to evaluate the existence and possible outcomes of a more equitable global vaccine distribution and explore a concrete incentive mechanism that promotes vaccine equity. Methods: We design a metapopulation epidemiological model that simultaneously considers global vaccine distribution and human mobility, which is then calibrated by the number of infections and real-world vaccination records during COVID-19 pandemic from March 2020 to July 2021. We explore the possibility of the enlightened self-interest incentive mechanism, i.e., improving one's own epidemic outcomes by sharing vaccines with other countries, by evaluating the number of infections and deaths under various vaccine sharing strategies using the proposed model. To understand how these strategies affect the national interests, we distinguish the imported and local cases for further cost-benefit analyses that rationalize the enlightened self-interest incentive mechanism behind vaccine sharing. ...

q-bio.PE

Multi-Scale Simulation of Complex Systems: A Perspective of Integrating Knowledge and Data

Complex system simulation has been playing an irreplaceable role in understanding, predicting, and controlling diverse complex systems. In the past few decades, the multi-scale simulation technique has drawn increasing attention for its remarkable ability to overcome the challenges of complex system simulation with unknown mechanisms and expensive computational costs. In this survey, we will systematically review the literature on multi-scale simulation of complex systems from the perspective of knowledge and data. Firstly, we will present background knowledge about simulating complex system simulation and the scales in complex systems. Then, we divide the main objectives of multi-scale modeling and simulation into five categories by considering scenarios with clear scale and scenarios with unclear scale, respectively. After summarizing the general methods for multi-scale simulation based on the clues of knowledge and data, we introduce the adopted methods to achieve different objectives. Finally, we introduce the applications of multi-scale simulation in typical matter systems and social systems.

eess.SY

Strategic COVID-19 vaccine distribution can simultaneously elevate social utility and equity

Balancing social utility and equity in distributing limited vaccines represents a critical policy concern for protecting against the prolonged COVID-19 pandemic. What is the nature of the trade-off between maximizing collective welfare and minimizing disparities between more and less privileged communities? To evaluate vaccination strategies, we propose a novel epidemic model that explicitly accounts for both demographic and mobility differences among communities and their association with heterogeneous COVID-19 risks, then calibrate it with large-scale data. Using this model, we find that social utility and equity can be simultaneously improved when vaccine access is prioritized for the most disadvantaged communities, which holds even when such communities manifest considerable vaccine reluctance. Nevertheless, equity among distinct demographic features are in tension due to their complex correlation in society. We design two behavior-and-demography-aware indices, community risk and societal harm, which capture the risks communities face and those they impose on society from not being vaccinated, to inform the design of comprehensive vaccine distribution strategies. Our study provides a framework for uniting utility and equity-based considerations in vaccine distribution, and sheds light on how to balance multiple ethical values in complex settings for epidemic control.

cs.CY

Policy-Aware Mobility Model Explains the Growth of COVID-19 in Cities

With the continued spread of coronavirus, the task of forecasting distinctive COVID-19 growth curves in different cities, which remain inadequately explained by standard epidemiological models, is critical for medical supply and treatment. Predictions must take into account non-pharmaceutical interventions to slow the spread of coronavirus, including stay-at-home orders, social distancing, quarantine and compulsory mask-wearing, leading to reductions in intra-city mobility and viral transmission. Moreover, recent work associating coronavirus with human mobility and detailed movement data suggest the need to consider urban mobility in disease forecasts. Here we show that by incorporating intra-city mobility and policy adoption into a novel metapopulation SEIR model, we can accurately predict complex COVID-19 growth patterns in U.S. cities ($R^2$ = 0.990). Estimated mobility change due to policy interventions is consistent with empirical observation from Apple Mobility Trends Reports (Pearson's R = 0.872), suggesting the utility of model-based predictions where data are limited. Our model also reproduces urban "superspreading", where a few neighborhoods account for most secondary infections across urban space, arising from uneven neighborhood populations and heightened intra-city churn in popular neighborhoods. Therefore, our model can facilitate location-aware mobility reduction policy that more effectively mitigates disease transmission at similar social cost. Finally, we demonstrate our model can serve as a fine-grained analytic and simulation framework that informs the design of rational non-pharmaceutical interventions policies.

q-bio.PE

Genetic Meta-Structure Search for Recommendation on Heterogeneous Information Network

In the past decade, the heterogeneous information network (HIN) has become an important methodology for modern recommender systems. To fully leverage its power, manually designed network templates, i.e., meta-structures, are introduced to filter out semantic-aware information. The hand-crafted meta-structure rely on intense expert knowledge, which is both laborious and data-dependent. On the other hand, the number of meta-structures grows exponentially with its size and the number of node types, which prohibits brute-force search. To address these challenges, we propose Genetic Meta-Structure Search (GEMS) to automatically optimize meta-structure designs for recommendation on HINs. Specifically, GEMS adopts a parallel genetic algorithm to search meaningful meta-structures for recommendation, and designs dedicated rules and a meta-structure predictor to efficiently explore the search space. Finally, we propose an attention based multi-view graph convolutional network module to dynamically fuse information from different meta-structures. Extensive experiments on three real-world datasets suggest the effectiveness of GEMS, which consistently outperforms all baseline methods in HIN recommendation. Compared with simplified GEMS which utilizes hand-crafted meta-paths, GEMS achieves over $6\%$ performance gain on most evaluation metrics. More importantly, we conduct an in-depth analysis on the identified meta-structures, which sheds light on the HIN based recommender system design.

cs.IR

Construction of a Single-Column Model in RegCM4 and its Preliminary Application for Evaluating PBL Schemes in Simulating the Dry Convection Boundary Layer

A single-column model (SCM) is constructed in the regional climate model RegCM4. The evolution of a dry convection boundary layer (DCBL) is used to evaluate this SCM and compare four planetary boundary layer (PBL) schemes, the Holtslag-Boville scheme (HB), Yonsei University scheme (YSU), and two University of Washington schemes (UW01, Grenier-Bretherton-McCaa scheme and UW09, Bretherton-Park scheme), using the SCM approach. A large-eddy simulation (LES) of the DCBL is performed as a benchmark to examine how well a PBL parameterization scheme reproduces the LES results, and several diagnostic outputs are compared to evaluate the schemes. In general, with the DCBL case, the YSU scheme performs best for reproducing the LES results, which include well-mixed features and vertical sensible heat fluxes; UW09 has the second best performance, UW01 has the third best performance, and the HB scheme has the worst performance. The results show that the SCM is proper constructed. Although more cases and further testing are required, these simulations show encouraging results towards the use of this SCM framework for studying the physical processes in RegCM4.

physics.ao-ph

Top-Tagging at the Energy Frontier

At proposed future hadron colliders and in the coming years at the LHC, top quarks will be produced at genuinely multi-TeV energies. Top-tagging at such high energies forces us to confront several new issues in terms of detector capabilities and jet physics. Here, we explore these issues in the context of some simple JHU/CMS-type declustering algorithms and the N-subjettiness jet-shape variable tau_32. We first highlight the complementarity between the two tagging approaches at particle-level with respect to discriminating top-jets against gluons and quarks, using multivariate optimization scans. We then introduce a basic fast detector simulation, including electromagnetic calorimeter showering patterns determined from GEANT. We consider a number of tricks for processing the fast detector output back to an approximate particle-level picture. Re-optimizing the tagger parameters, we demonstrate that the inevitable losses in discrimination power at very high energies can typically be ameliorated. For example, percent-scale mistag rates might be maintained even in extreme cases where an entire top decay would sit inside of one hadronic calorimeter cell and tracking information is completely absent. We then study three novel physics effects that will come up in the multi-TeV energy regime: gluon radiation off of boosted top quarks, mistags originating from g -> tt, and mistags originating from q -> (W/Z)q collinear electroweak splittings with subsequent hadronic decays. The first effect, while nominally a nuisance, can actually be harnessed to slightly improve discrimination against gluons. The second effect can lead to effective O(1) enhancements of gluon mistag rates for tight working points. And the third effect, while conceptually interesting, we show to be of highly subleading importance at all energies.

hep-ph

$J_{E_T}^{\rm II}$: A Two-prong Jet Finding Algorithm

We propose a new global jet-finding algorithm for reconstructing two-prong objects like hadronic weak gauge bosons at a hadron collider. The selection of particles in a two-prong jet is required to maximize a $J_{E_T}^{\rm II}$ function, which contains a modified second Fox-Wolfram moment and prefers a two-prong structure for a fixed jet mass. Compared to the traditional jet-substructure method, our algorithm can provide a similar or better performance for identifying boosted weak gauge bosons that are produced from either Standard Model processes or heavy di-boson resonance decays.

hep-ph

MT2 to the Rescue -- Searching for Sleptons in Compressed Spectra at the LHC

We propose a novel method for probing sleptons in compressed spectra at hadron colliders. The process under study is slepton pair production in R-parity conserving supersymmetry, where the slepton decays to a neutralino LSP of mass close to the slepton mass. In order to pass the trigger and obtain large missing energy, an energetic mono-jet is required. Both leptons need to be detected in order to suppress large standard model backgrounds with one charged lepton. We study variables that can be used to distinguish the signal from the remaining major backgrounds, which include tt, WW+jet, Z+jet, and single top production. We find that the dilepton MT2, bound by the mass difference, can be used as an upper bound to efficiently reduce the backgrounds. It is estimated that sleptons with masses up to about 150 GeV can be discovered at the 14 TeV LHC with 100/fb integrated luminosity.

hep-ph

J_{E_T}: A Global Jet Finding Algorithm

We introduce a new jet-finding algorithm for a hadron collider based on maximizing a J_{E_T} function for all possible combinations of particles in an event. This function prefers a larger value of the jet transverse energy and a smaller value of the jet mass. The jet shape is proved to be a circular cone in Cartesian coordinates with the geometric center shifted from the jet momentum toward the central region. The jet cone size shrinks for a more forward jet. We have implemented our J_{E_T} algorithm with a reasonable running time scaling as N n^3, where "N" is the total number of particles and "n" (much less than N) is the number of particles in a fiducial region. Many features of our J_{E_T} jets are similar to anti-k_t jets, including the reconstructed jet momentum and the "back-reaction" from soft contamination. Nevertheless, when the jet parameters in the two algorithms are matched using QCD jets, we find that the J_{E_T} algorithm has a larger efficiency than anti-k_t for identifying objects with hard splittings such as a W-jet.

hep-ph

New light on WW scattering at the LHC with W jet tagging

After the recent discovery of a 125 GeV Higgs-like particle at the Large Hadron Collider (LHC), it is crucial to examine its role in unitarizing high energy W_LW_L scattering, which may reveal its possible deviation from a Standard Model Higgs. We perform an updated study on WW scattering in the semileptonic channel at the LHC, improved by the recently developed W jet tagging method. The resultant statistical significance of a Strongly-Interacting Light Higgs (SILH) model is about 20% larger than that based on the conventionally "gold-plated" dileptonic channel, while 200% more signal events are retained. This allows the discovery of an SILH model if the signal strength is 100% (10%) of a pure Higgsless model, using about 40 (3000) fb^{-1} data at the 14 TeV LHC. Meanwhile, the excellent sensitivity to the anomalous Higgs-W boson coupling makes semileptonic WW scattering an important complement to precision measurements at the Higgs resonance.

hep-ph

Jet Radiation Radius

Jet radiation patterns are indispensable for the purpose of discriminating partons' with different quantum numbers. However, they are also vulnerable to various contaminations from the underlying event, pileup, and radiation of adjacent jets. In order to maximize the discrimination power, it is essential to optimize the jet radius used when analyzing the radiation patterns. We introduce the concept of jet radiation radius which quantifies how the jet radiation is distributed around the jet axes. We study the color and momentum dependence of the jet radiation radius, and discuss two applications: quark-gluon discrimination and $W$ jet tagging. In both cases, smaller (sub)jet radii are preferred for jets with higher PTs, albeit due to different mechanisms: the running of the QCD coupling constant and the boost to a color singlet system. A shrinking cone W jet tagging algorithm is proposed to achieve better discrimination than previous methods.

hep-ph

Hunting Quasi-Degenerate Higgsinos

We present a new strategy to uncover light, quasi-degenerate Higgsinos, a likely ingredient in a natural supersymmetric model. Our strategy focuses on Higgsinos with inter-state splittings of O(5-50) GeV that are produced in association with a hard, initial state jet and decay via off-shell gauge bosons to two or more leptons and missing energy, $pp \to j + \text{MET}\, + 2^+\, \ell$. The additional jet is used for triggering, allowing us to significantly loosen the lepton requirements and gain sensitivity to small inter-Higgsino splittings. Focusing on the two-lepton signal, we find the seemingly large backgrounds from diboson plus jet, $\bar tt$ and $Z/γ^* + j$ can be reduced with careful cuts, and that fake backgrounds appear minor. For Higgsino masses $m_χ$ just above the current LEP II bound ($μ\simeq 110\,$) GeV we find the significance can be as high as 3 sigma at the LHC using the existing 20 fb$^{-1}$ of 8 TeV data. Extrapolating to LHC at 14 TeV with 100 fb$^{-1}$ data, and as one example $M_1 = M_2 = 500$ GeV, we find 5 sigma evidence for $m_χ \lesssim\, 140\,$ GeV and 2 sigma evidence for $m_χ \lesssim\, 200\,$ GeV . We also present a reinterpretation of ATLAS/CMS monojet bounds in terms of degenerate Higgsino ($δm_χ \ll 5\,$) GeV plus jet production. We find the current monojet bounds on $m_χ$ are no better than the chargino bounds from LEP II.

hep-ph

Stealth Stops and Spin Correlation: A Snowmass White Paper

Stops with the mass nearly degenerate with the top mass, decaying into tops and soft neutralinos, are usually dubbed "stealth stops". Their kinematics looks very similar to that of the standard tops events, which leads to events with little or no excess of missing transverse energy. This complicates the probing of this region of the stop parameter space by hadron colliders, rendering the application of standard searching techniques challenging. In this Snowmass white paper we reanalyze the spin correlation approach to the search of stealth stops, focusing on the feasibility of this search at the 14 TeV LHC. We find, while the statistical limitations significantly shrink compared to the low-luminosity 8 TeV run, the systematic PDF uncertainties pose the main obstacle. We show that the current understanding of PDFs probably does not allow us to talk about top and stop discrimination via spin correlation in the inclusive sample. On the other hand the systematic uncertainties significantly shrink if only events with low center of mass energy are considered, rendering the search in this region feasible.

hep-ph