SearcharxivSearch

arXiv subjects

Hong Tian

Publications and source records attributed to Hong Tian.

9 recordsLinked to original sources

MemFine: Memory-Aware Fine-Grained Scheduling for MoE Training

The training of large-scale Mixture of Experts (MoE) models faces a critical memory bottleneck due to severe load imbalance caused by dynamic token routing. This imbalance leads to memory overflow on GPUs with limited capacity, constraining model scalability. Existing load balancing methods, which cap expert capacity, compromise model accuracy and fail on memory-constrained hardware. To address this, we propose MemFine, a memory-aware fine-grained scheduling framework for MoE training. MemFine decomposes the token distribution and expert computation into manageable chunks and employs a chunked recomputation strategy, dynamically optimized through a theoretical memory model to balance memory efficiency and throughput. Experiments demonstrate that MemFine reduces activation memory by 48.03% and improves throughput by 4.42% compared to full recomputation-based baselines, enabling stable large-scale MoE training on memory-limited GPUs.

cs.DC

Statistically Accurate and Robust Generative Prediction of Rock Discontinuities with A Tabular Foundation Model

Rock discontinuities critically govern the mechanical behavior and stability of rock masses. Their internal distributions remain largely unobservable and are typically inferred from surface-exposed discontinuities using generative prediction approaches. However, surface-exposed observations are inherently sparse, and existing generative prediction approaches either fail to capture the underlying complex distribution patterns or lack robustness under data-sparse conditions. Here, we proposed a simple yet robust approach for statistically accurate generative prediction of rock discontinuities by utilizing a tabular foundation model. By leveraging the powerful sample learning capability of the foundation model specifically designed for small data, our approach can effectively capture the underlying complex distribution patterns within limited measured discontinuities. Comparative experiments on ten datasets with diverse scales and distribution patterns of discontinuities demonstrate superior accuracy and robustness over conventional statistical models and deep generative approaches. This work advances quantitative characterization of rock mass structures, supporting safer and more reliable data-driven geotechnical design.

cs.LG

Quantifying Nonradiative Recombination and Resistive Losses in Perovskite Photovoltaics: A Modified Diode Model Approach

Pinpointing the origin of inefficiency can expedite the process of optimizing the efficiency of perovskite photovoltaics. However, it is challenging to discern and quantify the different loss pathways in a complete perovskite photovoltaic device under operational conditions. To address this challenge, we propose a modified diode model that can quantify bulk/interface defect-assisted recombination and series/shunt resistive losses. By adopting drift-diffusion simulation as the benchmark, we explore the physical meanings of the modified diode model parameters and evaluate the performance of the model for simulation parameters spanning many orders of magnitude. Our evaluation shows that, in most practical cases, the proposed model can accurately quantify all the aforementioned losses, and in some special cases, it is possible to identify the predominant loss pathway. Moreover, we apply the modified diode model to our lab-produced devices (based on Cs0.05FA0.95PbI3 perovskites), demonstrating its effectiveness in quantifying entangled losses in practice. Finally, we provide a set of guidelines for applying the modified diode model and interpreting the results. Source code available at https://github.com/WPT-Lab124/Modified-Diode-Model.

physics.app-ph

Considerations for Master Protocols Using External Controls

There has been an increasing use of master protocols in oncology clinical trials because of its efficiency and flexibility to accelerate cancer drug development. Depending on the study objective and design, a master protocol trial can be a basket trial, an umbrella trial, a platform trial, or any other form of trials in which multiple investigational products and/or subpopulations are studied under a single protocol. Master protocols can use external data and evidence (e.g., external controls) for treatment effect estimation, which can further improve efficiency of master protocol trials. This paper provides an overview of different types of external controls and their unique features when used in master protocols. Some key considerations in master protocols with external controls are discussed including construction of estimands, assessment of fit-for-use real-world data, and considerations for different types of master protocols. Similarities and differences between regular randomized controlled trials and master protocols when using external controls are discussed. A targeted learning-based causal roadmap is presented which constitutes three key steps: (1) define a target statistical estimand that aligns with the causal estimand for the study objective, (2) use an efficient estimator to estimate the target statistical estimand and its uncertainty, and (3) evaluate the impact of causal assumptions on the study conclusion by performing sensitivity analyses. Two illustrative examples for master protocols using external controls are discussed for their merits and possible improvement in causal effect estimation.

stat.AP

Principles of Conditionality and Layering of Error Rates with Application to Platform Trials

There has been a misconception that only one type of error rate control is necessary in clinical trials, leading to debates over whether to prioritize Familywise Error Rate (FWER) or False Discovery Rate (FDR). This misconception has led to misleading statements about FWER control and proposals to shift towards FDR control, which could be manipulated by the industry. In reality, since the early 2000s, biopharmaceutical statistics have implicitly applied two layers of Type I error rate control. This aligns with Tukey's 1953 invention of Error Rate per Family (ERpF) for controlling error across studies, while FWER applies within each study. Our paper clarifies this layering, using Platform trials to demonstrate the verifiable conditions needed across studies for the FDA to fulfill its regulatory mission. We show that controlling FWER within a study at $5\%$ inherently controls ERpF across studies at 5-per-100, regardless of study correlations. This supports current regulatory practices that protect public health while fostering innovation. We also address concerns about ERpF stability in Platform trials, where shared controls introduce dependencies. By applying the Conditionality Principle and utilizing an innovative Shiny app, we explore how correlations impact ERpF variability, providing deeper insights for informed decision-making. Our findings, supported by principles like Layering of Error Rate Controls and the Conditionality Principle, are particularly relevant as Platform trials gain popularity for their efficiency in testing multiple treatments simultaneously.

stat.ME

Remarks on a Paper by Leonetti and Siepe

In 2012, F.Leonetti and F.Siepe [1] considered solutions to boundary value problems of some anisotropic elliptic equations of the type $$ \left\{ \begin{array}{llll} \sum\limits _{i=1}\limits^{n} D_i (a_i(x,Du(x)))=0, &x\in Ω,\\ u(x)=θ(x), & x\in \partial Ω. \end{array} \right. $$ Under some suitable conditions, they obtained an integrability result, which shows that, higher integrability of the boundary datum $θ$ forces solutions $u$ to have higher integrability as well. In the present paper, we consider ${\cal K}_{ψ,θ}^{(p_i)}$-obstacle problems of the nonhomogeneous anisotropic elliptic equations $$ \sum_{i=1}^n D_i (a_i(x,Du(x)))=\sum_{i=1}^n D_i f^i(x). $$ Under some controllable growth and monotonicity conditions. We obtain an integrability result, which can be regarded as a generalization of the result due to Leonetti and Siepe.

math.AP

A Generalization of Exponential Class and its Applications

A function space, $L^{θ,\infty)}(Ω)$, $0 \leq θ<\infty$, is defined. It is proved that $L^{θ,\infty)}(Ω)$ is a Banach space which is a generalization of exponential class. An alternative definition of $L^{θ,\infty)}(Ω)$ space is given. As an application, we obtain weak monotonicity property for very weak solutions of $\mathcal{A}$-harmonic equation with variable coefficients under some suitable conditions related to $L^{θ,\infty)}(Ω)$, which provides a generalization of a known result due to Moscariello. A weighted space $L^{θ,\infty)}_w(Ω)$) is also defined, and the boundedness for the Hardy-Littlewood maximal operator $M_w$ and a Calderón-Zygmund operator $T$ with respect to $L^{θ,\infty)}_w(Ω)$ are obtained.

math.AP

A Parallel Solution to Finding Nodal Neighbors in Generic Meshes

In this paper we specifically present a parallel solution to finding the one-ring neighboring nodes and elements for each vertex in generic meshes. The finding of nodal neighbors is computationally straightforward but expensive for large meshes. To improve the efficiency, the parallelism is adopted by utilizing the modern Graphics Processing Unit (GPU). The presented parallel solution is heavily dependent on the parallel sorting, scan, and reduction, and can be applied to determine both the neighboring nodes and elements. To evaluate the performance, the parallel solution is compared to the corresponding serial solution. Experimental results show that: our parallel solution can achieve the speedups of approximately 55 and 90 over the corresponding serial solution for finding neighboring nodes and elements, respectively. Our parallel solution is efficient and easy to implement, but requires the allocation of large device memory.

cs.CG

Performance Impact of Data Layout on the GPU-accelerated IDW Interpolation

This paper focuses on evaluating the performance impact of different data layouts on the GPU-accelerated IDW interpolation. First, we redesign and improve our previous GPU implementation that was performed by exploiting the feature CUDA Dynamic Parallel (CDP). And then, we implement three versions of GPU implementations, i.e., the naive version, the tiled version, and the improved CDP version, based on five layouts including the Structure of Arrays (SoA), the Array of Sturcutes (AoS), the Array of aligned Sturcutes (AoaS), the Structure of Arrays of aligned Structures (SoAoS), and the Hybrid layout. Experimental results show that: the layouts AoS and AoaS achieve better performance than the layout SoA for both the naive version and tiled version, while the layout SoA is the best choice for the improved CDP version. We also observe that: for the two combined data layouts (the SoAoS and the Hybrid), there are no notable performance gains when compared to other three basic layouts. We recommend that: in practical applications, the layout AoaS is the best choice since the tiled version is the fastest one among the three versions of GPU implementations, especially on single precision.

cs.DC