SearcharxivSearch

arXiv subjects

Li Tao

Publications and source records attributed to Li Tao.

At least 19 recordsLinked to original sources

Scalable Training of Mixture-of-Experts Models with Megatron Core

Scaling Mixture-of-Experts (MoE) training introduces systems challenges absent in dense models. Because each token activates only a subset of experts, this sparsity allows total parameters to grow much faster than per-token computation, creating coupled constraints across memory, communication, and computation. Optimizing one dimension often shifts pressure to another, demanding co-design across the full system stack. We address these challenges for MoE training through integrated optimizations spanning memory (fine-grained recomputation, offloading, etc.), communication (optimized dispatchers, overlapping, etc.), and computation (Grouped GEMM, fusions, CUDA Graphs, etc.). The framework also provides Parallel Folding for flexible multi-dimensional parallelism, low-precision training support for FP8 and NVFP4, and efficient long-context training. On NVIDIA GB300 and GB200, it achieves 1,233/1,048 TFLOPS/GPU for DeepSeek-V3-685B and 974/919 TFLOPS/GPU for Qwen3-235B. As a performant, scalable, and production-ready open-source solution, it has been used across academia and industry for training MoE models ranging from billions to trillions of parameters on clusters scaling up to thousands of GPUs. This report explains how these techniques work, their trade-offs, and their interactions at the systems level, providing practical guidance for scaling MoE models with Megatron Core.

cs.DC

DS-HGCN: A Dual-Stream Hypergraph Convolutional Network for Predicting Student Engagement via Social Contagion

Student engagement is a critical factor influencing academic success and learning outcomes. Accurately predicting student engagement is essential for optimizing teaching strategies and providing personalized interventions. However, most approaches focus on single-dimensional feature analysis and assessing engagement based on individual student factors. In this work, we propose a dual-stream multi-feature fusion model based on hypergraph convolutional networks (DS-HGCN), incorporating social contagion of student engagement. DS-HGCN enables accurate prediction of student engagement states by modeling multi-dimensional features and their propagation mechanisms between students. The framework constructs a hypergraph structure to encode engagement contagion among students and captures the emotional and behavioral differences and commonalities by multi-frequency signals. Furthermore, we introduce a hypergraph attention mechanism to dynamically weigh the influence of each student, accounting for individual differences in the propagation process. Extensive experiments on public benchmark datasets demonstrate that our proposed method achieves superior performance and significantly outperforms existing state-of-the-art approaches.

cs.MM

HGC: A hybrid method combining gravity model and cycle structure for identifying influential spreaders in complex networks

Identifying influential spreaders in complex networks is a critical challenge in network science, with broad applications in disease control, information dissemination, and influence analysis in social networks. The gravity model, a distinctive approach for identifying influential spreaders, has attracted significant attention due to its ability to integrate node influence and the distance between nodes. However, the law of gravity is symmetric, whereas the influence between different nodes is asymmetric. Existing gravity model-based methods commonly rely on the topological distance as a metric to measure the distance between nodes. Such reliance neglects the strength or frequency of connections between nodes, resulting in symmetric influence values between node pairs, which ultimately leads to an inaccurate assessment of node influence. Moreover, these methods often overlook cycle structures within networks, which provide redundant pathways for nodes and contribute significantly to the overall connectivity and stability of the network. In this paper, we propose a hybrid method called HGC, which integrates the gravity model with effective distance and incorporates cycle structure to address the issues above. Effective distance, derived from probabilities, measures the distance between a source node and others by considering its connectivity, providing a more accurate reflection of actual relationships between nodes. To evaluate the accuracy and effectiveness of the proposed method, we conducted several experiments on eight real-world networks based on the Susceptible-Infected-Recovered model. The results demonstrate that HGC outperforms seven compared methods in accurately identifying influential nodes.

cs.CE

Angle measurement method of electronic speckle interferometry based on Michelson interferometer

{This paper proposes an angle measurement method based on Electronic Speckle Pattern Interferometry (ESPI) using a Michelson interferometer. By leveraging different principles within the same device, this method achieves complementary advantages across various angle ranges, enhancing measurement accuracy while maintaining high robustness. By utilizing CCD to record light field information in real time and combining geometric and ESPI methods, relationships between small angles and light field information are established, allowing for the design of relevant algorithms for real-time angle measurement. Numerical simulations and experiments were conducted to validate the feasibility and practicality of this method. Results indicate that it maintains measurement accuracy while offering a wide angle measurement range, effectively addressing the limitations of small angle measurements in larger ranges, showcasing significant potential for widespread applications in related fields.

physics.app-ph

End-to-end Graph Learning Approach for Cognitive Diagnosis of Student Tutorial

Cognitive diagnosis (CD) utilizes students' existing studying records to estimate their mastery of unknown knowledge concepts, which is vital for evaluating their learning abilities. Accurate CD is extremely challenging because CD is associated with complex relationships and mechanisms among students, knowledge concepts, studying records, etc. However, existing approaches loosely consider these relationships and mechanisms by a non-end-to-end learning framework, resulting in sub-optimal feature extractions and fusions for CD. Different from them, this paper innovatively proposes an End-to-end Graph Neural Networks-based Cognitive Diagnosis (EGNN-CD) model. EGNN-CD consists of three main parts: knowledge concept network (KCN), graph neural networks-based feature extraction (GNNFE), and cognitive ability prediction (CAP). First, KCN constructs CD-related interaction by comprehensively extracting physical information from students, exercises, and knowledge concepts. Second, a four-channel GNNFE is designed to extract high-order and individual features from the constructed KCN. Finally, CAP employs a multi-layer perceptron to fuse the extracted features to predict students' learning abilities in an end-to-end learning way. With such designs, the feature extractions and fusions are guaranteed to be comprehensive and optimal for CD. Extensive experiments on three real datasets demonstrate that our EGNN-CD achieves significantly higher accuracy than state-of-the-art models in CD.

cs.LG

Neuromorphic spatiotemporal optical flow: Enabling ultrafast visual perception beyond human capabilities

Optical flow, inspired by the mechanisms of biological visual systems, calculates spatial motion vectors within visual scenes that are necessary for enabling robotics to excel in complex and dynamic working environments. However, current optical flow algorithms, despite human-competitive task performance on benchmark datasets, remain constrained by unacceptable time delays (~0.6 seconds per inference, 4X human processing speed) in practical deployment. Here, we introduce a neuromorphic optical flow approach that addresses delay bottlenecks by encoding temporal information directly in a synaptic transistor array to assist spatial motion analysis. Compared to conventional spatial-only optical flow methods, our spatiotemporal neuromorphic optical flow offers the spatial-temporal consistency of motion information, rapidly identifying regions of interest in as little as 1-2 ms using the temporal motion cues derived from the embedded temporal information in the two-dimensional floating gate synaptic transistors. Thus, the visual input can be selectively filtered to achieve faster velocity calculations and various task execution. At the hardware level, due to the atomically sharp interfaces between distinct functional layers in two-dimensional van der Waals heterostructures, the synaptic transistor offers high-frequency response (~100 {\mu}s), robust non-volatility (>10000 s), and excellent endurance (>8000 cycles), enabling robust visual processing. In software benchmarks, our system outperforms state-of-the-art algorithms with a 400% speedup, frequently surpassing human-level performance while maintaining or enhancing accuracy by utilizing the temporal priors provided by the embedded temporal information.

cs.CV

GEM: A GEneral Memristive Transistor Model

Neuromorphic devices, with their distinct advantages in energy efficiency and parallel processing, are pivotal in advancing artificial intelligence applications. Among these devices, memristive transistors have attracted significant attention due to their superior stability and operation flexibility compared to two-terminal memristors. However, the lack of a robust model that accurately captures their complex electrical behavior has hindered further exploration of their potential. In this work, we introduce the GEneral Memristive transistor (GEM) model to address this challenge. The GEM model incorporates time-dependent differential equation, a voltage-controlled moving window function, and a nonlinear current output function, enabling precise representation of both switching and output characteristics in memristive transistors. Compared to previous models, the GEM model demonstrates a 300% improvement in modeling the switching behavior, while effectively capturing the inherent nonlinearities and physical limits of these devices. This advancement significantly enhances the realistic simulation of memristive transistors, thereby facilitating further exploration and application development.

physics.app-ph

Patient-Specific CT Doses Using DL-based Image Segmentation and GPU-based Monte Carlo Calculations for 10,281 Subjects

Computed tomography (CT) scans are a major source of medical radiation exposure worldwide. In countries like China, the frequency of CT scans has grown rapidly, particularly in routine physical examinations where chest CT scans are increasingly common. Accurate estimation of organ doses is crucial for assessing radiation risk and optimizing imaging protocols. However, traditional methods face challenges due to the labor-intensive process of manual organ segmentation and the computational demands of Monte Carlo (MC) dose calculations. In this study, we present a novel method that combines automatic image segmentation with GPU-accelerated MC simulations to compute patient-specific organ doses for a large cohort of 10,281 individuals undergoing CT examinations for physical examinations at a Chinese hospital. This is the first big-data study of its kind involving such a large population for CT dosimetry. The results show considerable inter-individual variability in CTDIvol-normalized organ doses, even among subjects with similar BMI or WED. Patient-specific organ doses vary widely, ranging from 33% to 164% normalized by the doses from ICRP Adult Reference Phantoms. Statistical analyses indicate that the "Reference Man" based average phantoms can lead to significant dosimetric uncertainties, with relative errors exceeding 50% in some cases. These findings underscore the fact that previous assessments of radiation risk may be inaccurate. It took our computational tool, on average, 135 seconds per subject, using a single NVIDIA RTX 3080 GPU card. The big-data analysis provides interesting data for improving CT dosimetry and risk assessment by avoiding uncertainties that were neglected in the past.

physics.med-ph

ResEnsemble-DDPM: Residual Denoising Diffusion Probabilistic Models for Ensemble Learning

Nowadays, denoising diffusion probabilistic models have been adapted for many image segmentation tasks. However, existing end-to-end models have already demonstrated remarkable capabilities. Rather than using denoising diffusion probabilistic models alone, integrating the abilities of both denoising diffusion probabilistic models and existing end-to-end models can better improve the performance of image segmentation. Based on this, we implicitly introduce residual term into the diffusion process and propose ResEnsemble-DDPM, which seamlessly integrates the diffusion model and the end-to-end model through ensemble learning. The output distributions of these two models are strictly symmetric with respect to the ground truth distribution, allowing us to integrate the two models by reducing the residual term. Experimental results demonstrate that our ResEnsemble-DDPM can further improve the capabilities of existing models. Furthermore, its ensemble learning strategy can be generalized to other downstream tasks in image generation and get strong competitiveness.

cs.CV

Aligning Language Models with Offline Learning from Human Feedback

Learning from human preferences is crucial for language models (LMs) to effectively cater to human needs and societal values. Previous research has made notable progress by leveraging human feedback to follow instructions. However, these approaches rely primarily on online learning techniques like Proximal Policy Optimization (PPO), which have been proven unstable and challenging to tune for language models. Moreover, PPO requires complex distributed system implementation, hindering the efficiency of large-scale distributed training. In this study, we propose an offline learning from human feedback framework to align LMs without interacting with environments. Specifically, we explore filtering alignment (FA), reward-weighted regression (RWR), and conditional alignment (CA) to align language models to human preferences. By employing a loss function similar to supervised fine-tuning, our methods ensure more stable model training than PPO with a simple machine learning system~(MLSys) and much fewer (around 9\%) computing resources. Experimental results demonstrate that conditional alignment outperforms other offline alignment methods and is comparable to PPO.

cs.CL

Transforming Graphs for Enhanced Attribute Clustering: An Innovative Graph Transformer-Based Method

Graph Representation Learning (GRL) is an influential methodology, enabling a more profound understanding of graph-structured data and aiding graph clustering, a critical task across various domains. The recent incursion of attention mechanisms, originally an artifact of Natural Language Processing (NLP), into the realm of graph learning has spearheaded a notable shift in research trends. Consequently, Graph Attention Networks (GATs) and Graph Attention Auto-Encoders have emerged as preferred tools for graph clustering tasks. Yet, these methods primarily employ a local attention mechanism, thereby curbing their capacity to apprehend the intricate global dependencies between nodes within graphs. Addressing these impediments, this study introduces an innovative method known as the Graph Transformer Auto-Encoder for Graph Clustering (GTAGC). By melding the Graph Auto-Encoder with the Graph Transformer, GTAGC is adept at capturing global dependencies between nodes. This integration amplifies the graph representation and surmounts the constraints posed by the local attention mechanism. The architecture of GTAGC encompasses graph embedding, integration of the Graph Transformer within the autoencoder structure, and a clustering component. It strategically alternates between graph embedding and clustering, thereby tailoring the Graph Transformer for clustering tasks, whilst preserving the graph's global structural information. Through extensive experimentation on diverse benchmark datasets, GTAGC has exhibited superior performance against existing state-of-the-art graph clustering methodologies.

cs.LG

Three-way causal attribute partial order structure analysis

As an emerging concept cognitive learning model, partial order formal structure analysis (POFSA) has been widely used in the field of knowledge processing. In this paper, we propose the method named three-way causal attribute partial order structure (3WCAPOS) to evolve the POFSA from set coverage to causal coverage in order to increase the interpretability and classification performance of the model. First, the concept of causal factor (CF) is proposed to evaluate the causal correlation between attributes and decision attributes in the formal decision context. Then, combining CF with attribute partial order structure, the concept of causal attribute partial order structure is defined and makes set coverage evolve into causal coverage. Finally, combined with the idea of three-way decision, 3WCAPOS is formed, which makes the purity of nodes in the structure clearer and the changes between levels more obviously. In addition, the experiments are carried out from the classification ability and the interpretability of the structure through the six datasets. Through these experiments, it is concluded the accuracy of 3WCAPOS is improved by 1% - 9% compared with classification and regression tree, and more interpretable and the processing of knowledge is more reasonable compared with attribute partial order structure.

cs.AI

Antibacterial Activity of Zinc Oxide Thin Films by Atomic Layer Deposition for Personal Protective Equipment Applications

The global pandemic has significantly increased the demand for personal protective equipment (PPE). The antimicrobial coating has been broadly applied to PPE to improve its prevention capability, especially after prolonged usage. However, antimicrobial coating by traditional methods, such as chemical vapor deposition, spraying, and slurry coating, suffers from drawbacks such as low efficiency, poor coverage, and loose adhesion to PPE. To overcome these limitations, this work adopted an atomic layer deposition (ALD) technique to deposit a zinc oxide (ZnO) thin film (~ 312.8 nm thick) as an antimicrobial coating and proved its advantage for depositing uniform ZnO coating on PPE with a fabric structure. Analysis by X-ray diffraction, Raman spectroscopy, and X-ray photoelectron spectroscopy confirmed the crystal structure and chemical composition of the ALD-ZnO. Ultraviolet-visible spectra disclosed a high absorption level of about 4.8 from 200 nm to 380 nm wavelength for ALD-ZnO, in contrast to 3.8 for commercial ZnO powders. Moreover, the ALD-ZnO exhibited strong antimicrobial properties when tested against Escherichia coli (E.Coli) in contrast to the control and bare glass samples. The colony-forming unit (CFU/ml) remained zero for all ALD-ZnO samples while varying between 2.30*109 and 4.97*109 with a median of 4.36*109 for the control and between 9.10*108 and 3.27*109 with a median of 2.04x 109 for the bare glass. Statistic analysis using null-hypothesis significance testing revealed that the calculated P value, between bare glass and control, ALD-ZnO and control, and bare glass and control, were all smaller than 0.0001 and significantly smaller than 0.05 alpha value, suggesting a high confidence level of ALD-ZnO as the main factor for preventing E.Coli growth.

physics.med-ph

Significant Ties Graph Neural Networks for Continuous-Time Temporal Networks Modeling

Temporal networks are suitable for modeling complex evolving systems. It has a wide range of applications, such as social network analysis, recommender systems, and epidemiology. Recently, modeling such dynamic systems has drawn great attention in many domains. However, most existing approaches resort to taking discrete snapshots of the temporal networks and modeling all events with equal importance. This paper proposes Significant Ties Graph Neural Networks (STGNN), a novel framework that captures and describes significant ties. To better model the diversity of interactions, STGNN introduces a novel aggregation mechanism to organize the most significant historical neighbors' information and adaptively obtain the significance of node pairs. Experimental results on four real networks demonstrate the effectiveness of the proposed framework.

cs.SI

Transmission-matrix Quantitative Phase Profilometry for Accurate and Fast Thickness Mapping of 2D Materials

The physical properties of two-dimensional (2D) materials may drastically vary with their thickness profiles. Current thickness profiling methods for 2D material (e.g., atomic force microscopy and ellipsometry) are limited in measurement throughput and accuracy. Here we present a novel high-speed and high-precision thickness profiling method, termed Transmission-Matrix Quantitative Phase Profilometry (TM-QPP). In TM-QPP, picometer-level optical pathlength sensitivity is enabled by extending the photon shot-noise limit of a high sensitivity common-path interferometric microscopy technique, while accurate thickness determination is realized by developing a transmission-matrix model that accounts for multiple refractions and reflections of light at sample interfaces. Using TM-QPP, the exact thickness profiles of monolayer and few-layered 2D materials (e.g., MoS2, MoSe2 and WSe2) are mapped over a wide field of view within seconds in a contact-free manner. Notably, TM-QPP is also capable of spatially resolving the number of layers of few-layered 2D materials.

physics.optics

A Spontaneously Formed Plasmonic-MoTe2 Hybrid Platform for Ultrasensitive Raman Enhancement

To develop highly sensitive, stable and repeatable surface-enhanced Raman scattering (SERS) substrates is crucial for analytical detection, which is a challenge for traditional metallic structures. Herein, by taking advantage of the high surface activity of 1T' transition metal telluride, we have fabricated high-density gold nanoparticles (AuNPs) that are spontaneously in-situ prepared on the 1T' MoTe2 atomic layers via a facile method, forming a plasmonic-2D material hybrid SERS substrate. This AuNP formation is unique to the 1T' phase, which is repressed in 2H MoTe2 with less surface activity. The hybrid structure generates coupling effects of electromagnetic and chemical enhancements, as well as excellent molecule adsorption, leading to the ultrasensitive (4*10^-17 M) and reproducible detection. Additionally, the immense fluorescence and photobleaching phenomena are mostly avoided. Flexible SERS tapes have been demonstrated in practical applications. Our approach facilitates the ultrasensitive SERS detection by a facile method, as well as the better mechanistic understanding of SERS beyond plasmonic effects.

physics.optics

Controlled synthesis of MoxW1-xTe2 atomic layers with emergent quantum states

Recently, new states of matter like superconducting or topological quantum states were found in transition metal dichalcogenides (TMDs) and manifested themselves in a series of exotic physical behaviors. Such phenomena have been demonstrated to exist in a series of transition metal tellurides including MoTe2, WTe2 and alloyed MoxW1-xTe2. However, the behaviors in the alloy system have been rarely addressed due to their difficulty in obtaining atomic layers with controlled composition, albeit the alloy offers a great platform to tune the quantum states. Here, we report a facile CVD method to synthesize the MoxW1-xTe2 with controllable thickness and chemical composition ratios. The atomic structure of monolayer MoxW1-xTe2 alloy was experimentally confirmed by scanning transmission electron microscopy (STEM). Importantly, two different transport behaviors including superconducting and Weyl semimetal (WSM) states were observed in Mo-rich Mo0.8W0.2Te2 and W-rich Mo0.2W0.8Te2 samples respectively. Our results show that the electrical properties of MoxW1-xTe2 can be tuned by controlling the chemical composition, demonstrating our controllable CVD growth method is an efficient strategy to manipulate the physical properties of TMDCs. Meanwhile, it provides a perspective on further comprehension and shed light on the design of device with topological multicomponent TMDCs materials.

cond-mat.mtrl-sci

Virtual Point Source Synthesis method for 3D Scintillation Detector Characterization

A novel data-processing method was developed to facilitate scintillation detector characterization. Combined with fan-beam calibration, this method can be used to quickly and conveniently calibrate gamma-ray detectors for SPECT, PET, homeland security or astronomy. Compared with traditional calibration methods, this new technique can accurately calibrate a photon-counting detector, including DOI information, with greatly reduced time. The enabling part of this technique is fan-beam scanning combined with a data-processing strategy called the common-data subset (CDS) method, which was used to synthesize the detector's mean detector response functions (MDRFs). Using this approach, $2N$ scans ($N$ in x and $N$ in y direction) are necessary to finish calibration of a 2D detector as opposed to $N^2$ scans with a pencil beam. For a 3D detector calibration, only $3N$ scans are necessary to achieve the 3D detector MDRFs that include DOI information. Moreover, this calibration technique can be used for detectors with complicated or irregular MDRFs. We present both Monte-Carlo simulations and experimental results that support the feasibility of this method.

physics.ins-det