SearcharxivSearch

arXiv subjects

Shaohua Chen

Publications and source records attributed to Shaohua Chen.

8 recordsLinked to original sources

Neural-Embedded Graphical Model for Self-Consistent Hierarchical Upscaling of Complex Composites

A persistent challenge in computational physical modeling is the substantial disparity between the characteristic length scales of microstructures and macroscopic structural components. Multiscale modeling has been widely adopted to bridge this gap by coupling methodologies tailored to different scales. However, conventional approaches, such as asymptotic homogenization (bottom-up) and submodeling (top-down), often entail rigorous mathematical prerequisites or intricate interfacing procedures. To address these limitations, we introduce a fully scalable neural-embedded graphical model (NEGM) that provides a unified framework for the progressive upscaling of highly heterogeneous composite materials. Specifically, NEGM encodes all microstructure- and material-related complexities into constituent neural network blocks, which are then organized into a hypergraph to simulate progressively larger domains. Extensive numerical benchmarks demonstrate that NEGM reliably predicts the physical responses of 2D and 3D composites exhibiting strong material nonlinearity, arbitrary boundary conditions, and irregular geometries. Crucially, because NEGM relies solely on neural network training and inference, it offers a scale-invariant formulation. This enables iterative application of NEGM to upscale from the microscale to arbitrarily large scales, circumventing the need for complex interfacing protocols between disparate modeling frameworks. We validate this progressive upscaling strategy on a large mosaic composite domain, showing that the accumulated error can be effectively contained provided the constituent blocks achieve sufficiently high prediction accuracy. Our findings suggest that artificial neural networks not only enhance the efficiency of direct single-scale simulations, as previously demonstrated, but also provide a clean and elegant pathway toward streamlined multiscale modeling.

cond-mat.dis-nn

A General-purpose Solver of Fourier Neural Swarm Operator Towards Accurate and Efficient Mechanical Modeling of Ultra Large Composite Materials

Composite media with complex microstructures exhibit highly tailorable mechanical properties but remain challenging to model efficiently and accurately. Conventional homogenization often oversimplifies microstructural effects, whereas multiscale approaches typically require costly coupling across spatial and temporal scales. To address these limitations, we propose a two-scale neural-swarm framework for large-scale mechanical modeling of heterogeneous composites. At the local scale, the mechanical characteristics of representative microstructural features are encoded into building-block Fourier neural operators (FNOs) using level-set representations. At the global scale, these pretrained FNOs are assembled into an FNO swarm according to the spatial distribution of microstructural constituents. A coarse-mesh finite element model is employed to provide global physical guidance, while Schwarz iteration is used to synchronize neighboring FNOs and enforce consistency across shared interfaces. The proposed framework is validated through nonlinear simulations of SiC-Al composites with diverse microstructural configurations. Compared with nonlinear finite element analysis, the FNO-swarm method achieves comparable accuracy while reducing computational cost by orders of magnitude. For an extreme dual-property SiC-Al composite containing more than a billion nodal points, the proposed approach predicts the mechanical response within approximately one hour, demonstrating exceptional scalability. Furthermore, the framework naturally accommodates arbitrary Dirichlet boundary conditions and complex domain geometries. The proposed neural-swarm strategy provides a robust and scalable paradigm for large-scale mechanics simulations, reconciling the longstanding trade-off between computational efficiency and physical fidelity in heterogeneous materials modeling.

cond-mat.dis-nn

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities

Recent advances in large language models (LLMs) have highlighted the potential of reinforcement learning with verifiable rewards (RLVR) to enhance reasoning capabilities through extended output sequences. However, traditional RL frameworks face inefficiencies when handling ultra-long outputs due to long-tail sequence distributions and entropy collapse during training. To address these challenges, we propose an Ultra-Long Output Reinforcement Learning (UloRL) approach for advancing large language models' reasoning abilities. Specifically, we divide ultra long output decoding into short segments, enabling efficient training by mitigating delays caused by long-tail samples. Additionally, we introduce dynamic masking of well-Mastered Positive Tokens (MPTs) to prevent entropy collapse. Experimental results demonstrate the effectiveness of our approach. On the Qwen3-30B-A3B model, RL with segment rollout achieved 2.06x increase in training speed, while RL training with 128k-token outputs improves the model's performance on AIME2025 from 70.9\% to 85.1\% and on BeyondAIME from 50.7\% to 61.9\%, even surpassing Qwen3-235B-A22B with remarkable gains. These findings underscore the potential of our methods to advance the reasoning capabilities of LLMs with ultra-long sequence generation. We will release our code and model for further use by the community.

cs.CL

Hunyuan-TurboS: Advancing Large Language Models through Mamba-Transformer Synergy and Adaptive Chain-of-Thought

As Large Language Models (LLMs) rapidly advance, we introduce Hunyuan-TurboS, a novel large hybrid Transformer-Mamba Mixture of Experts (MoE) model. It synergistically combines Mamba's long-sequence processing efficiency with Transformer's superior contextual understanding. Hunyuan-TurboS features an adaptive long-short chain-of-thought (CoT) mechanism, dynamically switching between rapid responses for simple queries and deep "thinking" modes for complex problems, optimizing computational resources. Architecturally, this 56B activated (560B total) parameter model employs 128 layers (Mamba2, Attention, FFN) with an innovative AMF/MF block pattern. Faster Mamba2 ensures linear complexity, Grouped-Query Attention minimizes KV cache, and FFNs use an MoE structure. Pre-trained on 16T high-quality tokens, it supports a 256K context length and is the first industry-deployed large-scale Mamba model. Our comprehensive post-training strategy enhances capabilities via Supervised Fine-Tuning (3M instructions), a novel Adaptive Long-short CoT Fusion method, Multi-round Deliberation Learning for iterative improvement, and a two-stage Large-scale Reinforcement Learning process targeting STEM and general instruction-following. Evaluations show strong performance: overall top 7 rank on LMSYS Chatbot Arena with a score of 1356, outperforming leading models like Gemini-2.0-Flash-001 (1352) and o4-mini-2025-04-16 (1345). TurboS also achieves an average of 77.9% across 23 automated benchmarks. Hunyuan-TurboS balances high performance and efficiency, offering substantial capabilities at lower inference costs than many reasoning models, establishing a new paradigm for efficient large-scale pre-trained models.

cs.CL

Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent

In this paper, we introduce Hunyuan-Large, which is currently the largest open-source Transformer-based mixture of experts model, with a total of 389 billion parameters and 52 billion activation parameters, capable of handling up to 256K tokens. We conduct a thorough evaluation of Hunyuan-Large's superior performance across various benchmarks including language understanding and generation, logical reasoning, mathematical problem-solving, coding, long-context, and aggregated tasks, where it outperforms LLama3.1-70B and exhibits comparable performance when compared to the significantly larger LLama3.1-405B model. Key practice of Hunyuan-Large include large-scale synthetic data that is orders larger than in previous literature, a mixed expert routing strategy, a key-value cache compression technique, and an expert-specific learning rate strategy. Additionally, we also investigate the scaling laws and learning rate schedule of mixture of experts models, providing valuable insights and guidances for future model development and optimization. The code and checkpoints of Hunyuan-Large are released to facilitate future innovations and applications. Codes: https://github.com/Tencent/Hunyuan-Large Models: https://huggingface.co/tencent/Tencent-Hunyuan-Large

cs.CL

Collective dynamics and self-organization of active particles on random networks

Collective cell migration in 3D extracellular matrix (ECM) is crucial to many physiological and pathological processes. Migrating cells can generate active pulling forces via actin filament contraction, which are transmitted to the ECM fibers and lead to a dynamically evolving force network in the system. Here, we elucidate the role of such force network in regulating collective cell behaviors using a minimal active-particle-on-network (APN) model, in which active particles can pull the fibers and hop between neighboring nodes of the network following local durotaxis. Our model reveals a dynamic transition as the particle number density approaches a critical value, from an "absorbing" state containing isolated stationary small particle clusters, to a "active" state containing a single large cluster undergone constant dynamic reorganization. This reorganization is dominated by a subset of highly dynamic "radical" particles in the cluster, whose number also exhibits a transition at the same critical density. The transition is underlaid by the percolation of "influence spheres" due to the particle pulling forces. Our results suggest a robust mechanism based on ECM-mediated mechanical coupling for collective cell behaviors in 3D ECM.

physics.bio-ph

Microstructure Representation and Reconstruction of Heterogeneous Materials via Deep Belief Network for Computational Material Design

Integrated Computational Materials Engineering (ICME) aims to accelerate optimal design of complex material systems by integrating material science and design automation. For tractable ICME, it is required that (1) a structural feature space be identified to allow reconstruction of new designs, and (2) the reconstruction process be property-preserving. The majority of existing structural presentation schemes rely on the designer's understanding of specific material systems to identify geometric and statistical features, which could be biased and insufficient for reconstructing physically meaningful microstructures of complex material systems. In this paper, we develop a feature learning mechanism based on convolutional deep belief network to automate a two-way conversion between microstructures and their lower-dimensional feature representations, and to achieves a 1000-fold dimension reduction from the microstructure space. The proposed model is applied to a wide spectrum of heterogeneous material systems with distinct microstructural features including Ti-6Al-4V alloy, Pb63-Sn37 alloy, Fontainebleau sandstone, and Spherical colloids, to produce material reconstructions that are close to the original samples with respect to 2-point correlation functions and mean critical fracture strength. This capability is not achieved by existing synthesis methods that rely on the Markovian assumption of material microstructures.

cond-mat.mtrl-sci

A maximal-information color to gray conversion method for document images: Toward an optimal grayscale representation for document image binarization

A novel method to convert color/multi-spectral images to gray-level images is introduced to increase the performance of document binarization methods. The method uses the distribution of the pixel data of the input document image in a color space to find a transformation, called the dual transform, which balances the amount of information on all color channels. Furthermore, in order to reduce the intensity variations on the gray output, a color reduction preprocessing step is applied. Then, a channel is selected as the gray value representation of the document image based on the homogeneity criterion on the text regions. In this way, the proposed method can provide a luminance-independent contrast enhancement. The performance of the method is evaluated against various images from two databases, the ICDAR'03 Robust Reading, the KAIST and the DIBCO'09 datasets, subjectively and objectively with promising results. The ground truth images for the images from the ICDAR'03 Robust Reading dataset have been created manually by the authors.

cs.CV