SearcharxivSearch

arXiv subjects

Zhenhua An

Publications and source records attributed to Zhenhua An.

4 recordsLinked to original sources

Resonant Context Anchoring: Decoupling Attention Routing and Signal Gain at Inference Time

Large Language Models (LLMs) frequently exhibit "contextual disregard" when faced with input evidence that conflicts with their internal parametric memory, leading to persistent factual hallucinations. Existing mitigation strategies primarily rely on suppressing specific neuron activations or employing computationally expensive contrastive decoding mechanisms, which often result in increased perplexity or significantly elevated inference latency. To address these limitations, we propose Resonant Context Anchoring (RCA), a lightweight inference-time intervention method grounded in the perspective of residual stream signal dynamics. RCA aims to resolve the signal attenuation of external evidence during its propagation through deep networks. The core mechanism involves the orthogonal decoupling of routing logic and information magnitude within the self-attention module. By utilizing raw pre-softmax attention scores as an instantaneous metric of semantic alignment, we construct a dynamic gain field via non-linear rectification to selectively amplify the norms of value vectors corresponding to context tokens, without altering the attention probability distribution. This mechanism effectively elevates the signal-to-noise ratio (SNR) of input evidence within the residual stream mixture, thereby robustly anchoring the generation trajectory to the truthful context during inference. Extensive experiments on the Llama-3 model series demonstrate that RCA significantly improves contextual faithfulness across multiple factual consistency and strong knowledge-conflict tasks, effectively suppressing parametric hallucinations. Furthermore, results confirm that as a training-free and computationally negligible plug-and-play module, RCA achieves a Pareto improvement in faithfulness and fluency while maintaining the model's general language understanding capabilities.

cs.CL

Benchmarking neural surrogates on realistic spatiotemporal multiphysics flows

Predicting multiphysics dynamics is computationally expensive and challenging due to the severe coupling of multi-scale, heterogeneous physical processes. While neural surrogates promise a paradigm shift, the field currently suffers from an "illusion of mastery", as repeatedly emphasized in top-tier commentaries: existing evaluations overly rely on simplified, low-dimensional proxies, which fail to expose the models' inherent fragility in realistic regimes. To bridge this critical gap, we present REALM (REalistic AI Learning for Multiphysics), a rigorous benchmarking framework designed to test neural surrogates on challenging, application-driven reactive flows. REALM features 11 high-fidelity datasets spanning from canonical multiphysics problems to complex propulsion and fire safety scenarios, alongside a standardized end-to-end training and evaluation protocol that incorporates multiphysics-aware preprocessing and a robust rollout strategy. Using this framework, we systematically benchmark over a dozen representative surrogate model families, including spectral operators, convolutional models, Transformers, pointwise operators, and graph/mesh networks, and identify three robust trends: (i) a scaling barrier governed jointly by dimensionality, stiffness, and mesh irregularity, leading to rapidly growing rollout errors; (ii) performance primarily controlled by architectural inductive biases rather than parameter count; and (iii) a persistent gap between nominal accuracy metrics and physically trustworthy behavior, where models with high correlations still miss key transient structures and integral quantities. Taken together, REALM exposes the limits of current neural surrogates on realistic multiphysics flows and offers a rigorous testbed to drive the development of next-generation physics-aware architectures.

cs.LG

Graphics Processing Unit/Artificial Neural Network-accelerated large-eddy simulation of turbulent combustion: Application to swirling premixed flames

Within the scope of reacting flow simulations, the real-time direct integration (DI) of stiff ordinary differential equations (ODE) for the computation of chemical kinetics stands as the primary demand on computational resources. Meanwhile, as the number of transport equations that need to be solved increases, the computational cost grows more substantially, particularly for those combustion models involving direct coupling of chemistry and flow such as the transported probability density function model. In the current study, an integrated Graphics Processing Unit-Artificial Neural Network (GPU-ANN) framework is introduced to comply with heavy computational costs while maintaining high fidelity. Within this framework, a GPU-based solver is employed to solve partial differential equations and compute thermal and transport properties, and an ANN is utilized to replace the calculation of reaction rates. Large eddy simulations of two swirling flames provide a robust validation, affirming and extending the GPU-ANN approach's applicability to challenging scenarios. The simulation results demonstrate a strong correlation in the macro flame structure and statistical characteristics between the GPU-ANN approach and the traditional Central Processing Unit (CPU)-based solver with DI. This comparison indicates that the GPU-ANN approach is capable of attaining the same degree of precision as the conventional CPU-DI solver, even in more complex scenarios. In addition, the overall speed-up factor for the GPU-ANN approach is over two orders of magnitude. This study establishes the potential groundwork for widespread application of the proposed GPU-ANN approach in combustion simulations, addressing various and complex scenarios based on detailed chemistry, while significantly reducing computational costs.

physics.flu-dyn

Evaluation of flamelet-based models for liquid ammonia combustion in a temporally evolving mixing layer

Liquid ammonia combustion can be enhanced by co-firing with small molecular fuels such as methane, and liquid ammonia will undergo flash evaporation due to its relatively low saturation pressure. These characteristics, involving the presence of multiple fuel streams, a rapid phase change process, and strong heat loss, pose challenges for flamelet modeling of liquid ammonia combustion. To address these issues, this study aims to evaluate the effectiveness of flamelet-based models for liquid ammonia combustion in a turbulent mixing layer. Specifically, the extended flamelet/progress variable (E-FPV), extended flamelet-generated manifolds (E-FGM), and extended hybrid (E-Hybrid) models are developed and assessed. Firstly, a three-dimensional Point-Particle Direct Numerical Simulation (PP-DNS) with detailed chemistry is performed, where the turbulent flow is fully resolved, and the ammonia droplets are described by the Lagrangian method, to investigate the combustion characteristics of a liquid ammonia/methane co-fired flame and to provide state-of-the-art validation data for flamelet modeling. The PP-DNS results reveal distinct stages in the liquid ammonia/methane co-fired flame. The phase change process introduces significant heat loss due to the high latent heat of liquid ammonia. Subsequently, flamelet-based models are developed to account for the complex fuel streams, rapid phase change process, and strong local heat loss. The performance of these models is evaluated through a priori analysis by comparing the predictions with the PP-DNS results. The a priori results show that the E-FGM model outperforms the E-FPV and E-Hybrid models. This superior performance can be attributed to the rapid flash evaporation and sufficient mixing of the superheated ammonia, resulting in the dominance of the premixed combustion mode in liquid ammonia combustion.

physics.flu-dyn