Searcharxiv⌕ Search

arXiv subjects

Xuguang Wang

Publications and source records attributed to Xuguang Wang.

9 recordsLinked to original sources

MAPCast: A Convection Allowing MPAS Emulator for Ensemble-based Background Error Covariance Estimation Toward Multi-Scale Data Assimilation

Machine learning (ML) emulators offer a cost-efficient alternative to numerical weather prediction models for generating convection-allowing background ensembles in ensemble-based data assimilation (DA). However, few studies have explored ML-based surrogate background ensembles for estimating background-error covariances (BECs). This study develops a convection-allowing emulator, MAPCast, trained on historical convection-allowing simulations from the Model for Prediction Across Scales (MPAS), and evaluates its ability to estimate BECs, paving the way toward multiscale DA. The evaluation uses 10 retrospective convective cases at 15- and 60-min forecast lead times corresponding to subhourly and hourly DA. MAPCast reproduces MPAS forecasts with good fidelity, including realistic storm coverage, temporal evolution, and similar spatial and spectral characteristics of state variables. Discrepancies are primarily confined to small spatial scales near sharp gradients and convective-scale features and variables. For BEC statistics, MAPCast captures ensemble spread magnitude and spatial distribution for most variables, although larger errors occur for storm-related fields that are vertical velocity and reflectivity. Correlation structures are reproduced most faithfully at mesoscale and above, followed by at convective scales, whereas cross-variable correlations are less accurately represented than univariate correlations, indicating that multivariate coupling remains the principal limitation. MAPCast shows weaker replication of full-scale versus decomposed large and small-scale correlations. BEC estimates derived from 15-min forecasts consistently outperform those from 60-min forecasts, suggesting that shorter lead times better preserve flow-dependent error structures.

physics.ao-ph↗

Evaluating data-driven background ensembles covariances from Graphcast: a case study for Hurricane Lee (2023)

Short-term background ensemble covariances (BEC) are crucial for ensemble-based data assimilation (DA). However, limited studies so far have examined the fidelity of the cost-effective data-driven model in producing the short-term BEC for hurricane data assimilation. In this study, we evaluate the background ensemble spread and correlations from GraphCast against those of GEFS for Hurricane Lee (2023) during both its intensification and non-intensification phases. Specifically, the BEC in the hurricane vortex, the hurricane environment, and the vortex-environment interactions are examined. Within the hurricane vortex, the background ensemble of GraphCast is less dispersive than GEFS. The two models agree well on the background ensemble correlations that are tied to the primary circulation but show a larger correlation difference associated with the secondary circulation, indicating the two models represent unbalanced and diabatic processes differently. In the hurricane environment, binned univariable correlations show linear relationships between the two models, with a weaker horizontal geopotential height correlation in GraphCast. GraphCast also shows a reduced spread and a flatter empirical orthogonal function spectrum of the 500 hPa geopotential height background ensemble, with more perturbation growth distributed to smaller-scale features such as shortwaves. For the vortex-environment interaction, the two models produce close background ensemble correlation patterns for Lee's track but differ more for intensity. Overall, GraphCast can produce broadly consistent short-term BEC for Hurricane Lee compared to those of GEFS. However, systematic difference exists in certain variables, scales, and processes, suggesting the need to further investigate its fidelity in a cycled hurricane DA context.

physics.ao-ph↗

Empowering Bridge Digital Twins by Bridging the Data Gap with a Unified Synthesis Framework

As critical transportation infrastructure, bridges face escalating challenges from aging and deterioration, while traditional manual inspection methods suffer from low efficiency. Although 3D point cloud technology provides a new data-driven paradigm, its application potential is often constrained by the incompleteness of real-world data, which results from missing labels and scanning occlusions. To overcome the bottleneck of insufficient generalization in existing synthetic data methods, this paper proposes a systematic framework for generating 3D bridge data. This framework can automatically generate complete point clouds featuring component-level instance annotations, high-fidelity color, and precise normal vectors. It can be further extended to simulate the creation of diverse and physically realistic incomplete point clouds, designed to support the training of segmentation and completion networks, respectively. Experiments demonstrate that a PointNet++ model trained with our synthetic data achieves a mean Intersection over Union (mIoU) of 84.2% in real-world bridge semantic segmentation. Concurrently, a fine-tuned KT-Net exhibits superior performance on the component completion task. This research offers an innovative methodology and a foundational dataset for the 3D visual analysis of bridge structures, holding significant implications for advancing the automated management and maintenance of infrastructure.

cs.CV↗

On methods for assessment of the influence and impact of observations in convection-permitting numerical weather prediction

In numerical weather prediction (NWP), a large number of observations are used to create initial conditions for weather forecasting through a process known as data assimilation. An assessment of the value of these observations for NWP can guide us in the design of future observation networks, help us to identify problems with the assimilation system, and allow us to assess changes to the assimilation system. However, the assessment can be challenging in convection-permitting NWP. First, the strong nonlinearity in the forecast model limits the methods available for the assessment. Second, convection-permitting NWP typically uses a limited area model and provides short forecasts, giving problems with verification and our ability to gather sufficient statistics. Third, convection-permitting NWP often makes use of novel observations, which can be difficult to simulate in an observing system simulation experiment (OSSE). We compare methods that can be used to assess the value of observations in convection-permitting NWP and discuss operational considerations when using these methods. We focus on their applicability to ensemble forecasting systems, as these systems are becoming increasingly dominant for convection-permitting NWP. We also identify several future research directions: comparison of forecast validation using analyses and observations, the effect of ensemble size on assessing the value of observations, flow-dependent covariance localization, and generation and validation of the nature run in an OSSE.

physics.ao-ph↗

FastAdaBelief: Improving Convergence Rate for Belief-based Adaptive Optimizers by Exploiting Strong Convexity

AdaBelief, one of the current best optimizers, demonstrates superior generalization ability compared to the popular Adam algorithm by viewing the exponential moving average of observed gradients. AdaBelief is theoretically appealing in that it has a data-dependent $O(\sqrt{T})$ regret bound when objective functions are convex, where $T$ is a time horizon. It remains however an open problem whether the convergence rate can be further improved without sacrificing its generalization ability. %on how to exploit strong convexity to further improve the convergence rate of AdaBelief. To this end, we make a first attempt in this work and design a novel optimization algorithm called FastAdaBelief that aims to exploit its strong convexity in order to achieve an even faster convergence rate. In particular, by adjusting the step size that better considers strong convexity and prevents fluctuation, our proposed FastAdaBelief demonstrates excellent generalization ability as well as superior convergence. As an important theoretical contribution, we prove that FastAdaBelief attains a data-dependant $O(\log T)$ regret bound, which is substantially lower than AdaBelief. On the empirical side, we validate our theoretical analysis with extensive experiments in both scenarios of strong and non-strong convexity on three popular baseline models. Experimental results are very encouraging: FastAdaBelief converges the quickest in comparison to all mainstream algorithms while maintaining an excellent generalization ability, in cases of both strong or non-strong convexity. FastAdaBelief is thus posited as a new benchmark model for the research community.

cs.LG↗

Observation of quantum spin Hall states in Ta$_2$Pd$_3$Te$_5$

Two-dimensional topological insulators (2DTIs), which host the quantum spin Hall (QSH) effect, are one of the key materials in next-generation spintronic devices. To date, experimental evidence of the QSH effect has only been observed in a few materials, and thus, the search for new 2DTIs is at the forefront of physical and materials science. Here, we report experimental evidence of a 2DTI in the van der Waals material Ta$_2$Pd$_3$Te$_5$. First-principles calculations show that each monolayer of Ta$_2$Pd$_3$Te$_5$ is a 2DTI with weak interlayer interactions. Combined transport, angle-resolved photoemission spectroscopy, and scanning tunneling microscopy measurements confirm the existence of a band gap at the Fermi level and topological edge states inside the gap. These results demonstrate that Ta$_2$Pd$_3$Te$_5$ is a promising material for fabricating spintronic devices based on the QSH effect.

cond-mat.mtrl-sci↗

No Answer is Better Than Wrong Answer: A Reflection Model for Document Level Machine Reading Comprehension

The Natural Questions (NQ) benchmark set brings new challenges to Machine Reading Comprehension: the answers are not only at different levels of granularity (long and short), but also of richer types (including no-answer, yes/no, single-span and multi-span). In this paper, we target at this challenge and handle all answer types systematically. In particular, we propose a novel approach called Reflection Net which leverages a two-step training procedure to identify the no-answer and wrong-answer cases. Extensive experiments are conducted to verify the effectiveness of our approach. At the time of paper writing (May.~20,~2020), our approach achieved the top 1 on both long and short answer leaderboard, with F1 scores of 77.2 and 64.1, respectively.

cs.CL↗

Dataset and Neural Recurrent Sequence Labeling Model for Open-Domain Factoid Question Answering

While question answering (QA) with neural network, i.e. neural QA, has achieved promising results in recent years, lacking of large scale real-word QA dataset is still a challenge for developing and evaluating neural QA system. To alleviate this problem, we propose a large scale human annotated real-world QA dataset WebQA with more than 42k questions and 556k evidences. As existing neural QA methods resolve QA either as sequence generation or classification/ranking problem, they face challenges of expensive softmax computation, unseen answers handling or separate candidate answer generation component. In this work, we cast neural QA as a sequence labeling problem and propose an end-to-end sequence labeling model, which overcomes all the above challenges. Experimental results on WebQA show that our model outperforms the baselines significantly with an F1 score of 74.69% with word-based input, and the performance drops only 3.72 F1 points with more challenging character-based input.

cs.CL↗

Deep Recurrent Models with Fast-Forward Connections for Neural Machine Translation

Neural machine translation (NMT) aims at solving machine translation (MT) problems using neural networks and has exhibited promising results in recent years. However, most of the existing NMT models are shallow and there is still a performance gap between a single NMT model and the best conventional MT system. In this work, we introduce a new type of linear connections, named fast-forward connections, based on deep Long Short-Term Memory (LSTM) networks, and an interleaved bi-directional architecture for stacking the LSTM layers. Fast-forward connections play an essential role in propagating the gradients and building a deep topology of depth 16. On the WMT'14 English-to-French task, we achieve BLEU=37.7 with a single attention model, which outperforms the corresponding single shallow model by 6.2 BLEU points. This is the first time that a single NMT model achieves state-of-the-art performance and outperforms the best conventional model by 0.7 BLEU points. We can still achieve BLEU=36.3 even without using an attention mechanism. After special handling of unknown words and model ensembling, we obtain the best score reported to date on this task with BLEU=40.4. Our models are also validated on the more difficult WMT'14 English-to-German task.

cs.CL↗