Searcharxiv⌕ Search

arXiv subjects

Rui Su

Publications and source records attributed to Rui Su.

At least 37 records · Page 2Linked to original sources

OpenSWI: A Massive-Scale Benchmark Dataset for Surface Wave Dispersion Curve Inversion

Surface wave dispersion curve inversion plays a critical role in both shallow resource exploration and deep geological studies, yet it remains hindered by sensitivity to initial models and low computational efficiency. Recently, data-driven deep learning methods, inspired by advances in computer vision, have shown promising potential to address these challenges. However, the lack of large-scale, diverse benchmark datasets remains a major obstacle to their development and evaluation. To bridge this gap, we present OpenSWI, a comprehensive benchmark dataset generated through the Surface Wave Inversion Dataset Preparation (SWIDP) pipeline. OpenSWI includes two synthetic datasets tailored to different research scales and scenarios, OpenSWI-shallow and OpenSWI-deep, and an AI-ready real-world dataset for generalization evaluation, OpenSWI-real. OpenSWI-shallow, derived from the 2-D OpenFWI geological model dataset, contains over 22 million 1-D velocity profiles paired with fundamental-mode phase and group velocity dispersion curves, spanning a wide range of shallow geological structures (e.g., flat layers, faults, folds, realistic stratigraphy). OpenSWI-deep, built from 14 global and regional 3-D geological models, comprises 1.26 million high-fidelity 1-D velocity-dispersion pairs for deep-Earth studies. OpenSWI-real, compiled from open-source projects, contains two sets of observed dispersion curves with corresponding reference models, serving as a benchmark for evaluating model generalization. To demonstrate utility, we trained models on OpenSWI-shallow and -deep and evaluated them on OpenSWI-real, demonstrating strong agreement between predictions and references, which confirms the diversity and representativeness of the dataset. To advance intelligent surface wave inversion, we release the SWIDP toolbox, OpenSWI datasets, and trained models for the research community.

physics.geo-ph↗

SeisMoLLM: Advancing Seismic Monitoring via Cross-modal Transfer with Pre-trained Large Language Model

Recent advances in deep learning have revolutionized seismic monitoring, yet developing a foundation model that performs well across multiple complex tasks remains challenging, particularly when dealing with degraded signals or data scarcity. This work presents SeisMoLLM, the first foundation model that utilizes cross-modal transfer for seismic monitoring, to unleash the power of large-scale pre-training from a large language model without requiring direct pre-training on seismic datasets. Through elaborate waveform tokenization and fine-tuning of pre-trained GPT-2 model, SeisMoLLM achieves state-of-the-art performance on the DiTing and STEAD datasets across five critical tasks: back-azimuth estimation, epicentral distance estimation, magnitude estimation, phase picking, and first-motion polarity classification. It attains 36 best results out of 43 task metrics and 12 top scores out of 16 few-shot generalization metrics, with many relative improvements ranging from 10% to 50%. In addition to its superior performance, SeisMoLLM maintains efficiency comparable to or even better than lightweight models in both training and inference. These findings establish SeisMoLLM as a promising foundation model for practical seismic monitoring and highlight cross-modal transfer as an exciting new direction for earthquake studies, showcasing the potential of advanced deep learning techniques to propel seismology research forward.

cs.LG↗

Progressive Cross-Stream Cooperation in Spatial and Temporal Domain for Action Localization

Spatio-temporal action localization consists of three levels of tasks: spatial localization, action classification, and temporal localization. In this work, we propose a new progressive cross-stream cooperation (PCSC) framework that improves all three tasks above. The basic idea is to utilize both spatial region (resp., temporal segment proposals) and features from one stream (i.e., the Flow/RGB stream) to help another stream (i.e., the RGB/Flow stream) to iteratively generate better bounding boxes in the spatial domain (resp., temporal segments in the temporal domain). In this way, not only the actions could be more accurately localized both spatially and temporally, but also the action classes could be predicted more precisely. Specifically, we first combine the latest region proposals (for spatial detection) or segment proposals (for temporal localization) from both streams to form a larger set of labelled training samples to help learn better action detection or segment detection models. Second, to learn better representations, we also propose a new message passing approach to pass information from one stream to another stream, which also leads to better action detection and segment detection models. By first using our newly proposed PCSC framework for spatial localization at the frame-level and then applying our temporal PCSC framework for temporal localization at the tube-level, the action localization results are progressively improved at both the frame level and the video level. Comprehensive experiments on two benchmark datasets UCF-101-24 and J-HMDB demonstrate the effectiveness of our newly proposed approaches for spatio-temporal action localization in realistic scenarios.

cs.CV↗

Improving Weakly Supervised Temporal Action Localization by Exploiting Multi-resolution Information in Temporal Domain

Weakly supervised temporal action localization is a challenging task as only the video-level annotation is available during the training process. To address this problem, we propose a two-stage approach to fully exploit multi-resolution information in the temporal domain and generate high quality frame-level pseudo labels based on both appearance and motion streams. Specifically, in the first stage, we generate reliable initial frame-level pseudo labels, and in the second stage, we iteratively refine the pseudo labels and use a set of selected frames with highly confident pseudo labels to train neural networks and better predict action class scores at each frame. We fully exploit temporal information at multiple scales to improve temporal action localization performance. Specifically, in order to obtain reliable initial frame-level pseudo labels, in the first stage, we propose an Initial Label Generation (ILG) module, which leverages temporal multi-resolution consistency to generate high quality class activation sequences (CASs), which consist of a number of sequences with each sequence measuring how likely each video frame belongs to one specific action class. In the second stage, we propose a Progressive Temporal Label Refinement (PTLR) framework. In our PTLR framework, two networks called Network-OTS and Network-RTS, which are respectively used to generate CASs for the original temporal scale and the reduced temporal scales, are used as two streams (i.e., the OTS stream and the RTS stream) to refine the pseudo labels in turn. By this way, the multi-resolution information in the temporal domain is exchanged at the pseudo label level, and our work can help improve each stream (i.e., the OTS/RTS stream) by exploiting the refined pseudo labels from another stream (i.e., the RTS/OTS stream).

cs.CV↗

Deep Reparameterization for Full Waveform Inversion: Architecture Benchmarking, Robust Inversion, and Multiphysics Extension

Full waveform inversion (FWI) is a high-resolution subsurface imaging technique, but its effectiveness is limited by challenges such as noise contamination, sparse acquisition, and artifacts from multiparameter coupling. To address these limitations, this study develops a deep reparameterized FWI (DR-FWI) framework, in which subsurface parameters are represented by a deep neural network. Instead of directly optimizing the parameters, DR-FWI optimizes the network weights to reconstruct them, thereby embedding structural priors and facilitating optimization. To provide benchmark guidelines for the design of DR-FWI, we conduct a comparative analysis of three representative architectures (U-Net, CNN, MLP) combined with two initial model embedding strategies: one pretraining the network to generate predefined initial models (pretraining-based), while the other directly adds network outputs to the initial models. Extensive ablation experiments show that combining CNN with pretraining-based initialization significantly enhances inversion accuracy, offering valuable insights into network design. To further understand the mechanism of DR-FWI, spectral bias analysis reveals that the network first captures low-frequency features and gradually reconstructs high-frequency details, enabling an adaptive multi-scale inversion strategy. Notably, the robustness of DR-FWI is validated under various noise levels and sparse acquisition scenarios, where its strong performance with limited shots and receivers demonstrates reduced reliance on dense observational data. Additionally, a backbone-branch structure is proposed to extend DR-FWI to multiparameter inversion, and its efficacy in mitigating cross-parameter interference is validated on a synthetic anomaly model and the Marmousi2 model. These results suggest a promising direction for joint inversion involving multiple parameters or multiphysics.

physics.geo-ph↗

Swarm Intelligence Enhanced Reasoning: A Density-Driven Framework for LLM-Based Multi-Agent Optimization

Recently, many approaches, such as Chain-of-Thought (CoT) prompting and Multi-Agent Debate (MAD), have been proposed to further enrich Large Language Models' (LLMs) complex problem-solving capacities in reasoning scenarios. However, these methods may fail to solve complex problems due to the lack of ability to find optimal solutions. Swarm Intelligence has been serving as a powerful tool for finding optima in the field of traditional optimization problems. To this end, we propose integrating swarm intelligence into the reasoning process by introducing a novel Agent-based Swarm Intelligence (ASI) paradigm. In this paradigm, we formulate LLM reasoning as an optimization problem and use a swarm intelligence scheme to guide a group of LLM-based agents in collaboratively searching for optimal solutions. To avoid swarm intelligence getting trapped in local optima, we further develop a Swarm Intelligence Enhancing Reasoning (SIER) framework, which develops a density-driven strategy to enhance the reasoning ability. To be specific, we propose to perform kernel density estimation and non-dominated sorting to optimize both solution quality and diversity simultaneously. In this case, SIER efficiently enhances solution space exploration through expanding the diversity of the reasoning path. Besides, a step-level quality evaluation is used to help agents improve solution quality by correcting low-quality intermediate steps. Then, we use quality thresholds to dynamically control the termination of exploration and the selection of candidate steps, enabling a more flexible and efficient reasoning process. Extensive experiments are ...

cs.MA↗

Room temperature spin-layer locking of exciton-polariton nonlinearities

Recent advancements in transition metal dichalcogenides (TMDs) have unveiled exceptional optical and electronic characteristics, opened up new opportunities, and provided a unique platform for exploring light-matter interactions under the strong coupling regime. The exploitation of exciton-polaritons, with their peculiar hybrid light-matter properties, for the development of spintronic customizable devices that enhance both the information capacity and functionality at ambient temperatures is often suggested as a promising route. However, although TMD polaritons have shown promising potential, the microscopic mechanisms leading to nonlinearities in TMD polaritons are complex and their spin-anisotropy, a crucial requirement for many proposed polaritonic devices, has been missing. Here, we demonstrate the absence of spin-anisotropic interaction in a monolayer WS2 microcavity (at room temperature) and show how spin-dependent interactions can be controlled and spin anisotropy recovered by engineering double WS2 layer structures with varied interlayer spacing. We attribute this phenomenon to a distinctive feature in exciton-polariton physics: layer-dependent polariton-phonon coupling. We use theoretical calculations of the phonon electrostatic potentials finding a drastically different coupling strength for single and double monolayer samples and discuss qualitatively how this explains the observed spin-anisotropic response. This is further consistent with experiments on multi WS2 layer samples and the identification of a critical separation distance, above which an effective single monolayer spin-anisotropic response is recovered, both in experiment and theory. Our work lays the groundwork for the development of spin-optronic polaritonic devices at room temperature.

physics.optics↗

Fast Information Streaming Handler (FisH): A Unified Seismic Neural Network for Single Station Real-Time Earthquake Early Warning

Existing EEW approaches often treat phase picking, location estimation, and magnitude estimation as separate tasks, lacking a unified framework. Additionally, most deep learning models in seismology rely on full three-component waveforms and are not suitable for real-time streaming data. To address these limitations, we propose a novel unified seismic neural network called Fast Information Streaming Handler (FisH). FisH is designed to process real-time streaming seismic data and generate simultaneous results for phase picking, location estimation, and magnitude estimation in an end-to-end fashion. By integrating these tasks within a single model, FisH simplifies the overall process and leverages the nonlinear relationships between tasks for improved performance. The FisH model utilizes RetNet as its backbone, enabling parallel processing during training and recurrent handling during inference. This capability makes FisH suitable for real-time applications, reducing latency in EEW systems. Extensive experiments conducted on the STEAD benchmark dataset provide strong validation for the effectiveness of our proposed FisH model. The results demonstrate that FisH achieves impressive performance across multiple seismic event detection and characterization tasks. Specifically, it achieves an F1 score of 0.99/0.96. Also, FisH demonstrates precise earthquake location estimation, with location error of only 6.0km, a distance error of 2.6km, and a back-azimuth error of 19°. The model also exhibits accurate earthquake magnitude estimation, with a magnitude error of just 0.14. Additionally, FisH is capable of generating real-time estimations, providing location and magnitude estimations with a location error of 8.06km and a magnitude error of 0.18 within a mere 3 seconds after the P-wave arrives.

cs.CV↗

Trembling Motion of Exciton-Polaritons Close to the Rashba-Dresselhaus Regime

We report the experimental emulation of trembling quantum motion, or Zitterbewegung, of exciton polaritons in a perovskite microcavity at room temperature. By introducing liquid crystal molecules into the microcavity, we achieve spinor states with synthetic Rashba-Dresselhaus spin-orbit coupling and tunable energy splitting. Under a resonant excitation, the polariton fluid exhibits clear trembling motion perpendicular to its flowing direction, accompanied by a unique spin pattern resembling interlocked fingers. Furthermore, leveraging on the sizable tunability of energy gaps by external electrical voltages, we observe the continuous transition of polariton Zitterbewegung from relativistic (small gaps) to non-relativistic (large gaps) regimes. Our findings pave the way for using exciton polaritons in the emulation of relativistic quantum physics.

cond-mat.mes-hall↗

Observation of perovskite topological valley exciton-polaritons at room temperature

Topological exciton-polaritons are a burgeoning class of topological photonic systems distinguished by their hybrid nature as part-light, part-matter quasiparticles. Their further control over novel valley degree of freedom (DOF) has offered considerable potential for developing active topological optical devices towards information processing. However, the experimental demonstration of propagating topological exciton-polaritons with valley DOF remains elusive at room temperature. Here, employing a two-dimensional (2D) valley-Hall perovskite lattice, we report the experimental observation of valley-polarized topological exciton-polaritons and their valley-dependent propagations at room temperature. The 2D valley-Hall perovskite lattice consists of two mutually inverted honeycomb lattices with broken inversion symmetry. By measuring their band structure with angle-resolved photoluminescence spectra, we experimentally verify the existence of valley-polarized polaritonic topological kink states with a large gap opening of ~ 9 meV in the bearded interface at room temperature. Moreover, these valley-polarized states exhibit counter-propagating behaviors under a resonant excitation at room temperature. Our results not only expand the landscape of realizing topological exciton-polaritons, but also pave the way for the development of topological valleytronic devices employing exciton-polaritons with valley DOF at room temperature

physics.optics↗

Perovskite topological exciton-polariton disclination laser at room temperature

Topologically nontrivial systems can be protected by band topology in momentum space, as seen in topological insulators and semimetals, or real-space topology, such as in lattice deformations known as topological disclinations (TDs). TDs, with inherent chiral symmetry, can support localized states pinned spectrally to the middle of the topological gap, preventing hybridization with bulk bands, and making them promising for topological lasers. Here, we experimentally realize a C4v symmetric TD laser based on perovskite exciton-polariton lattices at room temperature. Protected by the chiral and point group symmetries of the lattice, the TD state emerges in the middle of the gap and at the core of the perovskite lattice. Under a non-resonant pulsed excitation, coherent polariton lasing occurs precisely at the TD state with a low threshold of 9.5 uJ/cm2, as confirmed by momentum space and real space spectra measurements. This study not only introduces a class of symmetry-protected topological lasers, but also expands the landscape for exploring exciton-polariton light-matter interactions with novel topological structures.

physics.optics↗

Multifunctional magnetic oxide-MoS$_2$ heterostructures on silicon

Correlated oxides and related heterostructures are intriguing for developing future multifunctional devices by exploiting their exotic properties, but their integration with other materials, especially on Si-based platforms, is challenging. Here, van der Waals heterostructures of La$_{0.7}$Sr$_{0.3}$MnO$_3$ (LSMO), a correlated manganite perovskite, and MoS$_2$ are demonstrated on Si substrates with multiple functions. To overcome the problems due to the incompatible growth process, technologies involving freestanding LSMO membranes and van der Waals force-mediated transfer are used to fabricate the LSMO-MoS$_2$ heterostructures. The LSMO-MoS$_2$ heterostructures exhibit a gate-tunable rectifying behavior, based on which metal-semiconductor field-effect transistors (MESFETs) with on-off ratios of over 104 can be achieved. The LSMO-MoS$_2$ heterostructures can function as photodiodes displaying considerable open-circuit voltages and photocurrents. In addition, the colossal magnetoresistance of LSMO endows the LSMO-MoS$_2$ heterostructures with an electrically tunable magnetoresponse at room temperature. This work not only proves the applicability of the LSMO-MoS$_2$ heterostructure devices on Si-based platform but also demonstrates a paradigm to create multifunctional heterostructures from materials with disparate properties.

cond-mat.str-el↗

FengWu: Pushing the Skillful Global Medium-range Weather Forecast beyond 10 Days Lead

We present FengWu, an advanced data-driven global medium-range weather forecast system based on Artificial Intelligence (AI). Different from existing data-driven weather forecast methods, FengWu solves the medium-range forecast problem from a multi-modal and multi-task perspective. Specifically, a deep learning architecture equipped with model-specific encoder-decoders and cross-modal fusion Transformer is elaborately designed, which is learned under the supervision of an uncertainty loss to balance the optimization of different predictors in a region-adaptive manner. Besides this, a replay buffer mechanism is introduced to improve medium-range forecast performance. With 39-year data training based on the ERA5 reanalysis, FengWu is able to accurately reproduce the atmospheric dynamics and predict the future land and atmosphere states at 37 vertical levels on a 0.25° latitude-longitude resolution. Hindcasts of 6-hourly weather in 2018 based on ERA5 demonstrate that FengWu performs better than GraphCast in predicting 80\% of the 880 reported predictands, e.g., reducing the root mean square error (RMSE) of 10-day lead global z500 prediction from 733 to 651 $m^{2}/s^2$. In addition, the inference cost of each iteration is merely 600ms on NVIDIA Tesla A100 hardware. The results suggest that FengWu can significantly improve the forecast skill and extend the skillful global medium-range weather forecast out to 10.75 days lead (with ACC of z500 > 0.6) for the first time.

cs.AI↗

Efficient and accurate simulation of vitrification in multi-component metallic liquids with neural-network potentials

Constructing accurate interatomic potential and overcoming the exponential growth of structural equilibration time are challenges to the atomistic investigations of the composition-dependent structure and dynamics during the vitrification process of deeply supercooled multi-component metallic liquids. In this work, we describe a state-of-the-art strategy to address these challenges simultaneously. In the case of the representative Zr-Cu-Al system, in combination with a general algorithm for generating the neural-network potentials (NNP) of multi-component metallic glasses effectively and accurately, we propose a highly efficient atom-swapping hybrid Monte Carlo (SHMC) algorithm for accelerating the thermodynamic equilibration of deeply supercooled liquids. Extensive calculations demonstrate that the newly developed NNP faithfully reproduces the phase stabilities and structural characteristics obtained from the ab initio calculations and experiments. In the combined NNP-SHMC algorithm, the structure equilibration time in the deeply supercooled temperatures is accelerated by at least five orders of magnitudes, and the quenched glassy samples exhibit comparable stability to those prepared in the laboratory. Our results pave the way for the next-generation studies of the vitrification process and, thereby the composition-dependent glass-forming ability and physical properties of multi-component metallic glasses.

cond-mat.mtrl-sci↗

Slow Motion Matters: A Slow Motion Enhanced Network for Weakly Supervised Temporal Action Localization

Weakly supervised temporal action localization (WTAL) aims to localize actions in untrimmed videos with only weak supervision information (e.g. video-level labels). Most existing models handle all input videos with a fixed temporal scale. However, such models are not sensitive to actions whose pace of the movements is different from the ``normal" speed, especially slow-motion action instances, which complete the movements with a much slower speed than their counterparts with a normal speed. Here arises the slow-motion blurred issue: It is hard to explore salient slow-motion information from videos at ``normal" speed. In this paper, we propose a novel framework termed Slow Motion Enhanced Network (SMEN) to improve the ability of a WTAL network by compensating its sensitivity on slow-motion action segments. The proposed SMEN comprises a Mining module and a Localization module. The mining module generates mask to mine slow-motion-related features by utilizing the relationships between the normal motion and slow motion; while the localization module leverages the mined slow-motion features as complementary information to improve the temporal action localization results. Our proposed framework can be easily adapted by existing WTAL networks and enable them be more sensitive to slow-motion actions. Extensive experiments on three benchmarks are conducted, which demonstrate the high performance of our proposed framework.

cs.CV↗

3D-QueryIS: A Query-based Framework for 3D Instance Segmentation

Previous top-performing methods for 3D instance segmentation often maintain inter-task dependencies and the tendency towards a lack of robustness. Besides, inevitable variations of different datasets make these methods become particularly sensitive to hyper-parameter values and manifest poor generalization capability. In this paper, we address the aforementioned challenges by proposing a novel query-based method, termed as 3D-QueryIS, which is detector-free, semantic segmentation-free, and cluster-free. Specifically, we propose to generate representative points in an implicit manner, and use them together with the initial queries to generate the informative instance queries. Then, the class and binary instance mask predictions can be produced by simply applying MLP layers on top of the instance queries and the extracted point cloud embeddings. Thus, our 3D-QueryIS is free from the accumulated errors caused by the inter-task dependencies. Extensive experiments on multiple benchmark datasets demonstrate the effectiveness and efficiency of our proposed 3D-QueryIS method.

cs.CV↗

Atomic Origin of Annealing Embrittlement in Metallic Glasses

An atomistic understanding of annealing embrittlement is a longstanding issue for metallic glasses, which is still lacking due to the insurmountable gap between the thermal history of atomic models and laboratory-made samples. Here, based on a thermal-cycling annealing method that can vary the effective quenching rate over ten orders of magnitude, we perform an atomistic study of the ductile-brittle transition in a ternary model metallic glass, which can be keyed to the annealing embrittlement in bulk metallic glasses. We reveal that thermal annealing can effectively obliterate thermally active-able "defects", which are abundant in the hyper-quenched and ductile glass but gives rise to strain-created shear events in the well-annealed and brittle glass. While the activation of the strain-created events eventually causes single shear banding, other local structural disruptions can be "healed" by the same type of events upon stress reversal, thereby hindering shear band broadening or multiplication, and resulting in annealing embrittlement.

cond-mat.mtrl-sci↗

NSNet: Non-saliency Suppression Sampler for Efficient Video Recognition

It is challenging for artificial intelligence systems to achieve accurate video recognition under the scenario of low computation costs. Adaptive inference based efficient video recognition methods typically preview videos and focus on salient parts to reduce computation costs. Most existing works focus on complex networks learning with video classification based objectives. Taking all frames as positive samples, few of them pay attention to the discrimination between positive samples (salient frames) and negative samples (non-salient frames) in supervisions. To fill this gap, in this paper, we propose a novel Non-saliency Suppression Network (NSNet), which effectively suppresses the responses of non-salient frames. Specifically, on the frame level, effective pseudo labels that can distinguish between salient and non-salient frames are generated to guide the frame saliency learning. On the video level, a temporal attention module is learned under dual video-level supervisions on both the salient and the non-salient representations. Saliency measurements from both two levels are combined for exploitation of multi-granularity complementary information. Extensive experiments conducted on four well-known benchmarks verify our NSNet not only achieves the state-of-the-art accuracy-efficiency trade-off but also present a significantly faster (2.4~4.3x) practical inference speed than state-of-the-art methods. Our project page is at https://lawrencexia2008.github.io/projects/nsnet .

cs.CV↗