SearcharxivSearch

arXiv subjects

Shijun Cheng

Publications and source records attributed to Shijun Cheng.

At least 19 recordsLinked to original sources

Generative wave propagator

Seismic wavefield simulation is fundamental to seismology, but conventional finite-difference (FD) methods remain limited by numerical dispersion and stability constraints, which often require dense spatial grids and small time steps and thereby severely limit the effectiveness of iterative inversion workflows. We introduce a conditional diffusion-based wavefield propagator that advances seismic wavefields recursively from one time step to the next. Instead of learning an unconditional data distribution of wavefield evolution, the model is conditioned by a short history of recent wavefield time steps (snapshots), the velocity model, and the wavefield time step index, allowing it to represent the conditional transition between adjacent physical states. By training the network to directly predict the clean next wavefield snapshot, this strong physical conditioning makes it possible to replace the iterative reverse diffusion process with a single network evaluation for each predicted snapshot. To improve stability over long recursive rollouts, we further introduce a causal time-weighted loss, in which adaptive weights, accumulated as exponential moving averages of per-snapshot training errors, emphasize training directions that are consistent with the forward propagation sequence and reduce the amplification of one-step prediction errors. Because the learned propagator is tied to the temporal spacing of the training snapshots rather than to the FD stability limit, it can advance the wavefield using a physical time step ten times larger than that required by the underlying solver. Experiments on the Overthrust, SEG/EAGE, and Marmousi models show that the proposed method accurately reproduces wavefield snapshots and shot gathers and achieves an end-to-end speedup of 2.17 x over a GPU-accelerated tenth-order staggered-grid FD implementation under matched hardware conditions.

physics.geo-ph

Incorporating wave physical priors into diffusion models: A novel approach to seismic resolution enhancement

Seismic resolution enhancement remains a critical challenge in exploration geophysics, particularly when processing field data characterized by limited bandwidth, strong noise, and insufficient labeled training samples. Existing deep learning methods typically rely on supervised learning with synthetic training data, leading to distribution mismatch and poor generalization on real seismic acquisitions. To address these limitations, we develop a physics-guided self-supervised diffusion model (PG-SSDM) that learns directly from field observations without requiring paired high-resolution labels. The proposed framework combines three key innovations. First, a self-supervised training strategy constructs learning targets by progressively filtering the observed data itself, eliminating the need for high-resolution ground truth through iterative refinement across multiple stages. Second, seismic convolution model is embedded as a hard physical constraint in both the training loss function and the reverse sampling process, ensuring that generated high-resolution outputs respect fundamental seismic wave propagation physics. Third, the probabilistic nature of diffusion models enables uncertainty quantification, providing spatial confidence maps that identify regions where resolution enhancement may be less reliable. We validate PG-SSDM on synthetic data under various noise conditions and on a 3D post-stack field dataset. Experimental results demonstrate that the proposed method effectively recovers thin layers and subtle structures, suppresses noise, preserves structural continuity, thereby significantly improving the resolution and interpretability of seismic data.

physics.geo-ph

Meta-learning-enhanced implicit full waveform inversion

Implicit full waveform inversion (IFWI) introduces implicit neural representations to parameterize the subsurface velocity model as a continuous function of spatial coordinates, which alleviates the dependence on the initial model and improves inversion flexibility. However, IFWI still requires a large number of iterative updates for each new exploration area, leading to slow convergence, high computational cost, and a lack of mechanisms to share prior knowledge across different geological settings, thereby limiting its efficiency and generalization capability. To further accelerate convergence and enhance cross-area generalization, we propose a meta-learning-based implicit full waveform inversion method, referred to as Meta-learning-enhanced implicit full waveform inversion (Meta-IFWI). In this framework, the subsurface velocity model is represented using an implicit neural network with periodic activation functions (SIREN), while a meta-learning strategy is employed to pretrain a single network on multiple velocity inversion tasks. Through this process, the network learns shared inversion priors and rapid adaptation strategies across different geological scenarios. For a new inversion task, the pretrained Meta-IFWI model can be efficiently adapted to the observed seismic data with only a few gradient updates, significantly reducing the number of iterations required for inversion. Numerical experiments conducted on in-distribution models, including layered synthetic models and the Overthrust model, as well as out-of-distribution complex models such as Marmousi 2, demonstrate that, compared with conventional IFWI, the proposed Meta-IFWI achieves improved inversion accuracy while substantially accelerating convergence and reducing computational cost. Moreover, Meta-IFWI exhibits enhanced robustness and stronger cross-area generalization capability.

physics.geo-ph

Adaptive Self-Supervised Surface-Related Multiple Suppression

Effective suppression of surface-related multiples is essential to prevent imaging artifacts and erroneous structural interpretations. While conventional approaches rely on accurate priors or subsurface model knowledge, and supervised learning methods require labeled data that are impractical to obtain for real seismic data. To overcome these limitations, a recently proposed self-supervised learning (SSL) framework integrates multi-dimensional convolution (MDC) for multiple generation with a two-stage training strategy, eliminating the need for both prior knowledge and labeled data. However, their approach requires manual selection of a scaling factor to match the amplitudes between the MDC-generated multiples and the true multiples, thus introducing subjectivity and limiting its practical applicability. In this study, we propose an adaptive SSL method that treats the scaling factor as a learnable parameter, jointly optimized with the network weights in a unified single-stage training pipeline. This dynamic scaling implicitly introduces amplitude diversity into the training data, acting as an implicit regularizer that improves the network's robustness to amplitude variations of surface-related multiples. We further design a composite loss function with homoscedastic uncertainty-based adaptive weighting, which automatically balances the contributions of multiple loss terms without manual tuning. Synthetic and field data examples demonstrate that our method robustly and effectively suppresses surface-related multiples while preserving primary reflections, with migration results confirming improved subsurface imaging quality.

physics.geo-ph

Propagating the prior from far to near offset: A self-supervised diffusion framework for progressively recovering near-offsets of towed-streamer data

In marine towed-streamer seismic acquisition, the nearest hydrophone is often two hundred meter away from the source resulting in missing near-offset traces, which degrades critical processing workflows such as surface-related multiple elimination, velocity analysis, and full-waveform inversion. Existing reconstruction methods, like transform-domain interpolation, often produce kinematic inconsistencies and amplitude distortions, while supervised deep learning approaches require complete ground-truth near-offset data that are unavailable in realistic acquisition scenarios. To address these limitations, we propose a self-supervised diffusion-based framework that reconstructs missing near-offset traces without requiring near-offset reference data. Our method leverages overlapping patch extraction with single-trace shifts from the available far-offset section to train a conditional diffusion model, which learns offset-dependent statistical patterns governing event curvature, amplitude variation, and wavelet characteristics. At inference, we perform trace-by-trace recursive extrapolation from the nearest recorded offset toward zero offset, progressively propagating learned prior information from far to near offsets. The generative formulation further provides uncertainty estimates via ensemble sampling, quantifying prediction confidence where validation data are absent. Controlled validation experiments on synthetic and field datasets show substantial performance gains over conventional parabolic Radon transform baselines. Operational deployment on actual near-offset gaps demonstrates practical viability where ground-truth validation is impossible. Notably, the reconstructed waveforms preserve realistic amplitude-versus-offset trends despite training exclusively on far-offset observations, and uncertainty maps accurately identify challenging extrapolation regions.

physics.geo-ph

Physics-informed conditional diffusion model for generalizable elastic wave-mode separation

Traditional elastic wavefield separation methods, while accurate, often demand substantial computational resources, especially for large geological models or 3D scenarios. Purely data-driven neural network approaches can be more efficient, but may fail to generalize and maintain physical consistency due to the absence of explicit physical constraints. Here, we propose a physics-informed conditional diffusion model for elastic wavefield separation that seamlessly integrates domain-specific physics equations into both the training and inference stages of the reverse diffusion process. Conditioned on full elastic wavefields and subsurface P- and S-wave velocity profiles, our method directly predicts clean P-wave modes while enforcing Laplacian separation constraints through physics-guided loss and sampling corrections. Numerical experiments on diverse scenarios yield the separation results that closely match conventional numerical solutions but at a reduced cost, confirming the effectiveness and generalizability of our approach.

physics.geo-ph

Inhomogeneous plane waves in attenuative anisotropic porous media

We investigate the propagation of inhomogeneous plane waves in poro-viscoelastic media, explicitly incorporating both velocity and attenuation anisotropy. Starting from classical Biot theory, we present a fractional differential equation describing wave propagation in attenuative anisotropic porous media that accommodates arbitrary anisotropy in both velocity and attenuation. Then, instead of relying on the traditional complex wave vector approach, we derive new Christoffel and energy balance equations for general inhomogeneous waves by employing an alternative formulation based on the complex slowness vector. The phase velocities and complex slownesses of inhomogeneous fast and slow quasi-compressional (qP1 and qP2) and quasi-shear (qS1 and qS2) waves are determined by solving an eighth-degree algebraic equation. By invoking the derived energy balance equation along with the computed complex slowness, we present explicit and concise expressions for energy velocities. Additionally, we analyze dissipation factors defined by two alternative measures: the ratio of average dissipated energy density to either average strain energy density or average stored energy density. We clarify and discuss the implications of these definitional differences in the context of general poro-viscoelastic anisotropic media. Finally, our expressions are degenerated to give their counterparts of the homogeneous waves as a special case, and the reduced forms are identical to those presented by the existing poro-viscoelastic theory. Several examples are provided to illustrate the propagation characteristics of inhomogeneous plane waves in unbounded attenuative vertical transversely isotropic porous media.

physics.geo-ph

DiffPINN: Generative diffusion-initialized physics-informed neural networks for accelerating seismic wavefield representation

Physics-informed neural networks (PINNs) offer a powerful framework for seismic wavefield modeling, yet they typically require time-consuming retraining when applied to different velocity models. Moreover, their training can suffer from slow convergence due to the complexity of of the wavefield solution. To address these challenges, we introduce a latent diffusion-based strategy for rapid and effective PINN initialization. First, we train multiple PINNs to represent frequency-domain scattered wavefields for various velocity models, then flatten each trained network's parameters into a one-dimensional vector, creating a comprehensive parameter dataset. Next, we employ an autoencoder to learn latent representations of these parameter vectors, capturing essential patterns across diverse PINN's parameters. We then train a conditional diffusion model to store the distribution of these latent vectors, with the corresponding velocity models serving as conditions. Once trained, this diffusion model can generate latent vectors corresponding to new velocity models, which are subsequently decoded by the autoencoder into complete PINN parameters. Experimental results indicate that our method significantly accelerates training and maintains high accuracy across in-distribution and out-of-distribution velocity scenarios.

physics.geo-ph

Self-supervised surface-related multiple suppression with multidimensional convolution

Surface-related multiples pose significant challenges in seismic data processing, often obscuring primary reflections and reducing imaging quality. Traditional methods rely on computationally expensive algorithms, the prior knowledge of subsurface model, or accurate wavelet estimation, while supervised learning approaches require clean labels, which are impractical for real data. Thus, we propose a self-supervised learning framework for surface-related multiple suppression, leveraging multi-dimensional convolution to generate multiples from the observed data and a two-stage training strategy comprising a warm-up and an iterative data refinement stage, so the network learns to remove the multiples. The framework eliminates the need for labeled data by iteratively refining predictions using multiples augmented inputs and pseudo-labels. Numerical examples demonstrate that the proposed method effectively suppresses surface-related multiples while preserving primary reflections. Migration results confirm its ability to reduce artifacts and improve imaging quality.

physics.geo-ph

SeparationPINN: Physics-Informed Neural Networks for Seismic P- and S-Wave Mode Separation

Accurate separation of P- and S-waves is essential for multi-component seismic data processing, as it helps eliminate interference between wave modes during imaging or inversion, which leads to high-accuracy results. Traditional methods for separating P- and S-waves rely on the Christoffel equation to compute the polarization direction of the waves in the wavenumber domain, which is computationally expensive. Although machine learning has been employed to improve the computational efficiency of the separation process, most methods still require supervised learning with labeled data, which is often unavailable for field data. To address this limitation, we propose a wavefield separation technique based on the physics-informed neural network (PINN). This unsupervised machine learning approach is applicable to unlabeled data. Furthermore, the trained PINN model provides a mesh-free numerical solution that effectively captures wavefield features at multiple scales. Numerical tests demonstrate that the proposed PINN-based separation method can accurately separate P- and S-waves in both homogeneous and heterogeneous media.

physics.geo-ph

Seismic wavefield solutions via physics-guided generative neural operator

Current neural operators often struggle to generalize to complex, out-of-distribution conditions, limiting their ability in seismic wavefield representation. To address this, we propose a generative neural operator (GNO) that leverages generative diffusion models (GDMs) to learn the underlying statistical distribution of scattered wavefields while incorporating a physics-guided sampling process at each inference step. This physics guidance enforces wave equation-based constraints corresponding to specific velocity models, driving the iteratively generated wavefields toward physically consistent solutions. By training the diffusion model on wavefields corresponding to a diverse dataset of velocity models, frequencies, and source positions, our GNO enables to rapidly synthesize high-fidelity wavefields at inference time. Numerical experiments demonstrate that our GNO not only produces accurate wavefields matching numerical reference solutions, but also generalizes effectively to previously unseen velocity models and frequencies.

physics.geo-ph

Generating Reliable Initial Velocity Models for Full-waveform Inversion with Well and Structural Constraints

Full waveform inversion (FWI) plays an important role in velocity modeling due to its high-resolution advantages. However, its highly non-linear characteristic leads to numerous local minimums, which is known as the cycle-skipping problem. Therefore, effectively addressing the cycle-skipping issue is crucial to the success of FWI. Well-log data contain rich information about subsurface medium parameters, providing inherent advantages for velocity modeling. Traditional well-log data interpolation methods to build velocity models have limited accuracy and poor adaptability to complex geological structures. This study introduces a well interpolation algorithm based on a generative diffusion model (GDM) to generate initial models for FWI, addressing the cycle-skipping problem. Existing convolutional neural network (CNN)-based methods face difficulties in handling complex feature distributions and lack effective uncertainty quantification, limiting the reliability of their outputs. The proposed GDM-based approach overcomes these challenges by providing geologically consistent well interpolation while incorporating uncertainty assessment. Numerical experiments demonstrate that the method produces accurate and reliable initial models, enhancing FWI performance and mitigating cycle-skipping issues.

physics.geo-ph

A generative foundation model for an all-in-one seismic processing framework

Seismic data often face challenges in their utilization due to noise contamination, incomplete acquisition, and limited low-frequency information, which hinder accurate subsurface imaging and interpretation. Traditional processing methods rely heavily on task-specific designs to address these challenges and fail to account for the variability of data. To address these limitations, we present a generative seismic foundation model (GSFM), a unified framework based on generative diffusion models (GDMs), designed to tackle multi-task seismic processing challenges, including denoising, backscattered noise attenuation, interpolation, and low-frequency extrapolation. GSFM leverages a pre-training stage on synthetic data to capture the features of clean, complete, and broadband seismic data distributions and applies an iterative fine-tuning strategy to adapt the model to field data. By adopting a target-oriented diffusion process prediction, GSFM improves computational efficiency without compromising accuracy. Synthetic data tests demonstrate GSFM surpasses benchmarks with equivalent architectures in all tasks and achieves performance comparable to traditional pre-training strategies, even after their fine-tuning. Also, field data tests suggest that our iterative fine-tuning approach addresses the generalization limitations of conventional pre-training and fine-tuning paradigms, delivering significantly enhanced performance across diverse tasks. Furthermore, GSFM's inherent probabilistic nature enables effective uncertainty quantification, offering valuable insights into the reliability of processing results.

physics.geo-ph

Multi-frequency wavefield solutions for variable velocity models using meta-learning enhanced low-rank physics-informed neural network

Physics-informed neural networks (PINNs) face significant challenges in modeling multi-frequency wavefields in complex velocity models due to their slow convergence, difficulty in representing high-frequency details, and lack of generalization to varying frequencies and velocity scenarios. To address these issues, we propose Meta-LRPINN, a novel framework that combines low-rank parameterization using singular value decomposition (SVD) with meta-learning and frequency embedding. Specifically, we decompose the weights of PINN's hidden layers using SVD and introduce an innovative frequency embedding hypernetwork (FEH) that links input frequencies with the singular values, enabling efficient and frequency-adaptive wavefield representation. Meta-learning is employed to provide robust initialization, improving optimization stability and reducing training time. Additionally, we implement adaptive rank reduction and FEH pruning during the meta-testing phase to further enhance efficiency. Numerical experiments, which are presented on multi-frequency scattered wavefields for different velocity models, demonstrate that Meta-LRPINN achieves much fast convergence speed and much high accuracy compared to baseline methods such as Meta-PINN and vanilla PINN. Also, the proposed framework shows strong generalization to out-of-distribution frequencies while maintaining computational efficiency. These results highlight the potential of our Meta-LRPINN for scalable and adaptable seismic wavefield modeling.

cs.LG

Propagating the prior from shallow to deep with a pre-trained velocity-model Generative Transformer network

Building subsurface velocity models is essential to our goals in utilizing seismic data for Earth discovery and exploration, as well as monitoring. With the dawn of machine learning, these velocity models (or, more precisely, their distribution) can be stored accurately and efficiently in a generative model. These stored velocity model distributions can be utilized to regularize or quantify uncertainties in inverse problems, like full waveform inversion. However, most generators, like normalizing flows or diffusion models, treat the image (velocity model) uniformly, disregarding spatial dependencies and resolution changes with respect to the observation locations. To address this weakness, we introduce VelocityGPT, a novel implementation that utilizes Transformer decoders trained autoregressively to generate a velocity model from shallow subsurface to deep. Owing to the fact that seismic data are often recorded on the Earth's surface, a top-down generator can utilize the inverted information in the shallow as guidance (prior) to generating the deep. To facilitate the implementation, we use an additional network to compress the velocity model. We also inject prior information, like well or structure (represented by a migration image) to generate the velocity model. Using synthetic data, we demonstrate the effectiveness of VelocityGPT as a promising approach in generative model applications for seismic velocity model building.

physics.geo-ph

Generative Diffusion Model for Seismic Imaging Improvement of Sparsely Acquired Data and Uncertainty Quantification

Seismic imaging from sparsely acquired data faces challenges such as low image quality, discontinuities, and migration swing artifacts. Existing convolutional neural network (CNN)-based methods struggle with complex feature distributions and cannot effectively assess uncertainty, making it hard to evaluate the reliability of their processed results. To address these issues, we propose a new method using a generative diffusion model (GDM). Here, in the training phase, we use the imaging results from sparse data as conditional input, combined with noisy versions of dense data imaging results, for the network to predict the added noise. After training, the network can predict the imaging results for test images from sparse data acquisition, using the generative process with conditional control. This GDM not only improves image quality and removes artifacts caused by sparse data, but also naturally evaluates uncertainty by leveraging the probabilistic nature of the GDM. To overcome the decline in generation quality and the memory burden of large-scale images, we develop a patch fusion strategy that effectively addresses these issues. Synthetic and field data examples demonstrate that our method significantly enhances imaging quality and provides effective uncertainty quantification.

physics.geo-ph

Discovery of physically interpretable wave equations

Using symbolic regression to discover physical laws from observed data is an emerging field. In previous work, we combined genetic algorithm (GA) and machine learning to present a data-driven method for discovering a wave equation. Although it managed to utilize the data to discover the two-dimensional (x,z) acoustic constant-density wave equation u_tt=v^2(u_xx+u_zz) (subscripts of the wavefield, u, are second derivatives in time and space) in a homogeneous medium, it did not provide the complete equation form, where the velocity term is represented by a coefficient rather than directly given by v^2. In this work, we redesign the framework, encoding both velocity information and candidate functional terms simultaneously. Thus, we use GA to simultaneously evolve the candidate functional and coefficient terms in the library. Also, we consider here the physics rationality and interpretability in the randomly generated potential wave equations, by ensuring that both-hand sides of the equation maintain balance in their physical units. We demonstrate this redesigned framework using the acoustic wave equation as an example, showing its ability to produce physically reasonable expressions of wave equations from noisy and sparsely observed data in both homogeneous and inhomogeneous media. Also, we demonstrate that our method can effectively discover wave equations from a more realistic observation scenario.

physics.geo-ph

Meta-PINN: Meta learning for improved neural network wavefield solutions

Physics-informed neural networks (PINNs) provide a flexible and effective alternative for estimating seismic wavefield solutions due to their typical mesh-free and unsupervised features. However, their accuracy and training cost restrict their applicability. To address these issues, we propose a novel initialization for PINNs based on meta learning to enhance their performance. In our framework, we first utilize meta learning to train a common network initialization for a distribution of medium parameters (i.e. velocity models). This phase employs a unique training data container, comprising a support set and a query set. We use a dual-loop approach, optimizing network parameters through a bidirectional gradient update from the support set to the query set. Following this, we use the meta-trained PINN model as the initial model for a regular PINN training for a new velocity model in which the optimization of the network is jointly constrained by the physical and regularization losses. Numerical results demonstrate that, compared to the vanilla PINN with random initialization, our method achieves a much fast convergence speed, and also, obtains a significant improvement in the results accuracy. Meanwhile, we showcase that our method can be integrated with existing optimal techniques to further enhance its performance.

physics.geo-ph