SearcharxivSearch

arXiv subjects

Qiang Niu

Publications and source records attributed to Qiang Niu.

16 recordsLinked to original sources

Learning Earthquake Wave Arrival Time Picking from Labels with Inaccuracies

Inaccurately labeled training data, or "label noise", poses a significant threat to the integrity of supervised machine learning models. This corruption directly degrades performance by teaching the model erroneous mappings between features and labels, which leads to poor generalization and reduced accuracy on properly labeled validation and test data. Current seismological applications mainly rely on large-scale training sets or data augmentation to reduce the label-noise impact, which can be labor-intensive and costly. Here, we introduce a Label Noise-Contrastive Robust Learning (LaNCoR) approach that can effectively handle noisy labels in seismic signal processing tasks, without requiring large-scale training datasets. In this approach, the input waveform feature and label representation distributions are aligned in the feature space to correct mislabeling and reduce its impact on the training process. We present LaNCoR's performance on the task of P-phase arrival-time picking of real microseismic data using two baseline models and training approaches. Our results indicate that LaNCoR can improve performance by up to 28.8% across performance metrics. This approach holds great promise for model training in seismology and geosciences.

cs.LG

Arnoldi-Enhanced Multivariate Hermite Interpolation of Manifold-Valued Data

This paper presents a robust enhancement of the Tangent space Hermite Interpolation (THI) method for manifold-valued data by integrating the multivariate Arnoldi process. To circumvent the inherent numerical instability of multivariate confluent Vandermonde matrices, we use a $G$-Arnoldi-based recurrence to construct a discrete orthogonal polynomial basis directly on the tangent space. The method generates better numerical conditioning for high-order approximations. We analyze the convergence rates for both $C^0$ and $C^1$ errors in the multivariate setting. When only function values are used, the $C^0$ approximation error decays as $\mathcal{O}\left(\sqrt{M} n^{-m}\right)$. For the $C^1$ error without derivative data, the rate becomes $\mathcal{O}\left(\sqrt{M} h^{-1} n^{-m}\right)$, where $h$ is the fill distance of the sampling set. When derivative data are additionally available, the $C^1$ error is $\mathcal{O}\left(\sqrt{M} n^{-(m-1)}\right)$. In all cases, $n$ is the polynomial degree, $m$ denotes the regularity of the target function, and $M$ is the number of sampling points. Importantly, as $n$ increases, the required number of points $M$ must also increase. This reveals the interplay among approximation order, sampling density ($M$), fill distance ($h$), dimension ($d$), and the regularity ($m$) of the target function. Extensive numerical experiments conducted on the special orthogonal group $SO(3)$ and the unit sphere $S^2$ show that the Arnoldi-enhanced THI method outperforms the Kriging-based approaches in terms of both computational efficiency and accuracy.

math.NA

Stable High-Order Interpolation on the Grassmann Manifold by Maximum-Volume Coordinates and Arnoldi Orthogonalization

High-order interpolation on the Grassmann manifold $\Gr(n, p)$ is often hindered by the computational overhead and derivative instability of SVD-based geometric mappings. To solve the challenges, we propose a stabilized framework that combines Maximum-Volume (MV) local coordinates with Arnoldi-orthogonalized polynomial bases. First, manifold data are mapped to a well-conditioned Euclidean domain via MV coordinates. The approach bypasses the costly matrix factorizations inherent to traditional Riemannian normal coordinates. Within the coordinate space, we use the Vandermonde-with-Arnoldi (V+A) method for Lagrange interpolation and its confluent extension (CV+A) for derivative-enriched Hermite interpolation. By constructing discrete orthogonal bases directly from the parameter nodes, the solution of ill-conditioned linear system is avoided. Theoretical bounds are established to verify the stability of the geometric mapping and the polynomial approximation. Extensive numerical experiments demonstrate that the proposed MV-(C)V+A framework can produce highly accurate approximation in high-degree polynomial interpolation.

math.NA

Measurement and Modeling of Structure-Induced Surface Scattering on Terahertz Channel

As terahertz (THz) frequencies emerge as promising candidates for next-generation wireless networks, accurate characterization of propagation mechanisms in indoor/outdoor environments becomes essential for system design and performance optimization. This article presents an experimental and theoretical investigation of structure-induced indoor surface scattering on THz channels, examining how material properties and structural configurations jointly govern channel power and angular distribution. Six representative indoor surfaces are characterized, revealing that intrinsic structural inhomogeneity -- particularly the quasi-periodic earlywood-latewood arrangement in pine wood -- induces measurable angular scattering whose dominant lobes and angular shifts are reproduced by a beam-propagation modeling (BPM) framework. Material-covered surface configurations are further investigated, demonstrating that thin dielectric covering layers can substantially modify reflection characteristics through thickness- and frequency- dependent thin-film interference effects. Wide-angle bistatic measurements conducted in a conference-room environment reveal that structured indoor elements, such as folded curtains, can enhance angular scattering and extend spatial coverage. These findings establish that structure-induced surface scattering mechanisms offer potential for constructing non-line-of-sight THz links in indoor environments.

physics.app-ph

Eavesdropping Risk in Terahertz Channels by Covered Wavy Surfaces

Terahertz communications offer unprecedented data rates for next-generation wireless networks but suffer blockage susceptibility that restrict coverage and introduce physical-layer security vulnerabilities. Non-line-of-sight relay schemes using metallic wavy surfaces (MWS) address coverage limitations but require concealment beneath indoor materials for practical deployment. This work investigates THz channel characteristics and security vulnerabilities when MWS surfaces are covered with wallpaper, curtain, and wall plaster across 113-170 GHz. Results reveal that covering materials redistribute rather than eliminate eavesdropping threats, with persistent feasible interception scenarios remaining undetectable through conventional backscattering monitoring. These findings underscore the need for enhanced mechanisms designed for covered reflecting elements.

physics.app-ph

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data

The GPT-4o represents a significant milestone in enabling real-time interaction with large language models (LLMs) through speech, its remarkable low latency and high fluency not only capture attention but also stimulate research interest in the field. This real-time speech interaction is particularly valuable in scenarios requiring rapid feedback and immediate responses, dramatically enhancing user experience. However, there is a notable lack of research focused on real-time large speech language models, particularly for Chinese. In this work, we present KE-Omni, a seamless large speech language model built upon Ke-SpeechChat, a large-scale high-quality synthetic speech interaction dataset consisting of 7 million Chinese and English conversations, featuring 42,002 speakers, and totaling over 60,000 hours, This contributes significantly to the advancement of research and development in this field. The demos can be accessed at \url{https://huggingface.co/spaces/KE-Team/KE-Omni}.

cs.CL

SeisT: A foundational deep learning model for earthquake monitoring tasks

Seismograms, the fundamental seismic records, have revolutionized earthquake research and monitoring. Recent advancements in deep learning have further enhanced seismic signal processing, leading to even more precise and effective earthquake monitoring capabilities. This paper introduces a foundational deep learning model, the Seismogram Transformer (SeisT), designed for a variety of earthquake monitoring tasks. SeisT combines multiple modules tailored to different tasks and exhibits impressive out-of-distribution generalization performance, outperforming or matching state-of-the-art models in tasks like earthquake detection, seismic phase picking, first-motion polarity classification, magnitude estimation, back-azimuth estimation, and epicentral distance estimation. The performance scores on the tasks are 0.96, 0.96, 0.68, 0.95, 0.86, 0.55, and 0.81, respectively. The most significant improvements, in comparison to existing models, are observed in phase-P picking, phase-S picking, and magnitude estimation, with gains of 1.7%, 9.5%, and 8.0%, respectively. Our study, through rigorous experiments and evaluations, suggests that SeisT has the potential to contribute to the advancement of seismic signal processing and earthquake research.

physics.geo-ph

Adaptive Softassign via Hadamard-Equipped Sinkhorn

Softassign is a pivotal method in graph matching and other learning tasks. Many softassign-based algorithms exhibit performance sensitivity to a parameter in the softassign. However, tuning the parameter is challenging and almost done empirically. This paper proposes an adaptive softassign method for graph matching by analyzing the relationship between the objective score and the parameter. This method can automatically tune the parameter based on a given error bound to guarantee accuracy. The Hadamard-Equipped Sinkhorn formulas introduced in this study significantly enhance the efficiency and stability of the adaptive softassign. Moreover, these formulas can also be used in optimal transport problems. The resulting adaptive softassign graph matching algorithm enjoys significantly higher accuracy than previous state-of-the-art large graph matching algorithms while maintaining comparable efficiency.

math.OC

Towards Better Instruction Following Language Models for Chinese: Investigating the Impact of Training Data and Evaluation

Recently, significant public efforts have been directed towards developing low-cost models with capabilities akin to ChatGPT, thereby fostering the growth of open-source conversational models. However, there remains a scarcity of comprehensive and in-depth evaluations of these models' performance. In this study, we examine the influence of training data factors, including quantity, quality, and linguistic distribution, on model performance. Our analysis is grounded in several publicly accessible, high-quality instruction datasets, as well as our own Chinese multi-turn conversations. We assess various models using a evaluation set of 1,000 samples, encompassing nine real-world scenarios. Our goal is to supplement manual evaluations with quantitative analyses, offering valuable insights for the continued advancement of open-source chat models. Furthermore, to enhance the performance and training and inference efficiency of models in the Chinese domain, we extend the vocabulary of LLaMA - the model with the closest open-source performance to proprietary language models like GPT-3 - and conduct secondary pre-training on 3.4B Chinese words. We make our model, data, as well as code publicly available.

cs.CL

Exploring the Impact of Instruction Data Scaling on Large Language Models: An Empirical Study on Real-World Use Cases

The success of ChatGPT has recently attracted numerous efforts to replicate it, with instruction-tuning strategies being a key factor in achieving remarkable results. Instruction-tuning not only significantly enhances the model's performance and generalization but also makes the model's generated results more consistent with human speech patterns. However current research rarely studies the impact of different amounts of instruction data on model performance, especially in the real-world use cases. In this paper we explore the performance of large language models based on instruction tuning across different scales of instruction data. An evaluation dataset consisting of 12 major online use cases is constructed in the experiment. With Bloomz-7B1-mt as the base model, the results show that 1) merely increasing the amount of instruction data leads to continuous improvement in tasks such as open-ended generation, 2) in tasks such as math and code, the model performance curve remains quite flat while increasing data size. We further analyze the possible causes of these phenomena and propose potential future research directions such as effectively selecting high-quality training data, scaling base models and training methods specialized for hard tasks. We will release our training and evaluation datasets, as well as model checkpoints.

cs.CL

Modeling Randomly Walking Volatility with Chained Gamma Distributions

Volatility clustering is a common phenomenon in financial time series. Typically, linear models can be used to describe the temporal autocorrelation of the (logarithmic) variance of returns. Considering the difficulty in estimating this model, we construct a Dynamic Bayesian Network, which utilizes the conjugate prior relation of normal-gamma and gamma-gamma, so that its posterior form locally remains unchanged at each node. This makes it possible to find approximate solutions using variational methods quickly. Furthermore, we ensure that the volatility expressed by the model is an independent incremental process after inserting dummy gamma nodes between adjacent time steps. We have found that this model has two advantages: 1) It can be proved that it can express heavier tails than Gaussians, i.e., have positive excess kurtosis, compared to popular linear models. 2) If the variational inference(VI) is used for state estimation, it runs much faster than Monte Carlo(MC) methods since the calculation of the posterior uses only basic arithmetic operations. And its convergence process is deterministic. We tested the model, named Gam-Chain, using recent Crypto, Nasdaq, and Forex records of varying resolutions. The results show that: 1) In the same case of using MC, this model can achieve comparable state estimation results with the regular lognormal chain. 2) In the case of only using VI, this model can obtain accuracy that are slightly worse than MC, but still acceptable in practice; 3) Only using VI, the running time of Gam-Chain, in general case, can be reduced to below 5% of that based on the lognormal chain via MC.

q-fin.CP

Sub-GMN: The Neural Subgraph Matching Network Model

As one of the most fundamental tasks in graph theory, subgraph matching is a crucial task in many fields, ranging from information retrieval, computer vision, biology, chemistry and natural language processing. Yet subgraph matching problem remains to be an NP-complete problem. This study proposes an end-to-end learning-based approximate method for subgraph matching task, called subgraph matching network (Sub-GMN). The proposed Sub-GMN firstly uses graph representation learning to map nodes to node-level embedding. It then combines metric learning and attention mechanisms to model the relationship between matched nodes in the data graph and query graph. To test the performance of the proposed method, we applied our method on two databases. We used two existing methods, GNN and FGNN as baseline for comparison. Our experiment shows that, on dataset 1, on average the accuracy of Sub-GMN are 12.21\% and 3.2\% higher than that of GNN and FGNN respectively. On average running time Sub-GMN runs 20-40 times faster than FGNN. In addition, the average F1-score of Sub-GMN on all experiments with dataset 2 reached 0.95, which demonstrates that Sub-GMN outputs more correct node-to-node matches. Comparing with the previous GNNs-based methods for subgraph matching task, our proposed Sub-GMN allows varying query and data graphes in the test/application stage, while most previous GNNs-based methods can only find a matched subgraph in the data graph during the test/application for the same query graph used in the training stage. Another advantage of our proposed Sub-GMN is that it can output a list of node-to-node matches, while most existing end-to-end GNNs based methods cannot provide the matched node pairs.

cs.LG

CSGO: Constrained-Softassign Gradient Optimization For Large Graph Matching

Graph matching aims to find correspondences between two graphs. This paper integrates several well-known graph matching algorithms into a framework: the constrained gradient method. The primary difference among these algorithms lies in tuning a step size parameter and constraining operators. By leveraging these insights, we propose an adaptive step size parameter to guarantee the underlying algorithms' convergence, simultaneously enhancing their efficiency and robustness. For the constraining operator, we introduce a scalable softassign for large graph matching problems. Compared to the original softassign, our approach offers increased speed, improved robustness, and reduced risk of overflow. The advanced constraining operator enables a CSGO for large graph matching, which outperforms state-of-the-art methods in experiments. Notably, in attributed graph matching tasks, CSGO achieves an over 10X increase in speed compared to current constrained gradient algorithms.

math.CO

Confluent Vandermonde with Arnoldi

In this note, we extend the Vandermonde with Arnoldi method recently advocated by P. D. Brubeck, Y. Nakatsukasa and L. N. Trefethen to dealing with the confluent Vandermonde matrix. To apply the Arnoldi process, it is critical to find a Krylov subspace which generates the column space of the confluent Vandermonde matrix. A theorem is established for such Krylov subspaces for any order derivatives. This enables us to compute the derivatives of high degree polynomials to high precision. It also makes many applications involving derivatives possible, as illustrated by numerical examples. We note that one of the approaches orthogonalizes only the function values and is equivalent to the formula given by P. D. Brubeck and L. N. Trefethen. The other approach orthogonalizes the Hermite data. About which approach is preferable to another, we made the comparison, and the result is problem dependent.

math.NA

Fabricated Pictures Detection with Graph Matching

Fabricating experimental pictures in research work is a serious academic misconduct, which should better be detected in the reviewing process. However, due to large number of submissions, the detection whether a picture is fabricated or reused is laborious for reviewers, and sometimes is indistinct with human eyes. A tool for detecting similarity between images may help to alleviate this problem. Some methods based on local feature points matching work for most of the time, while these methods may result in mess of matchings due to ignorance of global relationship between features. We present a framework to detect similar, or perhaps fabricated, pictures with the graph matching techniques. A new iterative method is proposed, and experiments show that such a graph matching technique is better than the methods based only on local features for some cases.

cs.CV

A 3D Curve Offset Approach for Ruled Surface Generation in Engineering Design

Ruled surface is widely used in engineering design such as parting surface design of injection mold and checking surface design of checking fixture, which are usually generated by offsetting 3D curves. However, in 3D curve offset, there often exist break,interaction and overlapping problems which can't be solved by current CAD software automatically. This paper is targeted at developing a 3D curve offsetting algorithm for ruled surface generation, and three key technologies are introduced in details: An improved curve division method is proposed to reduce the offset accuracy error resulted from different offset distances and curvatures; An offsetting curve overlapping detection and elimination method is proposed; And then, a curve transition method is presented to improve curve offsetting quality for the break and intersection/overlapping regions, where a new algorithm for generating positive weights spherical rational quartic Bezier curve is proposed to bridge the breaks of offset curves to create a smooth ruled surface. Finally, two practical design cases, parting surface and checking surface generation, show that the proposed approach can enhance the efficiency and quality for ruled surface generation in engineering design.

cs.CG