Searcharxiv⌕ Search

arXiv subjects

Yu Pan

Publications and source records attributed to Yu Pan.

At least 109 records · Page 6Linked to original sources

Preparing Lessons for Progressive Training on Language Models

The rapid progress of Transformers in artificial intelligence has come at the cost of increased resource consumption and greenhouse gas emissions due to growing model sizes. Prior work suggests using pretrained small models to improve training efficiency, but this approach may not be suitable for new model structures. On the other hand, training from scratch can be slow, and progressively stacking layers often fails to achieve significant acceleration. To address these challenges, we propose a novel method called Apollo, which prep\textbf{a}res lessons for ex\textbf{p}anding \textbf{o}perations by \textbf{l}earning high-\textbf{l}ayer functi\textbf{o}nality during training of low layers. Our approach involves low-value-prioritized sampling (LVPS) to train different depths and weight sharing to facilitate efficient expansion. We also introduce an interpolation method for stable model depth extension. Experiments demonstrate that Apollo achieves state-of-the-art acceleration ratios, even rivaling methods using pretrained models, making it a universal and efficient solution for training deep models while reducing time, financial, and environmental costs.

cs.LG↗

Cumulative Distribution Function based General Temporal Point Processes

Temporal Point Processes (TPPs) hold a pivotal role in modeling event sequences across diverse domains, including social networking and e-commerce, and have significantly contributed to the advancement of recommendation systems and information retrieval strategies. Through the analysis of events such as user interactions and transactions, TPPs offer valuable insights into behavioral patterns, facilitating the prediction of future trends. However, accurately forecasting future events remains a formidable challenge due to the intricate nature of these patterns. The integration of Neural Networks with TPPs has ushered in the development of advanced deep TPP models. While these models excel at processing complex and nonlinear temporal data, they encounter limitations in modeling intensity functions, grapple with computational complexities in integral computations, and struggle to capture long-range temporal dependencies effectively. In this study, we introduce the CuFun model, representing a novel approach to TPPs that revolves around the Cumulative Distribution Function (CDF). CuFun stands out by uniquely employing a monotonic neural network for CDF representation, utilizing past events as a scaling factor. This innovation significantly bolsters the model's adaptability and precision across a wide range of data scenarios. Our approach addresses several critical issues inherent in traditional TPP modeling: it simplifies log-likelihood calculations, extends applicability beyond predefined density function forms, and adeptly captures long-range temporal patterns. Our contributions encompass the introduction of a pioneering CDF-based TPP model, the development of a methodology for incorporating past event information into future event prediction, and empirical validation of CuFun's effectiveness through extensive experimentation on synthetic and real-world datasets.

cs.LG↗

Quantity-Aware Coarse-to-Fine Correspondence for Image-to-Point Cloud Registration

Image-to-point cloud registration aims to determine the relative camera pose between an RGB image and a reference point cloud, serving as a general solution for locating 3D objects from 2D observations. Matching individual points with pixels can be inherently ambiguous due to modality gaps. To address this challenge, we propose a framework to capture quantity-aware correspondences between local point sets and pixel patches and refine the results at both the point and pixel levels. This framework aligns the high-level semantics of point sets and pixel patches to improve the matching accuracy. On a coarse scale, the set-to-patch correspondence is expected to be influenced by the quantity of 3D points. To achieve this, a novel supervision strategy is proposed to adaptively quantify the degrees of correlation as continuous values. On a finer scale, point-to-pixel correspondences are refined from a smaller search space through a well-designed scheme, which incorporates both resampling and quantity-aware priors. Particularly, a confidence sorting strategy is proposed to proportionally select better correspondences at the final stage. Leveraging the advantages of high-quality correspondences, the problem is successfully resolved using an efficient Perspective-n-Point solver within the framework of random sample consensus (RANSAC). Extensive experiments on the KITTI Odometry and NuScenes datasets demonstrate the superiority of our method over the state-of-the-art methods.

cs.CV↗

Tensorized Hypergraph Neural Networks

Hypergraph neural networks (HGNN) have recently become attractive and received significant attention due to their excellent performance in various domains. However, most existing HGNNs rely on first-order approximations of hypergraph connectivity patterns, which ignores important high-order information. To address this issue, we propose a novel adjacency-tensor-based \textbf{T}ensorized \textbf{H}ypergraph \textbf{N}eural \textbf{N}etwork (THNN). THNN is a faithful hypergraph modeling framework through high-order outer product feature message passing and is a natural tensor extension of the adjacency-matrix-based graph neural networks. The proposed THNN is equivalent to a high-order polynomial regression scheme, which enables THNN with the ability to efficiently extract high-order information from uniform hypergraphs. Moreover, in consideration of the exponential complexity of directly processing high-order outer product features, we propose using a partially symmetric CP decomposition approach to reduce model complexity to a linear degree. Additionally, we propose two simple yet effective extensions of our method for non-uniform hypergraphs commonly found in real-world applications. Results from experiments on two widely used {hypergraph datasets for 3-D visual object classification} show the model's promising performance.

cs.AI↗

PromptVC: Flexible Stylistic Voice Conversion in Latent Space Driven by Natural Language Prompts

Style voice conversion aims to transform the style of source speech to a desired style according to real-world application demands. However, the current style voice conversion approach relies on pre-defined labels or reference speech to control the conversion process, which leads to limitations in style diversity or falls short in terms of the intuitive and interpretability of style representation. In this study, we propose PromptVC, a novel style voice conversion approach that employs a latent diffusion model to generate a style vector driven by natural language prompts. Specifically, the style vector is extracted by a style encoder during training, and then the latent diffusion model is trained independently to sample the style vector from noise, with this process being conditioned on natural language prompts. To improve style expressiveness, we leverage HuBERT to extract discrete tokens and replace them with the K-Means center embedding to serve as the linguistic content, which minimizes residual style information. Additionally, we deduplicate the same discrete token and employ a differentiable duration predictor to re-predict the duration of each token, which can adapt the duration of the same linguistic content to different styles. The subjective and objective evaluation results demonstrate the effectiveness of our proposed system.

eess.AS↗

GEmo-CLAP: Gender-Attribute-Enhanced Contrastive Language-Audio Pretraining for Accurate Speech Emotion Recognition

Contrastive cross-modality pretraining has recently exhibited impressive success in diverse fields, whereas there is limited research on their merits in speech emotion recognition (SER). In this paper, we propose GEmo-CLAP, a kind of gender-attribute-enhanced contrastive language-audio pretraining (CLAP) method for SER. Specifically, we first construct an effective emotion CLAP (Emo-CLAP) for SER, using pre-trained text and audio encoders. Second, given the significance of gender information in SER, two novel multi-task learning based GEmo-CLAP (ML-GEmo-CLAP) and soft label based GEmo-CLAP (SL-GEmo-CLAP) models are further proposed to incorporate gender information of speech signals, forming more reasonable objectives. Experiments on IEMOCAP indicate that our proposed two GEmo-CLAPs consistently outperform Emo-CLAP with different pre-trained models. Remarkably, the proposed WavLM-based SL-GEmo-CLAP obtains the best WAR of 83.16\%, which performs better than state-of-the-art SER methods.

cs.CL↗

Assessing Deep Neural Networks as Probability Estimators

Deep Neural Networks (DNNs) have performed admirably in classification tasks. However, the characterization of their classification uncertainties, required for certain applications, has been lacking. In this work, we investigate the issue by assessing DNNs' ability to estimate conditional probabilities and propose a framework for systematic uncertainty characterization. Denoting the input sample as x and the category as y, the classification task of assigning a category y to a given input x can be reduced to the task of estimating the conditional probabilities p(y|x), as approximated by the DNN at its last layer using the softmax function. Since softmax yields a vector whose elements all fall in the interval (0, 1) and sum to 1, it suggests a probabilistic interpretation to the DNN's outcome. Using synthetic and real-world datasets, we look into the impact of various factors, e.g., probability density f(x) and inter-categorical sparsity, on the precision of DNNs' estimations of p(y|x), and find that the likelihood probability density and the inter-categorical sparsity have greater impacts than the prior probability to DNNs' classification uncertainty.

cs.LG↗

Transforming Agriculture with Intelligent Data Management and Insights

Modern agriculture faces grand challenges to meet increased demands for food, fuel, feed, and fiber with population growth under the constraints of climate change and dwindling natural resources. Data innovation is urgently required to secure and improve the productivity, sustainability, and resilience of our agroecosystems. As various sensors and Internet of Things (IoT) instrumentation become more available, affordable, reliable, and stable, it has become possible to conduct data collection, integration, and analysis at multiple temporal and spatial scales, in real-time, and with high resolutions. At the same time, the sheer amount of data poses a great challenge to data storage and analysis, and the \textit{de facto} data management and analysis practices adopted by scientists have become increasingly inefficient. Additionally, the data generated from different disciplines, such as genomics, phenomics, environment, agronomy, and socioeconomic, can be highly heterogeneous. That is, datasets across disciplines often do not share the same ontology, modality, or format. All of the above make it necessary to design a new data management infrastructure that implements the principles of Findable, Accessible, Interoperable, and Reusable (FAIR). In this paper, we propose Agriculture Data Management and Analytics (ADMA), which satisfies the FAIR principles. Our new data management infrastructure is intelligent by supporting semantic data management across disciplines, interactive by providing various data management/analysis portals such as web GUI, command line, and API, scalable by utilizing the power of high-performance computing (HPC), extensible by allowing users to load their own data analysis tools, trackable by keeping track of different operations on each file, and open by using a rich set of mature open source technologies.

cs.DB↗

Testing the spatial geometry of the universe with TianQin: the prospect of using supermassive black hole binaries

The determination of the spatial geometry of the universe plays an important role in modern cosmology. Any deviation from the cosmic curvature $Ω_K=0$ would have a profound impact on the primordial inflation paradigm and fundamental physics. In this paper, we carry out a systematic study of the prospect of measuring cosmic curvature with the inspiral signal of supermassive black hole binaries (SMBHBs) that could be detected with TianQin. The study is based on a cosmological-model-independent method that extended the application of gravitational wave (GW) standard sirens in cosmology. By comparing the distances from future simulated GW events and simulated $H(z)$ data, we evaluate if TianQin would produce robust constraints on the cosmic curvature parameter $Ω_{k}$. More specifically, we consider 3-yr to 10-yr observations of supermassive black hole binaries with total masses ranging from $10^{3}M_\odot$ to $10^{7}M_\odot$. Our results show that in the future, with the synergy of 10-yr high-quality observations, we can tightly constrain the curvature parameter at the level of $1σ$ $Ω_k=-0.002\pm0.061$. Moreover, our findings indicate that the total mass of SMBHB does influence the estimation of cosmic curvature, implied by the analysis performed on different subsamples of gravitational wave data. Therefore, TianQin is expected to provide a powerful and competitive probe of the spatial geometry of the universe, compared to future spaced-based detectors such as DECIGO.

astro-ph.CO↗

Reusing Pretrained Models by Multi-linear Operators for Efficient Training

Training large models from scratch usually costs a substantial amount of resources. Towards this problem, recent studies such as bert2BERT and LiGO have reused small pretrained models to initialize a large model (termed the ``target model''), leading to a considerable acceleration in training. Despite the successes of these previous studies, they grew pretrained models by mapping partial weights only, ignoring potential correlations across the entire model. As we show in this paper, there are inter- and intra-interactions among the weights of both the pretrained and the target models. As a result, the partial mapping may not capture the complete information and lead to inadequate growth. In this paper, we propose a method that linearly correlates each weight of the target model to all the weights of the pretrained model to further enhance acceleration ability. We utilize multi-linear operators to reduce computational and spacial complexity, enabling acceptable resource requirements. Experiments demonstrate that our method can save 76\% computational costs on DeiT-base transferred from DeiT-small, which outperforms bert2BERT by +12.0\% and LiGO by +20.7\%, respectively.

cs.LG↗

Local to Global: A Distributed Quantum Approximate Optimization Algorithm for Pseudo-Boolean Optimization Problems

With the rapid advancement of quantum computing, Quantum Approximate Optimization Algorithm (QAOA) is considered as a promising candidate to demonstrate quantum supremacy, which exponentially solves a class of Quadratic Unconstrained Binary Optimization (QUBO) problems. However, limited qubit availability and restricted coherence time challenge QAOA to solve large-scale pseudo-Boolean problems on currently available Near-term Intermediate Scale Quantum (NISQ) devices. In this paper, we propose a distributed QAOA which can solve a general pseudo-Boolean problem by converting it to a simplified Ising model. Different from existing distributed QAOAs' assuming that local solutions are part of a global one, which is not often the case, we introduce community detection using Louvian algorithm to partition the graph where subgraphs are further compressed by community representation and merged into a higher level subgraph. Recursively and backwards, local solutions of lower level subgraphs are updated by heuristics from solutions of higher level subgraphs. Compared with existing methods, our algorithm incorporates global heuristics into local solutions such that our algorithm is proven to achieve a higher approximation ratio and outperforms across different graph configurations. Also, ablation studies validate the effectiveness of each component in our method.

quant-ph↗

Quantum Approximate Optimization Algorithm in Non-Markovian Quantum Systems

Although quantum approximate optimization algorithm (QAOA) has demonstrated its quantum supremacy, its performance on Noisy Intermediate-Scale Quantum (NISQ) devices would be influenced by complicated noises, e.g., quantum colored noises. To evaluate the performance of QAOA under these noises, this paper presents a framework for running QAOA on non-Markovian quantum systems which are represented by an augmented system model. In this model, a non-Markovian environment carrying quantum colored noises is modelled as an ancillary system driven by quantum white noises which is directly coupled to the corresponding principal system; i.e., the computational unit for the algorithm. With this model, we mathematically formulate QAOA as piecewise Hamiltonian control of the augmented system, where we also optimize the control depth to fit into the circuit depth of current quantum devices. For efficient simulation of QAOA in non-Markovian quantum systems, a boosted algorithm using quantum trajectory is further presented. Finally, we show that non-Markovianity can be utilized as a quantum resource to achieve a relatively good performance of QAOA, which is characterized by our proposed exploration rate.

quant-ph↗

On statistical fluctuations in collective flows

In relativistic heavy-ion collisions, event-by-event fluctuations are known to have non-trivial implications. Even though the probability distribution is geometrically isotropic for the initial conditions, the anisotropic $\varepsilon_n$ still differs from zero owing to the statistical fluctuations in the energy profile. On the other hand, the flow harmonics extracted from the hadron spectrum using the multi-particle correlators are inevitably subjected to non-vanishing variance due to the finite number of hadrons emitted in individual events. As one aims to extract information on the fluctuations in the initial conditions via flow harmonics and their fluctuations, finite multiplicity may play a role in interfering with such an effort. In this study, we explore the properties and impacts of such fluctuations in the initial and final states, which both notably appear to be statistical ones originating from the finite number of quanta of the underlying system. We elaborate on the properties of the initial-state eccentricities for the smooth and event-by-event fluctuating initial conditions and their distinct impacts on the resulting flow harmonics. Numerical simulations are performed. The possible implications of the present study are also addressed.

nucl-th↗

HYBRIDFORMER: improving SqueezeFormer with hybrid attention and NSR mechanism

SqueezeFormer has recently shown impressive performance in automatic speech recognition (ASR). However, its inference speed suffers from the quadratic complexity of softmax-attention (SA). In addition, limited by the large convolution kernel size, the local modeling ability of SqueezeFormer is insufficient. In this paper, we propose a novel method HybridFormer to improve SqueezeFormer in a fast and efficient way. Specifically, we first incorporate linear attention (LA) and propose a hybrid LASA paradigm to increase the model's inference speed. Second, a hybrid neural architecture search (NAS) guided structural re-parameterization (SRep) mechanism, termed NSR, is proposed to enhance the ability of the model to extract local interactions. Extensive experiments conducted on the LibriSpeech dataset demonstrate that our proposed HybridFormer can achieve a 9.1% relative word error rate (WER) reduction over SqueezeFormer on the test-other dataset. Furthermore, when input speech is 30s, the HybridFormer can improve the model's inference speed up to 18%. Our source code is available online.

eess.AS↗

Single-photon Image Super-resolution via Self-supervised Learning

Single-Photon Image Super-Resolution (SPISR) aims to recover a high-resolution volumetric photon counting cube from a noisy low-resolution one by computational imaging algorithms. In real-world scenarios, pairs of training samples are often expensive or impossible to obtain. By extending Equivariant Imaging (EI) to volumetric single-photon data, we propose a self-supervised learning framework for the SPISR task. Particularly, using the Poisson unbiased Kullback-Leibler risk estimator and equivariance, our method is able to learn from noisy measurements without ground truths. Comprehensive experiments on simulated and real-world dataset demonstrate that the proposed method achieves comparable performance with supervised learning and outperforms interpolation-based methods.

eess.IV↗

Augmentations and immersed Lagrangian fillings

For a Legendrian link $Λ\subset J^1M$ with $M = \mathbb{R}$ or $S^1$, immersed exact Lagrangian fillings $L \subset \mbox{Symp}(J^1M) \cong T^*(\mathbb{R}_{>0} \times M)$ of $Λ$ can be lifted to conical Legendrian fillings $Σ\subset J^1(\mathbb{R}_{>0} \times M)$ of $Λ$. When $Σ$ is embedded, using the version of functoriality for Legendrian contact homology (LCH) from [30], for each augmentation $α: \mathcal{A}(Σ) \rightarrow \mathbb{Z}/2$ of the LCH algebra of $Σ$, there is an induced augmentation $ε_{(Σ,α)}: \mathcal{A}(Λ) \rightarrow \mathbb{Z}/2$. With $Σ$ fixed, the set of homotopy classes of all such induced augmentations, $I_Σ\subset \mathit{Aug}(Λ)/{\sim}$, is a Legendrian isotopy invariant of $Σ$. We establish methods to compute $I_Σ$ based on the correspondence between Morse complex families and augmentations. This includes developing a functoriality for the cellular DGA from [31] with respect to Legendrian cobordisms, and proving its equivalence to the functoriality for LCH. For arbitrary $n \geq 1$, we give examples of Legendrian torus knots with $2n$ distinct conical Legendrian fillings distinguished by their induced augmentation sets. We prove that when $ρ\neq 1$ and $Λ\subset J^1\mathbb{R}$ every $ρ$-graded augmentation of $Λ$ can be induced in this manner by an immersed Lagrangian filling. Alternatively, this is viewed as a computation of cobordism classes for an appropriate notion of $ρ$-graded augmented Legendrian cobordism.

math.SG↗

Using simulated Tianqin gravitational wave data and electromagnetic wave data to study the coincidence problem and Hubble tension problem

In this paper, we use electromagnetic wave data (H0LiCOW, $H(z)$, SNe) and gravitational wave data (Tianqin) to constrain the interacting dark energy (IDE) model and investigate the Hubble tension problem and coincidences problem. By combining these four kinds of data (Tianqin+H0LiCOW+SNe+$H(z)$), we obtained the parameter values at the confidence interval of $1σ$: $Ω_m=0.36\pm0.18$, $ω_x=-1.29^{+0.61}_{-0.23}$, $ξ=3.15^{+0.36}_{-1.1}$, and $H_0=70.04\pm0.42$ $kms^{-1}Mpc^{-1}$. According to our results, the best valve of $H_0$ show that the Hubble tension problem can be alleviated to some extent. In addition, the $ξ+3ω_x = -0.72^{+2.19}_{-1.19}(1σ)$ of which the center value indicates the coincidence problem is slightly alleviated. However, the $ξ+3ω_x = 0$ is still within the $1σ$ error range which indicates the $Λ$CDM model is still the model which is in best agreement with the observational data at present. Finally, we compare the constraint results of electromagnetic wave and gravitational wave on the model parameters and find that the constraint effect of electromagnetic wave data on model parameters is better than that of simulated Tianqin gravitational wave data.

astro-ph.CO↗

LMEC: Learnable Multiplicative Absolute Position Embedding Based Conformer for Speech Recognition

This paper proposes a Learnable Multiplicative absolute position Embedding based Conformer (LMEC). It contains a kernelized linear attention (LA) module called LMLA to solve the time-consuming problem for long sequence speech recognition as well as an alternative to the FFN structure. First, the ELU function is adopted as the kernel function of our proposed LA module. Second, we propose a novel Learnable Multiplicative Absolute Position Embedding (LM-APE) based re-weighting mechanism that can reduce the well-known quadratic temporal-space complexity of softmax self-attention. Third, we use Gated Linear Units (GLU) to substitute the Feed Forward Network (FFN) for better performance. Extensive experiments have been conducted on the public LibriSpeech datasets. Compared to the Conformer model with cosFormer style linear attention, our proposed method can achieve up to 0.63% word-error-rate improvement on test-other and improve the inference speed by up to 13% (left product) and 33% (right product) on the LA module.

eess.AS↗