Searcharxiv⌕ Search

arXiv subjects

Yu Pan

Publications and source records attributed to Yu Pan.

At least 127 records · Page 7Linked to original sources

LMEC: Learnable Multiplicative Absolute Position Embedding Based Conformer for Speech Recognition

This paper proposes a Learnable Multiplicative absolute position Embedding based Conformer (LMEC). It contains a kernelized linear attention (LA) module called LMLA to solve the time-consuming problem for long sequence speech recognition as well as an alternative to the FFN structure. First, the ELU function is adopted as the kernel function of our proposed LA module. Second, we propose a novel Learnable Multiplicative Absolute Position Embedding (LM-APE) based re-weighting mechanism that can reduce the well-known quadratic temporal-space complexity of softmax self-attention. Third, we use Gated Linear Units (GLU) to substitute the Feed Forward Network (FFN) for better performance. Extensive experiments have been conducted on the public LibriSpeech datasets. Compared to the Conformer model with cosFormer style linear attention, our proposed method can achieve up to 0.63% word-error-rate improvement on test-other and improve the inference speed by up to 13% (left product) and 33% (right product) on the LA module.

eess.AS↗

Deep Domain Adversarial Adaptation for Photon-efficient Imaging

Photon-efficient imaging with the single-photon light detection and ranging (LiDAR) captures the three-dimensional (3D) structure of a scene by only a few detected signal photons per pixel. However, the existing computational methods for photon-efficient imaging are pre-tuned on a restricted scenario or trained on simulated datasets. When applied to realistic scenarios whose signal-to-background ratios (SBR) and other hardware-specific properties differ from those of the original task, the model performance often significantly deteriorates. In this paper, we present a domain adversarial adaptation design to alleviate this domain shift problem by exploiting unlabeled real-world data, with significant resource savings. This method demonstrates superior performance on simulated and real-world experiments using our home-built up-conversion single-photon imaging system, which provides an efficient approach to bypass the lack of ground-truth depth information in implementing computational imaging algorithms for realistic applications.

eess.IV↗

Testing the coincidence problem with strong gravitational lens, Type Ia supernovae and Hubble parameter observational data

In this paper, we use three different kinds of observational data, including 130 strong gravitational lensing (SGL) systems, type Ia supernovae (SNeIa: Pantheon and Union2.1) and 31 Hubble parameter data points ($H(z)$) from cosmic chronometers to constrain the phenomenological model ($ρ_x\varproptoρ_m a^ξ$). By combining these three kinds of data (Union2.1+SGL+$H(z)$), we get the parameter value at the confidence interval of $2σ$, $Ω_{X,0} = 0.69\pm0.34$, $ω_x = -1.24\pm0.61$, $ξ= 3.8\pm3.9$ and $H_0 = 70.22\pm0.86$ kms$^{-1}$Mpc$^{-1}$. According to our results, we find that the $Λ$CDM model is still the model which is in best agreement with the observational data at present, and the coincidence problem is not alleviated. In addition, the $Ω_X$ and $Ω_m$ have the same order of magnitude in $0 0.645$, it is necessary to introduce the dark energy interacting with dark matter.

astro-ph.CO↗

Matrix method for perturbed black hole metric with discontinuity

Recent studies based on the notion of black hole pseudospectrum indicated substantial instability of the fundamental and high-overtone quasinormal modes. Besides its theoretical novelty, the details about the migration of the quasinormal mode spectrum due to specific perturbations may furnish valuable information on the properties of associated gravitational waves in a more realistic context. This work generalizes the matrix method for black hole quasinormal modes to cope with a specific class of perturbations to the metric featured by discontinuity, which is known to be intimately connected with the quasinormal mode structural instability. In practice, the presence of discontinuity poses a difficulty so that many well-known approaches for quasinormal modes cannot be straightforwardly applied. By comparing with other methods, we show that the modified matrix method is efficient, which can be used to solve for the low-lying modes with reasonable precision. Therefore, it might serve as an alternative gadget for relevant studies.

gr-qc↗

Residual Tensor Train: A Quantum-inspired Approach for Learning Multiple Multilinear Correlations

States of quantum many-body systems are defined in a high-dimensional Hilbert space, where rich and complex interactions among subsystems can be modelled. In machine learning, complex multiple multilinear correlations may also exist within input features. In this paper, we present a quantum-inspired multilinear model, named Residual Tensor Train (ResTT), to capture the multiple multilinear correlations of features, from low to high orders, within a single model. ResTT is able to build a robust decision boundary in a high-dimensional space for solving fitting and classification tasks. In particular, we prove that the fully-connected layer and the Volterra series can be taken as special cases of ResTT. Furthermore, we derive the rule for weight initialization that stabilizes the training of ResTT based on a mean-field analysis. We prove that such a rule is much more relaxed than that of TT, which means ResTT can easily address the vanishing and exploding gradient problem that exists in the existing TT models. Numerical experiments demonstrate that ResTT outperforms the state-of-the-art tensor network and benchmark deep learning models on MNIST and Fashion-MNIST datasets. Moreover, ResTT achieves better performance than other statistical methods on two practical examples with limited data which are known to have complex feature interactions.

cs.LG↗

Obstructions to reversing Lagrangian surgery in Lagrangian fillings

Given an immersed, Maslov-$0$, exact Lagrangian filling of a Legendrian knot, if the filling has a vanishing index and action double point, then through Lagrangian surgery it is possible to obtain a new immersed, Maslov-$0$, exact Lagrangian filling with one less double point and with genus increased by one. We show that it is not always possible to reverse the Lagrangian surgery: not every immersed, Maslov-$0$, exact Lagrangian filling with genus $g \geq 1$ and $p$ double points can be obtained from such a Lagrangian surgery on a filling of genus $g-1$ with $p+1$ double points. To show this, we establish the connection between the existence of an immersed, Maslov-$0$, exact Lagrangian filling of a Legendrian $Λ$ that has $p$ double points with action $0$ and the existence of an embedded, Maslov-$0$, exact Lagrangian cobordism from $p$ copies of a Hopf link to $Λ$. We then prove that a count of augmentations provides an obstruction to the existence of embedded, Maslov-$0$, exact Lagrangian cobordisms between Legendrian links.

math.SG↗

A Unified Weight Initialization Paradigm for Tensorial Convolutional Neural Networks

Tensorial Convolutional Neural Networks (TCNNs) have attracted much research attention for their power in reducing model parameters or enhancing the generalization ability. However, exploration of TCNNs is hindered even from weight initialization methods. To be specific, general initialization methods, such as Xavier or Kaiming initialization, usually fail to generate appropriate weights for TCNNs. Meanwhile, although there are ad-hoc approaches for specific architectures (e.g., Tensor Ring Nets), they are not applicable to TCNNs with other tensor decomposition methods (e.g., CP or Tucker decomposition). To address this problem, we propose a universal weight initialization paradigm, which generalizes Xavier and Kaiming methods and can be widely applicable to arbitrary TCNNs. Specifically, we first present the Reproducing Transformation to convert the backward process in TCNNs to an equivalent convolution process. Then, based on the convolution operators in the forward and backward processes, we build a unified paradigm to control the variance of features and gradients in TCNNs. Thus, we can derive fan-in and fan-out initialization for various TCNNs. We demonstrate that our paradigm can stabilize the training of TCNNs, leading to faster convergence and better results.

cs.LG↗

Efficient Depth Selection for the Implementation of Noisy Quantum Approximate Optimization Algorithm

Noise on near-term quantum devices will inevitably limit the performance of Quantum Approximate Optimization Algorithm (QAOA). One significant consequence is that the performance of QAOA may fail to monotonically improve with depth. In particular, optimal depth can be found at a certain point where the noise effects just outweigh the benefits brought by increasing the depth. In this work, we propose to use the model selection algorithm to identify the optimal depth with a few iterations of regularization parameters. Numerical experiments show that the algorithm can efficiently locate the optimal depth under relaxation and dephasing noises.

quant-ph↗

Automatic Depth Optimization for Quantum Approximate Optimization Algorithm

Quantum Approximate Optimization Algorithm (QAOA) is a hybrid algorithm whose control parameters are classically optimized. In addition to the variational parameters, the right choice of hyperparameter is crucial for improving the performance of any optimization model. Control depth, or the number of variational parameters, is considered as the most important hyperparameter for QAOA. In this paper we investigate the control depth selection with an automatic algorithm based on proximal gradient descent. The performances of the automatic algorithm are demonstrated on 7-node and 10-node Max-Cut problems, which show that the control depth can be significantly reduced during the iteration while achieving an sufficient level of optimization accuracy. With theoretical convergence guarantee, the proposed algorithm can be used as an efficient tool for choosing the appropriate control depth as a replacement of random search or empirical rules. Moreover, the reduction of control depth will induce a significant reduction in the number of quantum gates in circuit, which improves the applicability of QAOA on Noisy Intermediate-scale Quantum (NISQ) devices.

quant-ph↗

Robust optimization for quantum reinforcement learning control using partial observations

The current quantum reinforcement learning control models often assume that the quantum states are known a priori for control optimization. However, full observation of quantum state is experimentally infeasible due to the exponential scaling of the number of required quantum measurements on the number of qubits. In this paper, we investigate a robust reinforcement learning method using partial observations to overcome this difficulty. This control scheme is compatible with near-term quantum devices, where the noise is prevalent and predetermining the dynamics of quantum state is practically impossible. We show that this simplified control scheme can achieve similar or even better performance when compared to the conventional methods relying on full observation. We demonstrate the effectiveness of this scheme on examples of quantum state control and quantum approximate optimization algorithm. It has been shown that high-fidelity state control can be achieved even if the noise amplitude is at the same level as the control amplitude. Besides, an acceptable level of optimization accuracy can be achieved for QAOA with noisy control Hamiltonian. This robust control optimization model can be trained to compensate the uncertainties in practical quantum computing.

quant-ph↗

Reheating constraints on modified single-field Natural Inflation models

In this paper, we discuss three modified single-field natural inflation models in detail, including Special generalized Natural Inflation model(SNI), Extended Natural Inflation model(ENI) and Natural Inflation inspired model(NII). We derive the analytical expression of the tensor-to-scalar ratio $r$ and the spectral index $n_s$ for those models. Then the reheating temperature $T_{re}$ and reheating duration $N_{re}$ are analytically derived. Moreover, considering the CMB constraints, the feasible space of the SNI model in $(n_s, r)$ plane is almost covered by that of the NII, which means the NII is more general than the SNI. In addition, there is no overlapping space between the ENI and the other two models in $(n_s, r)$ plane, which indicates that the ENI and the other two models exclude each other, and more accurate experiments can verify them. Furthermore, the reheating brings tighter constraints to the inflation models, but they still work for a different reheating universe. Considering the constraints of $n_s$, $r$, $N_k$ and choosing $T_{re}$ near the electroweak energy scale, one can find that the decay constants of the three models have no overlapping area and the effective equations of state $ω_{re}$ should be within $\frac{1}{4}\lesssim ω_{re} \lesssim \frac{4}{5}$ for the three models.

hep-ph↗

Cosmological-model-independent tests of cosmic distance duality relation with Type Ia supernovae and radio quasars

In this paper, we investigate the possible deviations of the cosmic distance duality relation (CDDR) using the combination of the largest SNe Ia (Pantheon) and compact radio quasar (QSO) samples through two model-independent approaches. The deviation of CDDR is written as $D_L(z)/D_A(z)(1+z)^{-2}=η(z)$ and $η(z)=e^{τ(z)/2}$, with the parameterizations of $F_1$ ($τ(z) = 2ε_1 z$) and $F_2$ ($τ(z) = (1+z)^{2ε_2}-1$). Furthermore, in order to compare the two resulting distances, two cosmological-model-independent methods, i.e., the nearby SNe Ia method and the GP method are employed to match the two distinct data at the same redshift. Our findings indicate that, compared with the results obtained in the literature, there is an improvement in precision when the latest SNe Ia and QSO samples are used. Specially, in the framework of nearby SNe Ia method, the CDDR would be constrained at the precision of $Δε_{1} = 0.013$ in Model $F_1$ and $Δε_{2}=0.018$ in Model $F_2$. Regarding the GP method, one observes that a larger data size would produce more stringent constraints on the CDDR parameters. Therefore, accompanied by further developments in cosmological observations and the analysis methods, our analysis provides an insight into the evidence for unaccounted opacity sources at an earlier stage of the universe, or at the very least the new physics involved.

astro-ph.CO↗

Multiple Domain Cyberspace Attack and Defense Game Based on Reward Randomization Reinforcement Learning

The existing network attack and defense method can be regarded as game, but most of the game only involves network domain, not multiple domain cyberspace. To address this challenge, this paper proposed a multiple domain cyberspace attack and defense game model based on reinforcement learning. We define the multiple domain cyberspace include physical domain, network domain and digital domain. By establishing two agents, representing the attacker and the defender respectively, defender will select the multiple domain actions in the multiple domain cyberspace to obtain defender's optimal reward by reinforcement learning. In order to improve the defense ability of defender, a game model based on reward randomization reinforcement learning is proposed. When the defender takes the multiple domain defense action, the reward is randomly given and subject to linear distribution, so as to find the better defense policy and improve defense success rate. The experimental results show that the game model can effectively simulate the attack and defense state of multiple domain cyberspace, and the proposed method has a higher defense success rate than DDPG and DQN.

cs.AI↗

Robust photon-efficient imaging using a pixel-wise residual shrinkage network

Single-photon light detection and ranging (LiDAR) has been widely applied to 3D imaging in challenging scenarios. However, limited signal photon counts and high noises in the collected data have posed great challenges for predicting the depth image precisely. In this paper, we propose a pixel-wise residual shrinkage network for photon-efficient imaging from high-noise data, which adaptively generates the optimal thresholds for each pixel and denoises the intermediate features by soft thresholding. Besides, redefining the optimization target as pixel-wise classification provides a sharp advantage in producing confident and accurate depth estimation when compared with existing research. Comprehensive experiments conducted on both simulated and real-world datasets demonstrate that the proposed model outperforms the state-of-the-arts and maintains robust imaging performance under different signal-to-noise ratios including the extreme case of 1:100.

eess.IV↗

KGRGRL: A User's Permission Reasoning Method Based on Knowledge Graph Reward Guidance Reinforcement Learning

In general, multiple domain cyberspace security assessments can be implemented by reasoning user's permissions. However, while existing methods include some information from the physical and social domains, they do not provide a comprehensive representation of cyberspace. Existing reasoning methods are also based on expert-given rules, resulting in inefficiency and a low degree of intelligence. To address this challenge, we create a Knowledge Graph (KG) of multiple domain cyberspace in order to provide a standard semantic description of the multiple domain cyberspace. Following that, we proposed a user's permissions reasoning method based on reinforcement learning. All permissions in cyberspace are represented as nodes, and an agent is trained to find all permissions that user can have according to user's initial permissions and cyberspace KG. We set 10 reward setting rules based on the features of cyberspace KG in the reinforcement learning of reward information setting, so that the agent can better locate user's all permissions and avoid blindly finding user's permissions. The results of the experiments showed that the proposed method can successfully reason about user's permissions and increase the intelligence level of the user's permissions reasoning method. At the same time, the F1 value of the proposed method is 6% greater than that of the Translating Embedding (TransE) method.

cs.AI↗

Quantum-Inspired Solvers on Mixed-Integer Linear Programming Problem

Mixed-integer linear programming (MILP) plays a crucial role in artificial intelligence, biochemistry, finance, cryptography, etc. Notwithstanding popular for decades, the researches of MILP solvers are still limited by the resource consumption caused by complexity and failure of Moore's Law. Quantum-inspired Ising machines, as a new computing paradigm, can be used to solve integer programming problems by reducing them into Ising models. Therefore, it is necessary to understand the technical evolution of quantum inspired solvers to break the bottleneck. In this paper, the concept and traditional algorithms for MILP are introduced. Then, focused on Ising model, the principle and implementations of annealers and coherent Ising machines are summarized. Finally, the paper discusses the challenges and opportunities of miniaturized solvers in the future.

quant-ph↗

Semantically Proportional Patchmix for Few-Shot Learning

Few-shot learning aims to classify unseen classes with only a limited number of labeled data. Recent works have demonstrated that training models with a simple transfer learning strategy can achieve competitive results in few-shot classification. Although excelling at distinguishing training data, these models are not well generalized to unseen data, probably due to insufficient feature representations on evaluation. To tackle this issue, we propose Semantically Proportional Patchmix (SePPMix), in which patches are cut and pasted among training images and the ground truth labels are mixed proportionally to the semantic information of the patches. In this way, we can improve the generalization ability of the model by regional dropout effect without introducing severe label noise. To learn more robust representations of data, we further take rotate transformation on the mixed images and predict rotations as a rule-based regularizer. Extensive experiments on prevalent few-shot benchmarks have shown the effectiveness of our proposed method.

cs.CV↗

High precision measurement of cosmic curvature: from gravitational waves and cosmic chronometer

Although the spatial curvature has been measured with very high precision, it still suffers from the well known cosmic curvature tension. In this paper, we propose an improved method to determine the cosmic curvature, by using the simulated data of binary neutron star mergers observed by the second generation space-based DECi-hertz Interferometer Gravitational-wave Observatory (DECIGO). By applying the Hubble parameter observations of cosmic chronometers to the DECIGO standard sirens, we explore different possibilities of making measurements of the cosmic curvature referring to a distant past: one is to reconstruct the Hubble parameters through the Gaussian process without the influence of hypothetical models, and the other is deriving constraints on $Ω_K$ in the framework of non-flat $Λ$ cold dark matter model. It is shown that in the improved method DECIGO could provide a reliable and stringent constraint on the cosmic curvature ($Ω_{K} = -0.007\pm0.016$), while we could only expect the zero cosmic curvature to be established at the precision of $ΔΩ_K=0.12$ in the second model-dependent method. Therefore, our results indicate that in the framework of methodology proposed in this paper, the increasing number of well-measured standard sirens in DECIGO could significantly reduce the bias of estimations for cosmic curvature. Such constraint is also comparable to the precision of Planck 2018 results with the newest cosmic microwave background (CMB) observations ($ΔΩ_{K} \approx 0.018$), based on the concordance $Λ$CDM model.

astro-ph.CO↗