Searcharxiv⌕ Search

arXiv subjects

Li Tang

Publications and source records attributed to Li Tang.

At least 37 records · Page 2Linked to original sources

Cross-modal Prompts: Adapting Large Pre-trained Models for Audio-Visual Downstream Tasks

In recent years, the deployment of large-scale pre-trained models in audio-visual downstream tasks has yielded remarkable outcomes. However, these models, primarily trained on single-modality unconstrained datasets, still encounter challenges in feature extraction for multi-modal tasks, leading to suboptimal performance. This limitation arises due to the introduction of irrelevant modality-specific information during encoding, which adversely affects the performance of downstream tasks. To address this challenge, this paper proposes a novel Dual-Guided Spatial-Channel-Temporal (DG-SCT) attention mechanism. This mechanism leverages audio and visual modalities as soft prompts to dynamically adjust the parameters of pre-trained models based on the current multi-modal input features. Specifically, the DG-SCT module incorporates trainable cross-modal interaction layers into pre-trained audio-visual encoders, allowing adaptive extraction of crucial information from the current modality across spatial, channel, and temporal dimensions, while preserving the frozen parameters of large-scale pre-trained models. Experimental evaluations demonstrate that our proposed model achieves state-of-the-art results across multiple downstream tasks, including AVE, AVVP, AVS, and AVQA. Furthermore, our model exhibits promising performance in challenging few-shot and zero-shot scenarios. The source code and pre-trained models are available at https://github.com/haoyi-duan/DG-SCT.

cs.LG↗

Vanishing of the anomalous Hall effect and enhanced carrier mobility in the spin-gapless ferromagnetic Mn2CoGa1-xAlx alloys

Spin gapless semiconductor (SGS) has attracted long attention since its theoretical prediction, while concrete experimental hints are still lack in the relevant Heusler alloys. Here in this work, by preparing the series alloys of Mn2CoGa1-xAlx (x=0, 0.25, 0.5, 0.75 and 1), we identified the vanishing of anomalous Hall effect in the ferromagnetic Mn2CoGa (or x=0.25) alloy in a wide temperature interval, accompanying with growing contribution from the ordinary Hall effect. As a result, comparatively low carrier density (1020 cm-3) and high carrier mobility (150 cm2/Vs) are obtained in Mn2CoGa (or x=0.25) alloy in the temperature range of 10-200K. These also lead to a large dip in the related magnetoresistance at low fields. While in high Al content, despite the magnetization behavior is not altered significantly, the Hall resistivity is instead dominated by the anomalous one, just analogous to that widely reported in Mn2CoAl. The distinct electrical transport behavior of x=0 and x=0.75 (or 1) is presently understood by their possible different scattering mechanism of the anomalous Hall effect due to the differences in atomic order and conductivity. Our work can expand the existing understanding of the SGS properties and offer a better SGS candidate with higher carrier mobility that can facilitate the application in the spin-injected related devices.

cond-mat.mtrl-sci↗

Consistency of Pantheon+ supernovae with a large-scale isotropic universe

We investigate the possible anisotropy of the universe using the most up-to-date type Ia supernovae, i.e. the Pantheon+ compilation. We fit the full Pantheon+ data with the dipole-modulated $Λ$CDM model, and find that it is well consistent with a null dipole. We further divide the full sample into several subsamples with different high-redshift cutoff $z_c$. It is shown that the dipole appears at $2σ$ confidence level only if $z_c\leq 0.1$, and in this redshift region the dipole is very stable, almost independent of the specific value of $z_c$. For $z_c=0.1$, the dipole amplitude is $D=1.0_{-0.4}^{+0.4}\times 10^{-3}$, pointing towards $(l,b)=(334.5_{\ -21.6^{\circ}}^{\circ +25.7^{\circ}},16.0_{\ -16.8^{\circ}}^{\circ +27.1^{\circ}})$, which is about $65^{\circ}$ away from the CMB dipole. This implies that the full Pantheon+ is consistent with a large-scale isotropic universe, but the low-redshift anisotropy couldn't be purely explained by the peculiar motion of the local universe.

astro-ph.CO↗

Enhanced multilayer perceptron with feature selection and grid search for travel mode choice prediction

Accurate and reliable prediction of individual travel mode choices is crucial for developing multi-mode urban transportation systems, conducting transportation planning and formulating traffic demand management strategies. Traditional discrete choice models have dominated the modelling methods for decades yet suffer from strict model assumptions and low prediction accuracy. In recent years, machine learning (ML) models, such as neural networks and boosting models, are widely used by researchers for travel mode choice prediction and have yielded promising results. However, despite the superior prediction performance, a large body of ML methods, especially the branch of neural network models, is also limited by overfitting and tedious model structure determination process. To bridge this gap, this study proposes an enhanced multilayer perceptron (MLP; a neural network) with two hidden layers for travel mode choice prediction; this MLP is enhanced by XGBoost (a boosting method) for feature selection and a grid search method for optimal hidden neurone determination of each hidden layer. The proposed method was trained and tested on a real resident travel diary dataset collected in Chengdu, China.

econ.EM↗

Connecting Multi-modal Contrastive Representations

Multi-modal Contrastive Representation learning aims to encode different modalities into a semantically aligned shared space. This paradigm shows remarkable generalization ability on numerous downstream tasks across various modalities. However, the reliance on massive high-quality data pairs limits its further development on more modalities. This paper proposes a novel training-efficient method for learning MCR without paired data called Connecting Multi-modal Contrastive Representations (C-MCR). Specifically, given two existing MCRs pre-trained on (A, B) and (B, C) modality pairs, we project them to a new space and use the data from the overlapping modality B to aligning the two MCRs in the new space. Meanwhile, since the modality pairs (A, B) and (B, C) are already aligned within each MCR, the connection learned by overlapping modality can also be transferred to non-overlapping modality pair (A, C). To unleash the potential of C-MCR, we further introduce a semantic-enhanced inter- and intra-MCR connection method. We first enhance the semantic consistency and completion of embeddings across different modalities for more robust alignment. Then we utilize the inter-MCR alignment to establish the connection, and employ the intra-MCR alignment to better maintain the connection for inputs from non-overlapping modalities. To demonstrate the effectiveness of C-MCR, we connect CLIP and CLAP via texts to derive audio-visual representations, and integrate CLIP and ULIP via images for 3D-language representations. Remarkably, without using any paired data, C-MCR for audio-visual achieves state-of-the-art performance on audio-image retrieval, audio-visual source localization, and counterfactual audio-image recognition tasks. Furthermore, C-MCR for 3D-language also attains advanced zero-shot 3D point cloud classification accuracy on ModelNet40.

cs.LG↗

Constraining the spatial curvature of the local Universe with deep learning

We use the distance sum rule (DSR) method to constrain the spatial curvature of the Universe with a large sample of 161 strong gravitational lensing (SGL) systems, whose distances are calibrated from the Pantheon compilation of type Ia supernovae (SNe Ia) using deep learning. To investigate the possible influence of mass model of the lens galaxy on constraining the curvature parameter $Ω_k$, we consider three different lens models. Results show that a flat Universe is supported in the singular isothermal sphere (SIS) model with the parameter $Ω_k=0.049^{+0.147}_{-0.125}$. While in the power-law (PL) model, a closed Universe is preferred at $\sim 3σ$ confidence level, with the parameter $Ω_k=-0.245^{+0.075}_{-0.071}$. In extended power-law (EPL) model, the 95$\%$ confidence level upper limit of $Ω_k$ is $<0.011$. As for the parameters of the lens models, constrains on the three models indicate that the mass profile of the lens galaxy could not be simply described by the standard SIS model.

astro-ph.CO↗

Gloss Attention for Gloss-free Sign Language Translation

Most sign language translation (SLT) methods to date require the use of gloss annotations to provide additional supervision information, however, the acquisition of gloss is not easy. To solve this problem, we first perform an analysis of existing models to confirm how gloss annotations make SLT easier. We find that it can provide two aspects of information for the model, 1) it can help the model implicitly learn the location of semantic boundaries in continuous sign language videos, 2) it can help the model understand the sign language video globally. We then propose \emph{gloss attention}, which enables the model to keep its attention within video segments that have the same semantics locally, just as gloss helps existing models do. Furthermore, we transfer the knowledge of sentence-to-sentence similarity from the natural language model to our gloss attention SLT network (GASLT) to help it understand sign language videos at the sentence level. Experimental results on multiple large-scale sign language datasets show that our proposed GASLT model significantly outperforms existing methods. Our code is provided in \url{https://github.com/YinAoXiong/GASLT}.

cs.CV↗

Inferring redshift and energy distributions of fast radio bursts from the first CHIME/FRB catalog

We reconstruct the extragalactic dispersion measure \ -- redshift relation (${\rm DM_E}-z$ relation) from well-localized fast radio bursts (FRBs) using Bayesian inference method. Then the ${\rm DM_E}-z$ relation is used to infer the redshift and energy of the first CHIME/FRB catalog. We find that the distributions of extragalactic dispersion measure and inferred redshift of the non-repeating CHIME/FRBs follow cut-off power law, but with a significant excess at the low-redshift range. We apply a set of criteria to exclude events which are susceptible to selection effect, but find that the excess at low redshift still exists in the remaining FRBs (which we call Gold sample). The cumulative distributions of fluence and energy for both the full sample and the Gold sample do not follow the simple power law, but they can be well fitted by the bent power law. The underlying physical implications remain to be further investigated.

astro-ph.HE↗

Revised constraints on the photon mass from well-localized fast radio bursts

We constrain the photon mass from well-localized fast radio bursts (FRBs) using Bayes inference method. The probability distributions of dispersion measures (DM) of host galaxy and intergalactic medium are properly taken into account. The photon mass is tightly constrained from 17 well-localized FRBs in the redshift range $0<z<0.66$. Assuming that there is no redshift evolution of host DM, the $1σ$ and $2σ$ upper limits of photon mass are constrained to be $m_γ<4.8\times 10^{-51}$ kg and $m_γ<7.1\times 10^{-51}$ kg, respectively. Monte Carlo simulations show that, even enlarging the FRB sample to 200 and extending the redshift range to $0<z<3$ couldn't significantly improve the constraining ability on photon mass. This is because of the large uncertainty on the DM of intergalactic medium.

gr-qc↗

DeepRING: Learning Roto-translation Invariant Representation for LiDAR based Place Recognition

LiDAR based place recognition is popular for loop closure detection and re-localization. In recent years, deep learning brings improvements to place recognition by learnable feature extraction. However, these methods degenerate when the robot re-visits previous places with large perspective difference. To address the challenge, we propose DeepRING to learn the roto-translation invariant representation from LiDAR scan, so that robot visits the same place with different perspective can have similar representations. There are two keys in DeepRING: the feature is extracted from sinogram, and the feature is aggregated by magnitude spectrum. The two steps keeps the final representation with both discrimination and roto-translation invariance. Moreover, we state the place recognition as a one-shot learning problem with each place being a class, leveraging relation learning to build representation similarity. Substantial experiments are carried out on public datasets, validating the effectiveness of each proposed component, and showing that DeepRING outperforms the comparative methods, especially in dataset level generalization.

cs.CV↗

Deep learning method in testing the cosmic distance duality relation

The cosmic distance duality relation (DDR) is constrained from the combination of type-Ia supernovae (SNe Ia) and strong gravitational lensing (SGL) systems using deep learning method. To make use of the full SGL data, we reconstruct the luminosity distance from SNe Ia up to the highest redshift of SGL using deep learning, then it is compared with the angular diameter distance obtained from SGL. Considering the influence of lens mass profile, we constrain the possible violation of DDR in three lens mass models. Results show that in the SIS model and EPL model, DDR is violated at high confidence level, with the violation parameter $η_0=-0.193^{+0.021}_{-0.019}$ and $η_0=-0.247^{+0.014}_{-0.013}$, respectively. In the PL model, however, DDR is verified within 1$σ$ confidence level, with the violation parameter $η_0=-0.014^{+0.053}_{-0.045}$. Our results demonstrate that the constraints on DDR strongly depend on the lens mass models. Given a specific lens mass model, DDR can be constrained at a precision of $\textit{O}(10^{-2})$ using deep learning.

astro-ph.CO↗

Time Interval-enhanced Graph Neural Network for Shared-account Cross-domain Sequential Recommendation

Shared-account Cross-domain Sequential Recommendation (SCSR) task aims to recommend the next item via leveraging the mixed user behaviors in multiple domains. It is gaining immense research attention as more and more users tend to sign up on different platforms and share accounts with others to access domain-specific services. Existing works on SCSR mainly rely on mining sequential patterns via Recurrent Neural Network (RNN)-based models, which suffer from the following limitations: 1) RNN-based methods overwhelmingly target discovering sequential dependencies in single-user behaviors. They are not expressive enough to capture the relationships among multiple entities in SCSR. 2) All existing methods bridge two domains via knowledge transfer in the latent space, and ignore the explicit cross-domain graph structure. 3) None existing studies consider the time interval information among items, which is essential in the sequential recommendation for characterizing different items and learning discriminative representations for them. In this work, we propose a new graph-based solution, namely TiDA-GCN, to address the above challenges. Specifically, we first link users and items in each domain as a graph. Then, we devise a domain-aware graph convolution network to learn userspecific node representations. To fully account for users' domainspecific preferences on items, two effective attention mechanisms are further developed to selectively guide the message passing process. Moreover, to further enhance item- and account-level representation learning, we incorporate the time interval into the message passing, and design an account-aware self-attention module for learning items' interactive characteristics. Experiments demonstrate the superiority of our proposed method from various aspects.

cs.IR↗

Search for the correlations between host properties and ${\rm DM_{host}}$ of fast radio bursts: constraints on the baryon mass fraction in IGM

The application of fast radio bursts (FRBs) as probes to investigate astrophysics and cosmology requires the proper modelling of the dispersion measures of Milky Way (${\rm DM_{MW}}$) and host galaxy (${\rm DM_{host}}$). ${\rm DM_{MW}}$ can be estimated using the Milky Way electron models, such as NE2001 model and YMW16 model. However, ${\rm DM_{host}}$ is hard to model due to limited information on the local environment of FRBs. In this paper, using 17 well-localized FRBs, we search for the possible correlations between ${\rm DM_{host}}$ and the properties of host galaxies, such as the redshift, the stellar mass, the star-formation rate, the age of galaxy, the offset of FRB site from galactic center, and the half-light radius. We find no strong correlation between ${\rm DM_{host}}$ and any of the host property. Assuming that ${\rm DM_{host}}$ is a constant for all host galaxies, we constrain the fraction of baryon mass in the intergalactic medium today to be $f_{\rm IGM,0}=0.78_{-0.19}^{+0.15}$. If we model ${\rm DM_{host}}$ as a log-normal distribution, however, we obtain a larger value, $f_{\rm IGM,0}=0.83_{-0.17}^{+0.12}$. Based on the limited number of FRBs, no strong evidence for the redshift evolution of $f_{\rm IGM}$ is found.

astro-ph.CO↗

Reconstructing the Hubble diagram of gamma-ray bursts using deep learning

We calibrate the distance and reconstruct the Hubble diagram of gamma-ray bursts (GRBs) using deep learning. We construct an artificial neural network, which combines the recurrent neural network and Bayesian neural network, and train the network using the Pantheon compilation of type-Ia supernovae. The trained network is used to calibrate the distance of 174 GRBs based on the Combo-relation. We verify that there is no evident redshift evolution of Combo-relation, and obtain the slope and intercept parameters, $γ=0.856^{+0.083}_{-0.078}$ and $\log A=49.661^{+0.199}_{-0.217}$, with an intrinsic scatter $σ_{\rm int}=0.228^{+0.041}_{-0.040}$. Our calibrating method is independent of cosmological model, thus the calibrated GRBs can be directly used to constrain cosmological parameters. It is shown that GRBs alone can tightly constrain the $Λ$CDM model, with $Ω_{\rm M}=0.280^{+0.049}_{-0.057}$. However, the constraint on the $ω$CDM model is relatively looser, with $Ω_{\rm M}=0.345^{+0.059}_{-0.060}$ and $ω<-1.414$. The combination of GRBs and Pantheon can tightly constrain the $ω$CDM model, with $Ω_{\rm M}=0.336^{+0.055}_{-0.050}$ and $ω=-1.141^{+0.156}_{-0.135}$.

gr-qc↗

Image of the Schwarzschild black hole pierced by a cosmic string with a thin accretion disk

We study the optical appearance of a thin accretion disk around a Schwarzschild black hole pierced by a cosmic string with a semi-analytic method of Luminet [11]. Direct and secondary images with different parameters observed by a distant observer is plotted. The cosmic string parameter s can modify the shape and size of the thin disk image. We calculate and plot the distribution of both redshift and observed flux as seen by distant observers at different inclination angles. Those distributions are dependent on the inclination angel of the observer and cosmic parameter s.

gr-qc↗

Optimal Hardy inequalities associated with multipolar Schrödinger operators

We proved some optimal Hardy inequalities in RNwhich is closely related to multipolar Schrödinger operators with mean-value type potentials, these sharp inequalities imply some multipolar type Heisenberg inequalities. We also obtained someimproved multipolar Hardy inequalities on bounded domains, moreover, we got the range of the best Hardy constant for a specific Hardy inequality.

math.AP↗

Numerically stable coded matrix computations via circulant and rotation matrix embeddings

Polynomial based methods have recently been used in several works for mitigating the effect of stragglers (slow or failed nodes) in distributed matrix computations. For a system with $n$ worker nodes where $s$ can be stragglers, these approaches allow for an optimal recovery threshold, whereby the intended result can be decoded as long as any $(n-s)$ worker nodes complete their tasks. However, they suffer from serious numerical issues owing to the condition number of the corresponding real Vandermonde-structured recovery matrices; this condition number grows exponentially in $n$. We present a novel approach that leverages the properties of circulant permutation matrices and rotation matrices for coded matrix computation. In addition to having an optimal recovery threshold, we demonstrate an upper bound on the worst-case condition number of our recovery matrices which grows as $\approx O(n^{s+5.5})$; in the practical scenario where $s$ is a constant, this grows polynomially in $n$. Our schemes leverage the well-behaved conditioning of complex Vandermonde matrices with parameters on the complex unit circle, while still working with computation over the reals. Exhaustive experimental results demonstrate that our proposed method has condition numbers that are orders of magnitude lower than prior work.

cs.IT↗

Machine learning forecasts of the cosmic distance duality relation with strongly lensed gravitational wave events

We use simulated strongly lensed gravitational wave events from the Einstein Telescope to demonstrate how the luminosity and angular diameter distances, $d_L(z)$ and $d_A(z)$ respectively, can be combined to test in a model independent manner for deviations from the cosmic distance duality relation and the standard cosmological model. In particular, we use two machine learning approaches, the Genetic Algorithms and Gaussian Processes, to reconstruct the mock data and we show that both approaches are capable of correctly recovering the underlying fiducial model and can provide percent-level constraints at intermediate redshifts when applied to future Einstein Telescope data.

astro-ph.CO↗