SearcharxivSearch

arXiv subjects

Ye Yan

Publications and source records attributed to Ye Yan.

At least 19 recordsLinked to original sources

Exploring possible $^3_{\Lambda_c}\text{H}$ bound states through $p\Lambda_c$ femtoscopic correlations

The femtoscopic correlation technique in relativistic heavy-ion collisions provides a unique opportunity to investigate hadron-hadron interactions and possible exotic states. In this work, we study the $p\Lambda_c$ correlation function and its sensitivity to the low-energy $N\Lambda_c$ interaction related to possible $^3_{\Lambda_c}\mathrm{H}$ bound states. Based on the quark delocalization color screening model, three interaction scenarios with different strengths are constructed, and the corresponding spin-averaged $p\Lambda_c$ correlation functions are calculated within the Koonin--Pratt formalism. The results demonstrate that the correlation function is sensitive to the $p\Lambda_c$ interaction strength, with coupled-channel effects and $S$-$D$ wave mixing producing additional enhancements in the correlation signal. These findings suggest that future $p\Lambda_c$ femtoscopic measurements at relativistic heavy-ion collision experiments can provide valuable constraints on the interaction between charmed baryons and nucleons and offer guidance for exploring possible heavy-flavor hypernuclei.

hep-ph

A coupled-channel quark model study of possible $\Xi_{cc}^{(*)} K^{(*)}$ molecular states

Inspired by the recent experimental discovery of doubly charmed baryons, we investigate the possible $\Xi_{cc}^{(*)}K^{(*)}$ molecular systems within the framework of the quark delocalization color screening model. The energy spectra and scattering processes of the relevant baryon-meson systems are investigated to explore the dynamical properties of the possible molecular states. The spectrum calculations predict three bound states, namely the $I(J^P)=0(1/2^{-})$ $\Xi_{cc}K$, the $I(J^P)=0(3/2^{-})$ $\Xi_{cc}^{*}K$, and the $I(J^P)=0(5/2^{-})$ $\Xi_{cc}^{*}K^{*}$ molecular states. The scattering phase shift analysis further confirms two $\Xi_{cc}K^{*}$ resonance states with $I(J^P)=0(1/2^{-})$ and $0(3/2^{-})$, which originate from quasi-bound states through channel coupling. In particular, the $I(J^P)=0(1/2^{-})$ $\Xi_{cc}K$ bound state is consistent with previous theoretical studies, making it one of the most promising candidates for future experimental searches.

hep-ph

Investigation of fully heavy tetraquark within chiral quark model

In the framework of the Chiral quark model (ChQM), we investigate the fully charmed and fully bottomed tetraquark with $J^{PC}=2^{++}$ including two structures: $Q\bar{Q}-Q\bar{Q}$ and $QQ-\bar{Q}\bar{Q}$. The bound-state calculation shows that there is no bound state in either $cc\bar{c}\bar{c}$ or $bb\bar{b}\bar{b}$ systems. However, by using the real-scaling method, some resonance states are obtained. For the $cc\bar{c}\bar{c}$ system, when the channel-coupling includes only three $S$-wave channels, two resonant states are obtained: one with a mass around $7002$ MeV and decay width near $54$ MeV, and another with a mass around $7227$ MeV and a decay width near $66$ MeV. The former can be regarded as a candidate for the $X(6900)$, and the latter can be considered as a candidate for the $X(7200)$. Upon adding the $\chi_{c0}\chi_{c2}$, $\chi_{c1}\chi_{c1}$, $\chi_{c1}\chi_{c2}$, $\chi_{c2}\chi_{c2}$ channels, both resonant states still remain. For the $bb\bar{b}\bar{b}$ system, only one resonant state is obtained, regardless of whether the four channels composition of the excited mesons are included or excluded. The mass and width of this resonant state are around $19743$ MeV and $67$ MeV, respectively. We suggest that future experiments search for the possible resonance state in the invariant mass spectrum of $\Upsilon \Upsilon$ or $\Upsilon \Upsilon(2S)$.

hep-ph

DBMIF: a deep balanced multimodal iterative fusion framework for air- and bone-conduction speech enhancement

The performance of conventional speech enhancement systems degrades sharply in extremely low signal-to-noise ratio (SNR) environments where air-conduction (AC) microphones are overwhelmed by ambient noise. Although bone-conduction (BC) sensors offer complementary, noise-tolerant information, existing fusion approaches struggle to maintain consistent performance across a wide range of SNR conditions. To address this limitation, we propose the Deep Balanced Multimodal Iterative Fusion Framework (DBMIF), a three-branch architecture designed to reconstruct high-fidelity speech through rigorous cross-modal interaction. Specifically, grounded in a multi-scale interactive encoder-decoder backbone, the framework orchestrates an iterative attention module and a cross-branch gated module to facilitate adaptive weighting and bidirectional exchange. To complement this dynamic interaction, a balanced-interaction bottleneck is further integrated to learn a compact, stable fused representation. Extensive experiments demonstrate that DBMIF achieves competitive performance compared with recent unimodal and multimodal baselines in both speech quality and intelligibility across diverse noise types. In downstream ASR tasks, the proposed method reduces the character error rate by at least 2.5 percent compared to competing approaches. These results confirm that DBMIF effectively harnesses the robustness of BC speech while preserving the naturalness of AC speech, ensuring reliability in real-world scenarios. The source code is publicly available at github.com/wyl516w/dbmif.

eess.AS

Investigating $\Omega \phi$ Interaction and Correlation Functions

In this work, we investigate the interaction between the $\Omega$ baryon and the $s\bar{s}$ meson within the framework of the quark delocalization color screening model. The spectra calculations show that no bound state is formed in any of the considered channels, while the scattering indicates that the $\Omega\phi$ interaction with $J^{P}=1/2^{-}$ is weakly attractive. As for the $\Omega\phi$ interactions with $J^{P}=3/2^{-}$ and $5/2^{-}$, as well as the $\Omega\eta^{\prime}$ interaction with $J^{P}=3/2^{-}$, they are all repulsive. After an investigation on the femtoscopic correlation functions, we find that, due to the spin-averaging effect, the overall $\Omega\phi$ correlation function exhibits a weak dependence on the source size, which provides a crucial significance of our model for future experimental examinations in relativistic heavy-ion collisions.

hep-ph

Scene-aware SAR ship detection guided by unsupervised sea-land segmentation

DL based Synthetic Aperture Radar (SAR) ship detection has tremendous advantages in numerous areas. However, it still faces some problems, such as the lack of prior knowledge, which seriously affects detection accuracy. In order to solve this problem, we propose a scene-aware SAR ship detection method based on unsupervised sea-land segmentation. This method follows a classical two-stage framework and is enhanced by two models: the unsupervised land and sea segmentation module (ULSM) and the land attention suppression module (LASM). ULSM and LASM can adaptively guide the network to reduce attention on land according to the type of scenes (inshore scene and offshore scene) and add prior knowledge (sea land segmentation information) to the network, thereby reducing the network's attention to land directly and enhancing offshore detection performance relatively. This increases the accuracy of ship detection and enhances the interpretability of the model. Specifically, in consideration of the lack of land sea segmentation labels in existing deep learning-based SAR ship detection datasets, ULSM uses an unsupervised approach to classify the input data scene into inshore and offshore types and performs sea-land segmentation for inshore scenes. LASM uses the sea-land segmentation information as prior knowledge to reduce the network's attention to land. We conducted our experiments using the publicly available SSDD dataset, which demonstrated the effectiveness of our network.

cs.CV

MPFNet: A Multi-Prior Fusion Network with a Progressive Training Strategy for Micro-Expression Recognition

Micro-expression recognition (MER), a critical subfield of affective computing, presents greater challenges than macro-expression recognition due to its brief duration and low intensity. While incorporating prior knowledge has been shown to enhance MER performance, existing methods predominantly rely on simplistic, singular sources of prior knowledge, failing to fully exploit multi-source information. This paper introduces the Multi-Prior Fusion Network (MPFNet), leveraging a progressive training strategy to optimize MER tasks. We propose two complementary encoders: the Generic Feature Encoder (GFE) and the Advanced Feature Encoder (AFE), both based on Inflated 3D ConvNets (I3D) with Coordinate Attention (CA) mechanisms, to improve the model's ability to capture spatiotemporal and channel-specific features. Inspired by developmental psychology, we present two variants of MPFNet--MPFNet-P and MPFNet-C--corresponding to two fundamental modes of infant cognitive development: parallel and hierarchical processing. These variants enable the evaluation of different strategies for integrating prior knowledge. Extensive experiments demonstrate that MPFNet significantly improves MER accuracy while maintaining balanced performance across categories, achieving accuracies of 0.811, 0.924, and 0.857 on the SMIC, CASME II, and SAMM datasets, respectively. To the best of our knowledge, our approach achieves state-of-the-art performance on the SMIC and SAMM datasets.

cs.CV

MMME: A Spontaneous Multi-Modal Micro-Expression Dataset Enabling Visual-Physiological Fusion

Micro-expressions (MEs) are subtle, fleeting nonverbal cues that reveal an individual's genuine emotional state. Their analysis has attracted considerable interest due to its promising applications in fields such as healthcare, criminal investigation, and human-computer interaction. However, existing ME research is limited to single visual modality, overlooking the rich emotional information conveyed by other physiological modalities, resulting in ME recognition and spotting performance far below practical application needs. Therefore, exploring the cross-modal association mechanism between ME visual features and physiological signals (PS), and developing a multimodal fusion framework, represents a pivotal step toward advancing ME analysis. This study introduces a novel ME dataset, MMME, which, for the first time, enables synchronized collection of facial action signals (MEs), central nervous system signals (EEG), and peripheral PS (PPG, RSP, SKT, EDA, and ECG). By overcoming the constraints of existing ME corpora, MMME comprises 634 MEs, 2,841 macro-expressions (MaEs), and 2,890 trials of synchronized multimodal PS, establishing a robust foundation for investigating ME neural mechanisms and conducting multimodal fusion-based analyses. Extensive experiments validate the dataset's reliability and provide benchmarks for ME analysis, demonstrating that integrating MEs with PS significantly enhances recognition and spotting performance. To the best of our knowledge, MMME is the most comprehensive ME dataset to date in terms of modality diversity. It provides critical data support for exploring the neural mechanisms of MEs and uncovering the visual-physiological synergistic effects, driving a paradigm shift in ME research from single-modality visual analysis to multimodal fusion. The dataset will be publicly available upon acceptance of this paper.

cs.CV

Prediction of $p\bar{\Omega}$ states and femtoscopic study

Inspired by recent researches on the $p \Omega$ and $p \bar{\Lambda}$ systems, we investigate the $p \bar{\Omega}$ systems within the framework of a quark model. Our results show that the attraction between a nucleon and $\bar{\Omega}$ is slightly stronger than that between a nucleon and $\Omega$, suggesting that the $p \bar{\Omega}$ system is more likely to form bound states. The dynamic calculations indicate that the $p \bar{\Omega}$ systems with both $J^{P}=1^{-}$ and $2^{-}$ can form bound states, with binding energies deeper than those of the $p \Omega$ systems with $J^{P}=2^{+}$. The scattering phase shift and scattering parameter calculations also support the existence of $p \bar{\Omega}$ states. Additionally, we discuss the behavior of the femtoscopic correlation function for the $p \bar{\Omega}$ pairs for the first time. Considering the significant progress in experimental measurements of the correlation function of the $p \Omega$ system, the further study of the $p \bar{\Omega}$ systems using femtoscopic techniques will be a very valuable work.

hep-ph

PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation

Vision-and-language navigation (VLN) tasks require agents to navigate three-dimensional environments guided by natural language instructions, offering substantial potential for diverse applications. However, the scarcity of training data impedes progress in this field. This paper introduces PanoGen++, a novel framework that addresses this limitation by generating varied and pertinent panoramic environments for VLN tasks. PanoGen++ incorporates pre-trained diffusion models with domain-specific fine-tuning, employing parameter-efficient techniques such as low-rank adaptation to minimize computational costs. We investigate two settings for environment generation: masked image inpainting and recursive image outpainting. The former maximizes novel environment creation by inpainting masked regions based on textual descriptions, while the latter facilitates agents' learning of spatial relationships within panoramas. Empirical evaluations on room-to-room (R2R), room-for-room (R4R), and cooperative vision-and-dialog navigation (CVDN) datasets reveal significant performance enhancements: a 2.44% increase in success rate on the R2R test leaderboard, a 0.63% improvement on the R4R validation unseen set, and a 0.75-meter enhancement in goal progress on the CVDN validation unseen set. PanoGen++ augments the diversity and relevance of training environments, resulting in improved generalization and efficacy in VLN tasks.

cs.CV

DECAN: A Denoising Encoder via Contrastive Alignment Network for Dry Electrode EEG Emotion Recognition

EEG signal is important for brain-computer interfaces (BCI). Nevertheless, existing dry and wet electrodes are difficult to balance between high signal-to-noise ratio and portability in EEG recording, which limits the practical use of BCI. In this study, we propose a Denoising Encoder via Contrastive Alignment Network (DECAN) for dry electrode EEG, under the assumption of the EEG representation consistency between wet and dry electrodes during the same task. Specifically, DECAN employs two parameter-sharing deep neural networks to extract task-relevant representations of dry and wet electrode signals, and then integrates a representation-consistent contrastive loss to minimize the distance between representations from the same timestamp and category but different devices. To assess the feasibility of our approach, we construct an emotion dataset consisting of paired dry and wet electrode EEG signals from 16 subjects with 5 emotions, named PaDWEED. Results on PaDWEED show that DECAN achieves an average accuracy increase of 6.94$\%$ comparing to state-of-the art performance in emotion recognition of dry electrodes. Ablation studies demonstrate a decrease in inter-class aliasing along with noteworthy accuracy enhancements in the delta and beta frequency bands. Moreover, an inter-subject feature alignment can obtain an accuracy improvement of 5.99$\%$ and 5.14$\%$ in intra- and inter-dataset scenarios, respectively. Our proposed method may open up new avenues for BCI with dry electrodes. PaDWEED dataset used in this study is freely available at https://huggingface.co/datasets/peiyu999/PaDWEED.

cs.HC

Investigating the $p$-$\Omega$ Interaction and Correlation Functions

Motivated by experimental measurements, we investigate the $p$-$\Omega$ correlation functions and interactions on the basis of a quark model. By solving the inverse scattering problem with channel coupling, we renormalize the coupling to other channels into an effective single-channel $p$-$\Omega$ potentials. The effects of Coulomb interaction and spin-averaging are also discussed. According to our results, the depletion of the $p$-$\Omega$ correlation functions, which is attributed to the $J^P = 2^+$ bound state not observed in the ALICE Collaboration's measurements [Nature \textbf{588}, 232 (2020)], can be explained by the contribution of the attractive $J^P = 1^+$ component in spin-averaging. So far, we have provided a consistent description of the $p$-$\Omega$ system from the perspective of the quark model, including the energy spectrum, scattering phase shifts, and correlation functions. The existence of the $p$-$\Omega$ bound state has been supported by all three aspects. Additionally, a sign of the $p$-$\Omega$ correlation function's subtle sub-unity part can be seen in experimental measurements, which warrants more precise verification in the future.

hep-ph

Investigating $\Xi$ resonances from pentaquark perspective

We have investigated the $qss\bar{q}q$ ($q = u$ or $d$) system to find possible pentaquark explanations for the $\Xi$ resonances. The bound state calculation is carried out within the framework of the quark delocalization color screening model. The scattering processes are also studied to examine the possible resonance states. The current results indicate that the $\Xi(1950)$ can be interpreted as $\Lambda \bar{K}^*$ state with $J^P = 1/2^-$. Three states are identified that match the $\Xi(2250)$, which are $\Sigma^* \bar{K}^*$ state with $J^P = 3/2^-$, $\Sigma^* \bar{K}^*$ state with $J^P =5/2^-$, and $\Xi^* \rho$ state with $J^P =5/2^-$. This may explain the conflicting experimental values for the width of the $\Xi(2250)$. A new $\Xi$ resonance is predicted, whose mass and width are 2066--2079 MeV and 186--189 MeV, respectively. These results contribute to understanding the nature of the $\Xi$ resonances and to the future search for new $\Xi$ resonances. Moreover, it is meaningful to further investigate the $\Xi$ resonances from an unquenched picture on the basis of pentaquark investigation.

hep-ph

Landmark-Guided Cross-Speaker Lip Reading with Mutual Information Regularization

Lip reading, the process of interpreting silent speech from visual lip movements, has gained rising attention for its wide range of realistic applications. Deep learning approaches greatly improve current lip reading systems. However, lip reading in cross-speaker scenarios where the speaker identity changes, poses a challenging problem due to inter-speaker variability. A well-trained lip reading system may perform poorly when handling a brand new speaker. To learn a speaker-robust lip reading model, a key insight is to reduce visual variations across speakers, avoiding the model overfitting to specific speakers. In this work, in view of both input visual clues and latent representations based on a hybrid CTC/attention architecture, we propose to exploit the lip landmark-guided fine-grained visual clues instead of frequently-used mouth-cropped images as input features, diminishing speaker-specific appearance characteristics. Furthermore, a max-min mutual information regularization approach is proposed to capture speaker-insensitive latent representations. Experimental evaluations on public lip reading datasets demonstrate the effectiveness of the proposed approach under the intra-speaker and inter-speaker conditions.

cs.AI

HDA-LVIO: A High-Precision LiDAR-Visual-Inertial Odometry in Urban Environments with Hybrid Data Association

To enhance localization accuracy in urban environments, an innovative LiDAR-Visual-Inertial odometry, named HDA-LVIO, is proposed by employing hybrid data association. The proposed HDA_LVIO system can be divided into two subsystems: the LiDAR-Inertial subsystem (LIS) and the Visual-Inertial subsystem (VIS). In the LIS, the LiDAR pointcloud is utilized to calculate the Iterative Closest Point (ICP) error, serving as the measurement value of Error State Iterated Kalman Filter (ESIKF) to construct the global map. In the VIS, an incremental method is firstly employed to adaptively extract planes from the global map. And the centroids of these planes are projected onto the image to obtain projection points. Then, feature points are extracted from the image and tracked along with projection points using Lucas-Kanade (LK) optical flow. Next, leveraging the vehicle states from previous intervals, sliding window optimization is performed to estimate the depth of feature points. Concurrently, a method based on epipolar geometric constraints is proposed to address tracking failures for feature points, which can improve the accuracy of depth estimation for feature points by ensuring sufficient parallax within the sliding window. Subsequently, the feature points and projection points are hybridly associated to construct reprojection error, serving as the measurement value of ESIKF to estimate vehicle states. Finally, the localization accuracy of the proposed HDA-LVIO is validated using public datasets and data from our equipment. The results demonstrate that the proposed algorithm achieves obviously improvement in localization accuracy compared to various existing algorithms.

cs.RO

Prediction of $P_{cc}$ states in quark model

Inspired by the observation of hidden-charm pentaquark $P_c$ and $P_{cs}$ states by the LHCb Collaboration, we explore the $qqc\bar{c}c$ ($q~=~u$ or $d$) pentaquark systems in the quark delocalization color screening model. The interaction between baryons and mesons and the influence of channel coupling are studied in this work. Three compact $qqc\bar{c}c$ pentaquark states are obtained, whose masses are 5259 MeV with $I(J^P)$ = $0(1/2^-)$, 5396 MeV with $I(J^P)$ = $1(1/2^-)$, and 5465 MeV with $I(J^P)$ = $1(3/2^-)$. Two molecular states are obtained, which are $I(J^P)$ = $0(1/2^-)$ $\Lambda_c J/\psi$ with 5367 MeV and $I(J^P)$ = $0(5/2^-)$ $\Xi_{cc}^* \bar{D}^*$ with 5690 MeV. These predicted states may provide important information for future experimental search.

hep-ph

Exotic $Qq\bar{q}\bar{q}$ states in the chiral quark model

In the framework of the chiral quark model, we investigate the $Qq\bar{q}\bar{q}$ ($Q= c, b$ and $q= u, d$) tetraquark system with two structures: $Q\bar{q}$-$q\bar{q}$ and $Qq$-$\bar{q}\bar{q}$. The bound-state calculation shows that for the single channel, there is no evidence for any bound state below the minimum threshold in both $cq\bar{q}\bar{q}$ and $bq\bar{q}\bar{q}$ systems. However, after coupling all channels of two structures, we obtain a bound state below the minimum threshold in the $cq\bar{q}\bar{q}$ system with the energy of $1998$ MeV, and the quantum number is $IJ^{P}=\frac{1}{2}0^{+}$. Meanwhile, in the $bq\bar{q}\bar{q}$ system, two bound states with energies of $5414$ MeV and $5456$ MeV are obtained, and the quantum numbers are $IJ^{P}=\frac{1}{2}0^{+}$ and $IJ^{P}=\frac{1}{2}1^{+}$, respectively. Besides, we also employe the real-scaling method to search for resonance states in the $cq\bar{q}\bar{q}$ and $bq\bar{q}\bar{q}$ systems. Unfortunately, no genuine resonance states were obtained in both systems. We suggest future experiments to search for these three possible bound states.

hep-ph

Safe-VLN: Collision Avoidance for Vision-and-Language Navigation of Autonomous Robots Operating in Continuous Environments

The task of vision-and-language navigation in continuous environments (VLN-CE) aims at training an autonomous agent to perform low-level actions to navigate through 3D continuous surroundings using visual observations and language instructions. The significant potential of VLN-CE for mobile robots has been demonstrated across a large number of studies. However, most existing works in VLN-CE focus primarily on transferring the standard discrete vision-and-language navigation (VLN) methods to continuous environments, overlooking the problem of collisions. Such oversight often results in the agent deviating from the planned path or, in severe instances, the agent being trapped in obstacle areas and failing the navigational task. To address the above-mentioned issues, this paper investigates various collision scenarios within VLN-CE and proposes a classification method to predicate the underlying causes of collisions. Furthermore, a new VLN-CE algorithm, named Safe-VLN, is proposed to bolster collision avoidance capabilities including two key components, i.e., a waypoint predictor and a navigator. In particular, the waypoint predictor leverages a simulated 2D LiDAR occupancy mask to prevent the predicted waypoints from being situated in obstacle-ridden areas. The navigator, on the other hand, employs the strategy of `re-selection after collision' to prevent the robot agent from becoming ensnared in a cycle of perpetual collisions. The proposed Safe-VLN is evaluated on the R2R-CE, the results of which demonstrate an enhanced navigational performance and a statistically significant reduction in collision incidences.

cs.RO