SearcharxivSearch

arXiv subjects

Yingying Zhou

Publications and source records attributed to Yingying Zhou.

9 recordsLinked to original sources

MIDAS: Mutual Information Disentanglement with Uncertainty-Aware Fusion for Incomplete Multimodal Sentiment Analysis

Most existing multimodal sentiment analysis approaches assume access to complete multimodal inputs. However, real-world applications frequently encounter incomplete or corrupted modalities, posing a critical challenge. Although several methods have been proposed to tackle this issue, they mainly rely on data imputation and heuristic coordination constraints, which fail to effectively extract and leverage task-relevant information from the incomplete multimodal data. To address this challenge, we propose a unified framework termed Mutual Information Disentanglement with uncertainty-Aware fuSion (MIDAS), which effectively restructures multimodal representations under incomplete conditions. MIDAS adopts a variational modeling strategy to represent each modality with multivariate Gaussian latent variables and further decomposes them into shared and exclusive factors. To obtain reliable representations, we design a minimax objective that minimizes the mutual information between shared and exclusive spaces for stable disentanglement, while maximizing the mutual information among shared spaces across modalities to enhance semantic alignment. In addition, an uncertainty-aware fusion mechanism is introduced, where posterior variance is leveraged as a reliability indicator to adaptively weight latent features during fusion, ensuring robust integration even when modalities are incomplete. Extensive experiments on three widely used datasets show that MIDAS achieves strong and consistent performance gains over competitive baselines across a wide range of incomplete settings, demonstrating its effectiveness and robustness for incomplete data scenarios.

cs.AI

DashFusion: Dual-stream Alignment with Hierarchical Bottleneck Fusion for Multimodal Sentiment Analysis

Multimodal sentiment analysis (MSA) integrates various modalities, such as text, image, and audio, to provide a more comprehensive understanding of sentiment. However, effective MSA is challenged by alignment and fusion issues. Alignment requires synchronizing both temporal and semantic information across modalities, while fusion involves integrating these aligned features into a unified representation. Existing methods often address alignment or fusion in isolation, leading to limitations in performance and efficiency. To tackle these issues, we propose a novel framework called Dual-stream Alignment with Hierarchical Bottleneck Fusion (DashFusion). Firstly, dual-stream alignment module synchronizes multimodal features through temporal and semantic alignment. Temporal alignment employs cross-modal attention to establish frame-level correspondences among multimodal sequences. Semantic alignment ensures consistency across the feature space through contrastive learning. Secondly, supervised contrastive learning leverages label information to refine the modality features. Finally, hierarchical bottleneck fusion progressively integrates multimodal information through compressed bottleneck tokens, which achieves a balance between performance and computational efficiency. We evaluate DashFusion on three datasets: CMU-MOSI, CMU-MOSEI, and CH-SIMS. Experimental results demonstrate that DashFusion achieves state-of-the-art performance across various metrics, and ablation studies confirm the effectiveness of our alignment and fusion techniques. The codes for our experiments are available at https://github.com/ultramarineX/DashFusion.

cs.CV

CSST Slitless Spectra: Target Detection and Classification with YOLO

Addressing the spatial uncertainty and spectral blending challenges in CSST slitless spectroscopy, we present a deep learning-driven, end-to-end framework based on the You Only Look Once (YOLO) models. This approach directly detects, classifies, and analyzes spectral traces from raw 2D images, bypassing traditional, error-accumulating pipelines. YOLOv5 effectively detects both compact zero-order and extended first-order traces even in highly crowded fields. Building on this, YOLO11 integrates source classification (star/galaxy) and discrete astrophysical parameter estimation (e.g., redshift bins), showcasing complete spectral trace analysis without other manual preprocessing. Our framework processes large images rapidly, learning spectral-spatial features holistically to minimize errors. We achieve high trace detection precision (YOLOv5) and demonstrate successful quasar identification and binned redshift estimation (YOLO11). This study establishes machine learning as a paradigm shift in slitless spectroscopy, unifying detection, classification, and preliminary parameter estimation in a scalable system. Future research will concentrate on direct, continuous prediction of astrophysical parameters from raw spectral traces.

astro-ph.IM

Slitless Spectroscopy Source Detection Using YOLO Deep Neural Network

Slitless spectroscopy eliminates the need for slits, allowing light to pass directly through a prism or grism to generate a spectral dispersion image that encompasses all celestial objects within a specified area. This technique enables highly efficient spectral acquisition. However, when processing CSST slitless spectroscopy data, the unique design of its focal plane introduces a challenge: photometric and slitless spectroscopic images do not have a one-to-one correspondence. As a result, it becomes essential to first identify and count the sources in the slitless spectroscopic images before extracting spectra. To address this challenge, we employed the You Only Look Once (YOLO) object detection algorithm to develop a model for detecting targets in slitless spectroscopy images. This model was trained on 1,560 simulated CSST slitless spectroscopic images. These simulations were generated from the CSST Cycle 6 and Cycle 9 main survey data products, representing the Galactic and nearby galaxy regions and the high galactic latitude regions, respectively. On the validation set, the model achieved a precision of 88.6% and recall of 90.4% for spectral lines, and 87.0% and 80.8% for zeroth-order images. In testing, it maintained a detection rate >80% for targets brighter than 21 mag (medium-density regions) and 20 mag (low-density regions) in the Galactic and nearby galaxies regions, and >70% for targets brighter than 18 mag in high galactic latitude regions.

astro-ph.IM

Deep Learning Approaches for Multimodal Intent Recognition: A Survey

Intent recognition aims to identify users' underlying intentions, traditionally focusing on text in natural language processing. With growing demands for natural human-computer interaction, the field has evolved through deep learning and multimodal approaches, incorporating data from audio, vision, and physiological signals. Recently, the introduction of Transformer-based models has led to notable breakthroughs in this domain. This article surveys deep learning methods for intent recognition, covering the shift from unimodal to multimodal techniques, relevant datasets, methodologies, applications, and current challenges. It provides researchers with insights into the latest developments in multimodal intent recognition (MIR) and directions for future research.

cs.CL

AGCo-MATA: Air-Ground Collaborative Multi-Agent Task Allocation in Mobile Crowdsensing

Rapid progress in intelligent unmanned systems has presented new opportunities for mobile crowd sensing (MCS). Today, heterogeneous air-ground collaborative multi-agent framework, which comprise unmanned aerial vehicles (UAVs) and unmanned ground vehicles (UGVs), have presented superior flexibility and efficiency compared to traditional homogeneous frameworks in complex sensing tasks. Within this context, task allocation among different agents always play an important role in improving overall MCS quality. In order to better allocate tasks among heterogeneous collaborative agents, in this paper, we investigated two representative complex multi-agent task allocation scenarios with dual optimization objectives: (1) For AG-FAMT (Air-Ground Few Agents More Tasks) scenario, the objectives are to maximize the task completion while minimizing the total travel distance; (2) For AG-MAFT (Air-Ground More Agents Few Tasks) scenario, where the agents are allocated based on their locations, has the optimization objectives of minimizing the total travel distance while reducing travel time cost. To achieve this, we proposed a Multi-Task Minimum Cost Maximum Flow (MT-MCMF) optimization algorithm tailored for AG-FAMT, along with a multi-objective optimization algorithm called W-ILP designed for AG-MAFT, with a particular focus on optimizing the charging path planning of UAVs. Our experiments based on a large-scale real-world dataset demonstrated that the proposed two algorithms both outperform baseline approaches under varying experimental settings, including task quantity, task difficulty, and task distribution, providing a novel way to improve the overall quality of mobile crowdsensing tasks.

cs.MA

Understanding the velocity distribution of the Galactic Bulge with APOGEE and Gaia

We revisit the stellar velocity distribution in the Galactic bulge/bar region with APOGEE DR16 and {\it Gaia} DR2, focusing in particular on the possible high-velocity (HV) peaks and their physical origin. We fit the velocity distributions with two different models, namely with Gauss-Hermite polynomial and Gaussian mixture model (GMM). The result of the fit using Gauss-Hermite polynomials reveals a positive correlation between the mean velocity ($\bar{V}$) and the "skewness" ($h_{3}$) of the velocity distribution, possibly caused by the Galactic bar. The $n=2$ GMM fitting reveals a symmetric longitudinal trend of $|μ_{2}|$ and $σ_{2}$ (the mean velocity and the standard deviation of the secondary component), which is inconsistent to the $x_{2}$ orbital family predictions. Cold secondary peaks could be seen at $|l|\sim6^\circ$. However, with the additional tangential information from {\it Gaia}, we find that the HV stars in the bulge show similar patterns in the radial-tangential velocity distribution ($V_{\rm R}-V_{\rm T}$), regardless of the existence of a distinct cold HV peak. The observed $V_{\rm R}-V_{\rm T}$ (or $V_{\rm GSR}-μ_{l}$) distributions are consistent with the predictions of a simple MW bar model. The chemical abundances and ages inferred from ASPCAP and CANNON suggest that the HV stars in the bulge/bar are generally as old as, if not older than, the other stars in the bulge/bar region.

astro-ph.GA

Shape of LOSVD\lowercase{s} in barred disks: Implications for future IFU surveys

The shape of LOSVDs (line-of-sight velocity distributions) carries important information about the internal dynamics of galaxies. The skewness of LOSVDs represents their asymmetric deviation from a Gaussian profile. Correlations between the skewness parameter ($h_3$) and the mean velocity (\vm) of a Gauss-Hermite series reflect the underlying stellar orbital configurations of different morphological components. Using two self-consistent $N$-body simulations of disk galaxies with different bar strengths, we investigate \hv\ correlations at different inclination angles. Similar to previous studies, we find anticorrelations in the disk area, and positive correlations in the bar area when viewed edge-on. However, at intermediate inclinations, the outer parts of bars exhibit anticorrelations, while the core areas dominated by the boxy/peanut-shaped (B/PS) bulges still maintain weak positive correlations. When viewed edge-on, particles in the foreground/background disk (the wing region) in the bar area constitute the main velocity peak, whereas the particles in the bar contribute to the high-velocity tail, generating the \hv\ correlation. If we remove the wing particles, the LOSVDs of the particles in the outer part of the bar only exhibit a low-velocity tail, resulting in a negative \hv\ correlation, whereas the core areas in the central region still show weakly positive correlations. We discuss implications for IFU observations on bars, and show that the variation of the \hv\ correlation in the disk galaxy may be used as a kinematic indicator of the bar and the B/PS bulge.

astro-ph.GA

Chemical abundances and ages of the bulge stars in APOGEE high-velocity peaks

A cold high-velocity (HV, $\sim$ 200 km/s) peak was first reported in several Galactic bulge fields based on the APOGEE commissioning observations. Both the existence and the nature of the high-velocity peak are still under debate. Here we revisit this feature with the latest APOGEE DR13 data. We find that most of the low latitude bulge fields display a skewed Gaussian distribution with a HV shoulder. However, only 3 out of 53 fields show distinct high-velocity peaks around 200 km/s. The velocity distribution can be well described by Gauss-Hermite polynomials, except the three fields showing clear HV peaks. We find that the correlation between the skewness parameter ($h_{3}$) and the mean velocity ($\bar{v}$), instead of a distinctive HV peak, is a strong indicator of the bar. It was recently suggested that the HV peak is composed of preferentially young stars. We choose three fields showing clear HV peaks to test this hypothesis using the metallicity, [$α$/M] and [C/N] as age proxies. We find that both young and old stars show HV features. The similarity between the chemical abundances of stars in the HV peaks and the main component indicates that they are not systematically different in terms of chemical abundance or age. In contrast, there are clear differences in chemical space between stars in the Sagittarius dwarf and the bulge stars. The strong HV peaks off-plane are still to be explained properly, and could be different in nature.

astro-ph.GA