SearcharxivSearch

arXiv subjects

Qihan Hu

Publications and source records attributed to Qihan Hu.

3 recordsLinked to original sources

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models

Recent advances in reasoning models have driven significant progress in text and multimodal domains, yet audio reasoning remains relatively limited. Only a few Large Audio Language Models (LALMs) incorporate explicit Chain-of-Thought (CoT) reasoning, and their capabilities are often inconsistent and insufficient for complex tasks. To bridge this gap, we introduce Audio-Cogito, a fully open-source solution for deep audio reasoning. We develop Cogito-pipe for high-quality audio reasoning data curation, producing 545k reasoning samples. Based on this dataset, we adopt a self-distillation strategy for model fine-tuning. Experiments on the MMAR benchmark, the only audio benchmark evaluating the CoT process, show that our model achieves the best performance among open-source models and matches or surpasses certain closed-source models in specific metrics. Our approach also ranks among the top-tier systems in the Interspeech 2026 Audio Reasoning Challenge.

eess.AS

Efficient Multi-View Fusion and Flexible Adaptation to View Missing in Cardiovascular System Signals

The progression of deep learning and the widespread adoption of sensors have facilitated automatic multi-view fusion (MVF) about the cardiovascular system (CVS) signals. However, prevalent MVF model architecture often amalgamates CVS signals from the same temporal step but different views into a unified representation, disregarding the asynchronous nature of cardiovascular events and the inherent heterogeneity across views, leading to catastrophic view confusion. Efficient training strategies specifically tailored for MVF models to attain comprehensive representations need simultaneous consideration. Crucially, real-world data frequently arrives with incomplete views, an aspect rarely noticed by researchers. Thus, the View-Centric Transformer (VCT) and Multitask Masked Autoencoder (M2AE) are specifically designed to emphasize the centrality of each view and harness unlabeled data to achieve superior fused representations. Additionally, we systematically define the missing-view problem for the first time and introduce prompt techniques to aid pretrained MVF models in flexibly adapting to various missing-view scenarios. Rigorous experiments involving atrial fibrillation detection, blood pressure estimation, and sleep staging-typical health monitoring tasks-demonstrate the remarkable advantage of our method in MVF compared to prevailing methodologies. Notably, the prompt technique requires finetuning less than 3% of the entire model's data, substantially fortifying the model's resilience to view missing while circumventing the need for complete retraining. The results demonstrate the effectiveness of our approaches, highlighting their potential for practical applications in cardiovascular health monitoring. Codes and models are released at URL.

cs.LG

Data-driven approach for modeling Reynolds stress tensor with invariance preservation

The present study represents a data-driven turbulent model with Galilean invariance preservation based on machine learning algorithm. The fully connected neural network (FCNN) and tensor basis neural network (TBNN) [Ling et al. (2016)] are established. The models are trained based on five kinds of flow cases with Reynolds Averaged Navier-Stokes (RANS) and high-fidelity data. The mappings between two invariant sets, mean strain rate tensor and mean rotation rate tensor as well as additional consideration of invariants of turbulent kinetic energy gradients, and the Reynolds stress anisotropy tensor are trained. The prediction of the Reynolds stress anisotropy tensor is treated as user's defined RANS turbulent model with a modified turbulent kinetic energy transport equation. The results show that both FCNN and TBNN models can provide more accurate predictions of the anisotropy tensor and turbulent state in square duct flow and periodic flow cases compared to the RANS model. The machine learning based turbulent model with turbulent kinetic energy gradient related invariants can improve the prediction precision compared with only mean strain rate tensor and mean rotation rate tensor based models. The TBNN model is able to predict a better flow velocity profile compared with FCNN model due to a prior physical knowledge.

physics.flu-dyn