SearcharxivSearch

arXiv subjects

Siyuan Hou

Publications and source records attributed to Siyuan Hou.

8 recordsLinked to original sources

SALMONN-2: Advancing General-Purpose Hearing Abilities with Self-Supervised Representations

Recent audio large language models (ALLMs) are typically built upon audio encoders trained with large amounts of supervised data. Since self-supervised learning (SSL) audio encoder models are known to learn general-purpose and transferable representations, we investigate whether general-purpose SSL audio representations can serve as an effective foundation for ALLMs. We present SALMONN-2, an ALLM built upon a unified SSL encoder. To better exploit the hierarchical representations learned by SSL encoders, we propose a multi-layer feature fusion (MLF) adapter that aggregates information from all encoder layers before projecting them into the language model. Beyond conventional audio understanding tasks, we further explore multimodal in-context learning (MICL) in ALLMs and study how this capability can be acquired through contextual biasing training. Experimental results show that a general-purpose SSL encoder achieves performance comparable to, or better than, specialised supervised audio encoders while providing a more balanced capability across speech, audio, music and paralinguistic tasks. SALMONN-2 further achieves state-of-the-art performance among comparable-scale open-weight models on ALLM understanding benchmarks, obtaining the best results on MMAU-Pro, MMAR and MMSU. We also show that MICL does not emerge naturally in ALLMs, but can be effectively acquired through targeted contextual biasing training.

eess.AS

SciTS: Scientific Time Series Understanding and Generation with LLMs

The scientific reasoning ability of large language models (LLMs) has recently attracted significant attention. Time series, as a fundamental modality in scientific data, presents unique challenges that are often overlooked in current multimodal LLMs, which either encode numerical sequences as text or convert them into images. Such approaches may be insufficient for comprehensive scientific time series understanding and generation. Existing unified time series models typically specialise in either forecasting or analysis, and their effectiveness on non-periodic, heterogeneous scientific signals remains unclear. To address these gaps, we introduce SciTS, a benchmark spanning 12 scientific domains and 43 tasks, with over 50k+ instances, both univariate and multivariate signals ranging from $10^0$ to $10^7$ in length and up to 10~MHz in frequency. We benchmark 17 models, including text-only LLMs, multimodal LLMs, and unified time series models, and find that general-purpose LLMs exhibit stronger generalisability than specialised time series models, while representing time series as text or images limits their performance due to excessively long sequences and loss of numerical precision, respectively. We then introduce TimeOmni, a framework that equips LLMs with the ability to understand and generate time series while remaining compatible with general-purpose LLM training. This work fills a gap in both dedicated benchmarks and modelling frameworks for scientific time series, paving the way for LLMs to understand and generate complex temporal scientific data.

cs.LG

Self-Interacting Dark Matter with Mass Segregation: A Unified Explanation of Dwarf Cores and Small-Scale Lenses

In two-component self-interacting dark matter (SIDM) models with inter-species interactions, mass segregation arises naturally from collisional relaxation, enhancing central densities and gravothermal evolution. We demonstrate that models with velocity-dependent interactions, both within and between species, can connect several small-scale observations while remaining consistent with cluster-scale constraints. This combination enables core formation in dwarf halos, where the presence of baryons increases the inner densities and enhances the predicted strong lensing signatures. Using cosmological and controlled simulations alongside an accurate parametric model, we present proof-of-principle examples showing that this framework can explain the structure of dark perturbers observed in strong lensing systems, and can enhance the efficiency of small-scale lenses by a factor of a few, in line with the excess reported in galaxy-galaxy strong-lensing observations. Importantly, mass segregation can enhance the Einstein radii of SIDM halos relative to their cold dark matter (CDM) counterparts, overcoming a key challenge in one-component SIDM scenarios. Our results present mass segregation in two-component SIDM as a self-consistent, testable framework with the potential to address multiple small-scale challenges in structure formation.

astro-ph.CO

Flux-ratio anomalies in cusp quasars reveal dark matter beyond CDM

Strongly lensed quasars in cusp configurations provide a uniquely sensitive probe of small-scale dark matter structure. Using the largest microlensing-free flux ratios for 17 quadruply imaged cusps, we combine these with extensive Monte Carlo simulations of mock lens realizations under cold dark matter (CDM), self-interacting dark matter (SIDM), and fuzzy dark matter (FDM) scenarios. Building on this, we propose a region (minor-axis and narrow major-axis cusp lenses) where flux-ratio anomalies persist even under globally parameterized models ("macromodels") with multipole freedom (capturing disk, asymmetric, or merger-driven structures). Within this region, J1042+1641 is $>3σ$ incompatible with both CDM and SIDM. Our results yield a Bayes factor exceeding $100$, providing very strong evidence for FDM over even the most optimistic CDM and SIDM scenarios. As only 11 cusp lenses lie within this region, extending to larger samples will be essential for assessing its statistical generality and for decisively confirming these findings with future microlensing-free flux ratio data.

astro-ph.CO

Intern-S1: A Scientific Multimodal Foundation Model

In recent years, a plethora of open-source foundation models have emerged, achieving remarkable progress in some widely attended fields, with performance being quite close to that of closed-source models. However, in high-value but more challenging scientific professional fields, either the fields still rely on expert models, or the progress of general foundation models lags significantly compared to those in popular areas, far from sufficient for transforming scientific research and leaving substantial gap between open-source models and closed-source models in these scientific domains. To mitigate this gap and explore a step further toward Artificial General Intelligence (AGI), we introduce Intern-S1, a specialized generalist equipped with general understanding and reasoning capabilities with expertise to analyze multiple science modal data. Intern-S1 is a multimodal Mixture-of-Experts (MoE) model with 28 billion activated parameters and 241 billion total parameters, continually pre-trained on 5T tokens, including over 2.5T tokens from scientific domains. In the post-training stage, Intern-S1 undergoes offline and then online reinforcement learning (RL) in InternBootCamp, where we propose Mixture-of-Rewards (MoR) to synergize the RL training on more than 1000 tasks simultaneously. Through integrated innovations in algorithms, data, and training systems, Intern-S1 achieved top-tier performance in online RL training. On comprehensive evaluation benchmarks, Intern-S1 demonstrates competitive performance on general reasoning tasks among open-source models and significantly outperforms open-source models in scientific domains, surpassing closed-source state-of-the-art models in professional tasks, such as molecular synthesis planning, reaction condition prediction, predicting thermodynamic stabilities for crystals. Our models are available at https://huggingface.co/internlm/Intern-S1.

cs.LG

MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Audio comprehension-including speech, non-speech sounds, and music-is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio understanding to qualify as generally intelligent. However, evaluating auditory intelligence comprehensively remains challenging. To address this gap, we introduce MMAU-Pro, the most comprehensive and rigorously curated benchmark for assessing audio intelligence in AI systems. MMAU-Pro contains 5,305 instances, where each instance has one or more audios paired with human expert-generated question-answer pairs, spanning speech, sound, music, and their combinations. Unlike existing benchmarks, MMAU-Pro evaluates auditory intelligence across 49 unique skills and multiple complex dimensions, including long-form audio comprehension, spatial audio reasoning, multi-audio understanding, among others. All questions are meticulously designed to require deliberate multi-hop reasoning, including both multiple-choice and open-ended response formats. Importantly, audio data is sourced directly ``from the wild" rather than from existing datasets with known distributions. We evaluate 22 leading open-source and proprietary multimodal AI models, revealing significant limitations: even state-of-the-art models such as Gemini 2.5 Flash and Audio Flamingo 3 achieve only 59.2% and 51.7% accuracy, respectively, approaching random performance in multiple categories. Our extensive analysis highlights specific shortcomings and provides novel insights, offering actionable perspectives for the community to enhance future AI systems' progression toward audio general intelligence. The benchmark and code is available at https://sonalkum.github.io/mmau-pro.

eess.AS

A Universal Analytic Model for Gravitational Lensing by Self-Interacting Dark Matter Halos

We present a model for analytically calculating gravitational lensing by self-interacting dark matter (SIDM) halos. Leveraging the universal behavior of SIDM halos during gravothermal evolution, we calibrate the lensing potential using a fluid simulation, normalizing the evolution time to align with established scenarios. From this potential, we derive explicit equations for the deflection angle and surface density profile, quantifying their deviations from numerical results. Our model builds on the parametric approach of arXiv:2305.16176, providing refinements in the deep core-collapse regime and enabling more comprehensive lensing studies. We explore characteristic lensing features, including critical curves and caustics, for SIDM halos in isolation and within a main halo, tracking their evolution through the gravothermal phase. We also examine signatures in the self-similar regime of core collapsed halos and highlight the role of baryonic effects in realistic halos. The application of our model extends to generic halos, whose profiles fit one or a superposition of our parametric forms. We make our implementation publicly available on https://github.com/HouSiyuan2001/SIDM_Lensing_Model to support further research.

astro-ph.CO

Editing Music with Melody and Text: Using ControlNet for Diffusion Transformer

Despite the significant progress in controllable music generation and editing, challenges remain in the quality and length of generated music due to the use of Mel-spectrogram representations and UNet-based model structures. To address these limitations, we propose a novel approach using a Diffusion Transformer (DiT) augmented with an additional control branch using ControlNet. This allows for long-form and variable-length music generation and editing controlled by text and melody prompts. For more precise and fine-grained melody control, we introduce a novel top-$k$ constant-Q Transform representation as the melody prompt, reducing ambiguity compared to previous representations (e.g., chroma), particularly for music with multiple tracks or a wide range of pitch values. To effectively balance the control signals from text and melody prompts, we adopt a curriculum learning strategy that progressively masks the melody prompt, resulting in a more stable training process. Experiments have been performed on text-to-music generation and music-style transfer tasks using open-source instrumental recording data. The results demonstrate that by extending StableAudio, a pre-trained text-controlled DiT model, our approach enables superior melody-controlled editing while retaining good text-to-music generation performance. These results outperform a strong MusicGen baseline in terms of both text-based generation and melody preservation for editing. Audio examples can be found at https://stable-audio-control.github.io.

eess.AS