SearcharxivSearch

arXiv subjects

Ruixi Zhang

Publications and source records attributed to Ruixi Zhang.

11 recordsLinked to original sources

SenseNova-U1.5: Towards Native Unified Visual Intelligence

We launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture. We strengthen its visual interface through spatially coherent patch reconstruction and scale its training with carefully curated generation and editing data, improved task formulation, structural prompt enhancement, and native resolutions of up to 4K. For post-training, we optimize specialized experts for visual aesthetics, bilingual text rendering, infographic generation, and image editing, and consolidate their capabilities through multi-expert on-policy distillation. Across extensive evaluations, SenseNova-U1.5 largely advances image fidelity, text rendering, complex composition, multi-reference editing, and interleaved generation, while improving instruction following and preserving subject identity, geometry, and unmodified regions. Despite limited exposure to structured formats in its generation data, SenseNova-U1.5 generalizes effectively to long, complex, and structured visual instructions, further proving that multimodal understanding can transfer to visual planning and creation. Together, these findings position native unified modelling as a promising path towards systems that perceive, reason and create within a fully end-to-end framework. We will open-source training code, including supervised fine-tuning, reinforcement learning, and on-policy distillation.

cs.CV

Convergence of a two-parameter hyperbolic relaxation system toward the incompressible Navier-Stokes equations

We investigate a two-parameter hyperbolic relaxation approximation to the incompressible Navier-Stokes equations, incorporating a first-order relaxation and the artificial compressibility method. With vanishingly small perturbations of initial velocity, we rigorously prove the simultaneous convergence of fluid velocity and pressure toward the Navier-Stokes limit in the three-dimensional case by constructing an intermediate affine system to obtain the necessary error estimates for the pressure. Furthermore, we extend the velocity convergence analysis to the case of $\mathcal O(1)$ initial velocity perturbations, and establish the global-in-time recovery of the velocity field using a modulated energy structure and delicate bootstrap arguments in both two- and three-dimensional settings.

math.AP

Evaluating Hydro-Science and Engineering Knowledge of Large Language Models

Hydro-Science and Engineering (Hydro-SE) is a critical and irreplaceable domain that secures human water supply, generates clean hydropower energy, and mitigates flood and drought disasters. Featuring multiple engineering objectives, Hydro-SE is an inherently interdisciplinary domain that integrates scientific knowledge with engineering expertise. This integration necessitates extensive expert collaboration in decision-making, which poses difficulties for intelligence. With the rapid advancement of large language models (LLMs), their potential application in the Hydro-SE domain is being increasingly explored. However, the knowledge and application abilities of LLMs in Hydro-SE have not been sufficiently evaluated. To address this issue, we propose the Hydro-SE LLM evaluation benchmark (Hydro-SE Bench), which contains 4,000 multiple-choice questions. Hydro-SE Bench covers nine subfields and enables evaluation of LLMs in aspects of basic conceptual knowledge, engineering application ability, and reasoning and calculation ability. The evaluation results on Hydro-SE Bench show that the accuracy values vary among 0.74 to 0.80 for commercial LLMs, and among 0.41 to 0.68 for small-parameter LLMs. While LLMs perform well in subfields closely related to natural and physical sciences, they struggle with domain-specific knowledge such as industry standards and hydraulic structures. Model scaling mainly improves reasoning and calculation abilities, but there is still great potential for LLMs to better handle problems in practical engineering application. This study highlights the strengths and weaknesses of LLMs for Hydro-SE tasks, providing model developers with clear training targets and Hydro-SE researchers with practical guidance for applying LLMs.

cs.CL

Free boundary problem of low Mach number magnetohydrodynamic flows

In this paper we consider a free boundary problem of low Mach number magnetohydrodynamic flow in spatial dimension n $\geq$ 2. A priori estimates of the second fundamental form and various flow quantities in Sobolev norms are obtained by adopting the geometrical point of view introduced by Christodoulou and Lindblad [3]. Moreover, a blow up criterion is derived by using the method of Beale, Kato and Majda [2].

math.AP

A hyperbolic relaxation system of the incompressible Navier-Stokes equations with artificial compressibility

We introduce a new hyperbolic approximation to the incompressible Navier-Stokes equations by incorporating a first-order relaxation and using the artificial compressibility method. With two relaxation parameters in the model, we rigorously prove the asymptotic limit of the system towards the incompressible Navier-Stokes equations as both parameters tend to zero. Notably, the convergence of the approximate pressure variable is achieved by the help of a linear `auxiliary' system and energy-type error estimates of its differences with the two-parameter model and the Navier-Stokes equations.

math.AP

Poisson quadrature method of moments for 2D kinetic equations with velocity of constant magnitude

This work is concerned with kinetic equations with velocity of constant magnitude. We propose a quadrature method of moments based on the Poisson kernel, called Poisson-EQMOM. The derived moment closure systems are well defined for all physically relevant moments and the resultant approximations of the distribution function converge as the number of moments goes to infinity. The convergence makes our method stand out from most existing moment methods. Moreover, we devise a delicate moment inversion algorithm. As an application, the Vicsek model is studied for overdamped active particles. Then the Poisson-EQMOM is validated with a series of numerical tests including spatially homogeneous, one-dimensional and two-dimensional problems.

physics.comp-ph

Dissipativeness of the hyperbolic quadrature method of moments for kinetic equations

This paper presents a dissipativeness analysis of a quadrature method of moments (called HyQMOM) for the one-dimensional BGK equation. The method has exhibited its good performance in numerous applications. However, its mathematical foundation has not been clarified. Here we present an analytical proof of the strict hyperbolicity of the HyQMOM-induced moment closure systems by introducing a polynomial-based closure technique. As a byproduct, a class of numerical schemes for the HyQMOM system is shown to be realizability preserving under CFL-type conditions. We also show that the system preserves the dissipative properties of the kinetic equation by verifying a certain structural stability condition. The proof uses a newly introduced affine invariance and the homogeneity of the HyQMOM and heavily relies on the theory of orthogonal polynomials associated with realizable moments, in particular, the moments of the standard normal distribution.

math.NA

Are Soft Prompts Good Zero-shot Learners for Speech Recognition?

Large self-supervised pre-trained speech models require computationally expensive fine-tuning for downstream tasks. Soft prompt tuning offers a simple parameter-efficient alternative by utilizing minimal soft prompt guidance, enhancing portability while also maintaining competitive performance. However, not many people understand how and why this is so. In this study, we aim to deepen our understanding of this emerging method by investigating the role of soft prompts in automatic speech recognition (ASR). Our findings highlight their role as zero-shot learners in improving ASR performance but also make them vulnerable to malicious modifications. Soft prompts aid generalization but are not obligatory for inference. We also identify two primary roles of soft prompts: content refinement and noise information enhancement, which enhances robustness against background noise. Additionally, we propose an effective modification on noise prompts to show that they are capable of zero-shot learning on adapting to out-of-distribution noise environments.

cs.SD

Stability analysis of an extended quadrature method of moments for kinetic equations

This paper performs a stability analysis of a class of moment closure systems derived with an extended quadrature method of moments (EQMOM) for the one-dimensional BGK equation. The class is characterized with a kernel function. A sufficient condition on the kernel is identified for the EQMOM-derived moment systems to be strictly hyperbolic. We also investigate the realizability of the moment method. Moreover, sufficient and necessary conditions are established for the two-node systems to be well-defined and strictly hyperbolic, and to preserve the dissipation property of the kinetic equation.

math.NA

Contrastive Speech Mixup for Low-resource Keyword Spotting

Most of the existing neural-based models for keyword spotting (KWS) in smart devices require thousands of training samples to learn a decent audio representation. However, with the rising demand for smart devices to become more personalized, KWS models need to adapt quickly to smaller user samples. To tackle this challenge, we propose a contrastive speech mixup (CosMix) learning algorithm for low-resource KWS. CosMix introduces an auxiliary contrastive loss to the existing mixup augmentation technique to maximize the relative similarity between the original pre-mixed samples and the augmented samples. The goal is to inject enhancing constraints to guide the model towards simpler but richer content-based speech representations from two augmented views (i.e. noisy mixed and clean pre-mixed utterances). We conduct our experiments on the Google Speech Command dataset, where we trim the size of the training set to as small as 2.5 mins per keyword to simulate a low-resource condition. Our experimental results show a consistent improvement in the performance of multiple models, which exhibits the effectiveness of our method.

cs.SD

deHuBERT: Disentangling Noise in a Self-supervised Model for Robust Speech Recognition

Existing self-supervised pre-trained speech models have offered an effective way to leverage massive unannotated corpora to build good automatic speech recognition (ASR). However, many current models are trained on a clean corpus from a single source, which tends to do poorly when noise is present during testing. Nonetheless, it is crucial to overcome the adverse influence of noise for real-world applications. In this work, we propose a novel training framework, called deHuBERT, for noise reduction encoding inspired by H. Barlow's redundancy-reduction principle. The new framework improves the HuBERT training algorithm by introducing auxiliary losses that drive the self- and cross-correlation matrix between pairwise noise-distorted embeddings towards identity matrix. This encourages the model to produce noise-agnostic speech representations. With this method, we report improved robustness in noisy environments, including unseen noises, without impairing the performance on the clean set.

cs.SD