SearcharxivSearch

arXiv subjects

Yu Tu

Publications and source records attributed to Yu Tu.

7 recordsLinked to original sources

MiMo-Audio: Audio Language Models are Few-Shot Learners

Existing audio language models typically rely on task-specific fine-tuning to accomplish particular audio tasks. In contrast, humans are able to generalize to new audio tasks with only a few examples or simple instructions. GPT-3 has shown that scaling next-token prediction pretraining enables strong generalization capabilities in text, and we believe this paradigm is equally applicable to the audio domain. By scaling MiMo-Audio's pretraining data to over one hundred million of hours, we observe the emergence of few-shot learning capabilities across a diverse set of audio tasks. We develop a systematic evaluation of these capabilities and find that MiMo-Audio-7B-Base achieves SOTA performance on both speech intelligence and audio understanding benchmarks among open-source models. Beyond standard metrics, MiMo-Audio-7B-Base generalizes to tasks absent from its training data, such as voice conversion, style transfer, and speech editing. MiMo-Audio-7B-Base also demonstrates powerful speech continuation capabilities, capable of generating highly realistic talk shows, recitations, livestreaming and debates. At the post-training stage, we curate a diverse instruction-tuning corpus and introduce thinking mechanisms into both audio understanding and generation. MiMo-Audio-7B-Instruct achieves open-source SOTA on audio understanding benchmarks (MMSU, MMAU, MMAR, MMAU-Pro), spoken dialogue benchmarks (Big Bench Audio, MultiChallenge Audio) and instruct-TTS evaluations, approaching or surpassing closed-source models. Model checkpoints and full evaluation suite are available at https://github.com/XiaomiMiMo/MiMo-Audio.

cs.CL

MiMo-VL Technical Report

We open-source MiMo-VL-7B-SFT and MiMo-VL-7B-RL, two powerful vision-language models delivering state-of-the-art performance in both general visual understanding and multimodal reasoning. MiMo-VL-7B-RL outperforms Qwen2.5-VL-7B on 35 out of 40 evaluated tasks, and scores 59.4 on OlympiadBench, surpassing models with up to 78B parameters. For GUI grounding applications, it sets a new standard with 56.1 on OSWorld-G, even outperforming specialized models such as UI-TARS. Our training combines four-stage pre-training (2.4 trillion tokens) with Mixed On-policy Reinforcement Learning (MORL) integrating diverse reward signals. We identify the importance of incorporating high-quality reasoning data with long Chain-of-Thought into pre-training stages, and the benefits of mixed RL despite challenges in simultaneous multi-domain optimization. We also contribute a comprehensive evaluation suite covering 50+ tasks to promote reproducibility and advance the field. The model checkpoints and full evaluation suite are available at https://github.com/XiaomiMiMo/MiMo-VL.

cs.CL

MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining

We present MiMo-7B, a large language model born for reasoning tasks, with optimization across both pre-training and post-training stages. During pre-training, we enhance the data preprocessing pipeline and employ a three-stage data mixing strategy to strengthen the base model's reasoning potential. MiMo-7B-Base is pre-trained on 25 trillion tokens, with additional Multi-Token Prediction objective for enhanced performance and accelerated inference speed. During post-training, we curate a dataset of 130K verifiable mathematics and programming problems for reinforcement learning, integrating a test-difficulty-driven code-reward scheme to alleviate sparse-reward issues and employing strategic data resampling to stabilize training. Extensive evaluations show that MiMo-7B-Base possesses exceptional reasoning potential, outperforming even much larger 32B models. The final RL-tuned model, MiMo-7B-RL, achieves superior performance on mathematics, code and general reasoning tasks, surpassing the performance of OpenAI o1-mini. The model checkpoints are available at https://github.com/xiaomimimo/MiMo.

cs.CL

An adaptive switch strategy for acquisition functions in Bayesian optimization of wind farm layout

Wind farm layout optimization (WFLO), which seeks to maximizing annual energy production by strategically adjusting wind turbines' location, is essential for the development of large-scale wind farms. While low-fidelity methods dominate WFLO studies, high-fidelity methods are less commonly applied due to their significant computational costs. This paper introduces a Bayesian optimization framework that leverages a novel adaptive acquisition function switching strategy to enhance the efficiency and effectiveness of WFLO using high-fidelity modeling methods. The proposed switch acquisition functions strategy alternates between MSP and MES acquisition functions, dynamically balancing exploration and exploitation. By iteratively retraining the Kriging model with intermediate optimal layouts, the framework progressively refines its predictions to accelerate convergence to optimal solutions. The performance of the switch-acquisition-function-based Bayesian optimization framework is first validated using 4- and 10-dimensional Ackley benchmark functions, where it demonstrates superior optimization efficiency compared to using MSP or MES alone. The framework is then applied to WFLO problems using Gaussian wake models for three varying wind farm cases. Results show that the switch-acquisition-function-based Bayesian optimization framework outperforms traditional heuristic algorithms, achieving near-optimal annual energy output with significantly fewer calculations. Finally, the framework is extended to high-fidelity WFLO by coupling it with CFD simulations, where turbine rotors are modeled as actuator disks. The novel switch-acquisition-function-based Bayesian optimization enables more effective exploration to achieve higher annual energy production in WFLO, advancing the design of more effective wind farm layouts.

math.OC

An optimization framework for wind farm layout design using CFD-based Kriging model

Wind farm layout optimization (WFLO) seeks to alleviate the wake loss and maximize wind farm power output efficiency, and is a crucial process in the design of wind energy projects.Since the optimization algorithms typically require thousands of numerical evaluations of the wake effects, conventional WFLO studies are usually carried out with the low-fidelity analytical wake models.In this paper, we develop an optimization framework for wind farm layout design using CFD-based Kriging model to maximize the annual energy production (AEP) of wind farms. This surrogate-based optimization (SBO) framework uses latin hypercube sampling to generate a group of wind farm layout samples, based on which CFD simulations are carried out to obtain the corresponding AEPs.This wind farm layout dataset is used to train the Kriging model, which is then integrated with an optimizer based on genetic algorithm (GA). As the optimization progresses, the intermediate optimal layout designs are again fed into the dataset.Such adaptive update of wind farm layout dataset continues until the algorithm converges.To evaluate the performance of the proposed SBO framework, we apply it to three representative wind farm cases.Compared to the conventional staggered layout, the optimized wind farm produces significantly higher total AEP.In particular, the SBO framework requires significantly smaller number of CFD calls to yield the optimal layouts that generates almost the same AEP with the direct CFD-GA method.Further analysis on the velocity fields show that the optimization framework attempts to locate the downstream turbines away from the the wakes of upstream ones.The proposed CFD-based surrogate model provides a more accurate and flexible alternative to the conventional analytical-wake-model-based methods in WFLO tasks, and has the potential to be used for designing efficient wind farm projects.

physics.flu-dyn

Aerodynamic performance enhancement and noise reduction for Darrieus vertical axis wind turbines: V-shaped blades and trailing-edge serrations

The Darrieus vertical axis wind turbines (VAWTs) are one of the mainstream devices for wind energy utilization in the urban areas. The market's pursuit of high wind energy conversion efficiency promotes the research on improving the wind turbine power coefficient while reducing noise. This paper makes further investigation on the aerodynamics and aeroacoustics of a small VAWT with V-shaped blades and trailing-edge serrations. The feasibility of utilizing the Reynolds-Averaged Navier-Stokes SST k-omega turbulence model and the FW-H method is verified against experiments. The studied V-shaped blades can effectively improve the power performance of VAWT over a wide range of tip speed ratios under normal wind speed conditions, and the trailing-edge serrations will also slightly increase the power output of V-bladed VAWT at the optimal tip speed ratio. The power coefficient of the V-bladed wind turbine with trailing-edge serrations is about 28.3% higher than that of the original turbine. In addition, a dumbbell-shaped noise directivity distribution was first discovered in the VAWT compared with the traditional elliptical distribution. The V-bladed VAWT generated less low-frequency noise and the trailing-edge serrations realized the expected noise reduction effect. Practically, this study proposes a feasible solution for the design of high-efficiency and low-noise wind turbines.

physics.flu-dyn

Aerodynamic characterization of two tandem wind turbines under yaw misalignment control using actuator line model

Yaw control has proven to be promising in alleviating the wake effects that plague the efficiency of wind farms. In this work, the actuator line modeling (ALM) method is adopted to simulate the flows over two tandem turbines distanced by $3 - 7$ rotor diameters, with the yaw angle of the upstream rotor varying from $\gamma_1=0^{\circ}$ to $50^{\circ}$. The aim is to provide a comprehensive aerodynamic characterization of a simple wind farm under yaw misalignment. With increasing yaw angle, the power generated by the downstream rotor increases, compensating the power loss in the upstream rotor, and resulting in higher total power of the two turbines than that without yaw control. The maximum power output is achieved as the upstream wake of the yawed rotor is redirected away from the downstream rotor plane. Behind the downstream rotor, the secondary steering phenomenon is observed, where the wake is also redirected from the centerline. The use of the actuator line model also reveal unsteady aerodynamic characteristics that can not be captured by lower-fidelity models. For the upstream rotor, the yaw misalignment results in time-varying change in the local angle of attack on the blade, giving rise to unsteady loading. The downstream rotor is partially submerged in the deflected wake incurred by the yawed upstream rotor. As the blade revolves into and out of the wake deficit, the blade experiences cyclic loading, leading to even stronger fluctuations in the aerodynamic loads than the upstream rotor. These analysis provides a comprehensive understanding of the yaw control effects on the two tandem rotors from the perspectives of aerodynamic performance, wake profiles, and unsteady characteristics. The insights gained from the study can aid the design of collective yaw control strategies of wind farms, and lay the foundation for assessing the fatigue damage associated with yaw misalignment.

physics.flu-dyn