SearcharxivSearch

arXiv subjects

Jinbin Fu

Publications and source records attributed to Jinbin Fu.

2 recordsLinked to original sources

YouZhi: Towards High-Concurrency Financial LLMs via Adaptive GQA-to-MLA Transition

Large language models (LLMs) drive significant financial innovations, yet their high-concurrency deployment is severely bottlenecked by KV cache memory overhead, which inflates infrastructure costs and throttles scalability. To address this, we propose YouZhi-LLM, a highly efficient financial LLM empowered by a comprehensive structural transition and training pipeline natively built on the Huawei Ascend ecosystem. At its algorithmic core, YouZhi-LLM features a layer-adaptive GQA-to-MLA transition framework that dynamically assigns per-layer FreqFold sizes, maximizing KV-cache compression while minimizing perplexity degradation. To recover representation capacity and inject domain expertise, the Ascend-based training pipeline seamlessly integrates generalized knowledge distillation with financial-specific supervised fine-tuning. Evaluations demonstrate the superiority of this systematic approach, with the adaptive transition reducing perplexity degradation by up to 35% over uniform baselines. Crucially, when evaluated on Ascend NPUs via vLLM-Ascend, the massive KV-cache reduction translates directly into deployment efficiency. Compared to their respective base models, YouZhi-7B yields a 12.3% improvement in average financial benchmark score alongside a 2.69$\times$ increase in maximum concurrency; similarly, YouZhi-14B achieves a 7.0% accuracy gain and a 2.43$\times$ concurrency boost, establishing a new paradigm for cost-effective, high-throughput financial inference.

cs.CL

Nonlinear Unsteady Vortex-Lattice Vortex-Particle Method with Adaptive Wake Conversion for Rotorcraft Aerodynamics

Nonlinear unsteady vortex lattice-vortex particle methods (NL-UVLM-VPM) provide medium-fidelity predictions of rotorcraft aerodynamics with explicit three-dimensional wake representations at a moderate computational cost. This study presents an NL-UVLM-VPM approach with a scale-consistent adaptive wake panel-particle conversion strategy that mitigates the inherent temporal-spatial resolution coupling of conventional wake treatments in rotorcraft aerodynamic simulations. Numerical assessment shows that this strategy preserves the near third-order temporal convergence of the underlying time-integration scheme while improving robustness under coarsened temporal resolution. For a representative hover case, computational time is reduced by 29% relative to the conventional conversion strategy at identical temporal resolution and by nearly 70% compared with a fine-resolution reference simulation over 20 rotor revolutions, while maintaining thrust and torque predictions within 1% of the reference solution. Based on these analyzes, practical recommendations for particle conversion parameters and wake resolution are provided. The methodology is further validated for increasingly complex scenarios, including hover, forward flight with blade-vortex interaction, and multirotor interaction. Predictions show good agreement with experimental data and dedicated unsteady Reynolds-averaged Navier-Stokes simulations (URANS), while computational speedups exceeding two orders of magnitude relative to URANS are achieved.

physics.flu-dyn