SearcharxivSearch

arXiv subjects

Hui Xiang

Publications and source records attributed to Hui Xiang.

8 recordsLinked to original sources

KD-NVC: A Search-and-Distill Framework to Accelerate Neural Video Coding

While neural video coding (NVC) has achieved remarkable rate-distortion performance, real-time decoding on edge devices has become an important demand but remains limited by high complexity. Knowledge distillation (KD) is widely used for model acceleration, yet its application to NVC faces critical challenges. Specifically, the heterogeneity of NVC sub-modules renders uniform architectural reduction suboptimal, necessitating a per-module design for better rate-distortion-speed trade-off. However, searching for diverse architectures via existing neural architecture search (NAS) algorithms is unaffordable due to the expensive training cost of neural video codecs. Moreover, after the lightweight architecture is determined, existing distillation methods overlook the feature-energy sparsity induced by the rate-constraint, which is essential for maintaining compression performance. To address these issues, we propose a two-stage distillation framework KD-NVC. In the first stage, we introduce an acceleration-efficiency-based neural architecture search (AE-NAS) algorithm. It explores the module-wise Pareto frontier to adaptively allocate the acceleration budget across heterogeneous modules. Also, it introduces the acceleration-efficiency metric to determine the final student architecture without practically training all architecture-level candidates. In the second stage, we design an energy-aware feature distillation (EFD) loss that aligns the spatially-aggregated feature-energy signatures between the teacher and student codecs, transferring the rate-induced sparsity patterns for better compression efficiency. Experimental results demonstrate that the proposed framework consistently outperforms existing codec-oriented distillation methods, and achieves 69 FPS decoding at 1080p on RTX 5060 while maintaining comparable RD performance to VTM-LDB.

eess.IV

Recent Advances of End-to-End Video Coding Technologies for AVS Standard Development

Video coding standards are essential to enable the interoperability and widespread adoption of efficient video compression technologies. In pursuit of greater video compression efficiency, the AVS video coding working group launched the standardization exploration of end-to-end intelligent video coding, establishing the AVS End-to-End Intelligent Video Coding Exploration Model (AVS-EEM) project. A core design principle of AVS-EEM is its focus on practical deployment, featuring inherently low computational complexity and requiring strict adherence to the common test conditions of conventional video coding. This paper details the development history of AVS-EEM and provides a systematic introduction to its key technical framework, covering model architectures, training strategies, and inference optimizations. These innovations have collectively driven the project's rapid performance evolution, enabling continuous and significant gains under strict complexity constraints. Through over two years of iterative refinement and collaborative effort, the coding performance of AVS-EEM has seen substantial improvement. Experimental results demonstrate that its latest model achieves superior compression efficiency compared to the conventional AVS3 reference software, marking a significant step toward a deployable intelligent video coding standard.

eess.IV

Real-Time Neural Video Compression with Unified Intra and Inter Coding

Neural video compression (NVC) technologies have advanced rapidly in recent years, yielding state-of-the-art schemes such as DCVC-RT that offer superior compression efficiency to H.266/VVC and real-time encoding/decoding capabilities. Nonetheless, existing NVC schemes have several limitations, including inefficiency in dealing with disocclusion and new content, interframe error propagation and accumulation, among others. To eliminate these limitations, we borrow the idea from classic video coding schemes, which allow intra coding within inter-coded frames. With the intra coding tool enabled, disocclusion and new content are properly handled, and interframe error propagation is naturally intercepted without the need for manual refresh mechanisms. We present an NVC framework with unified intra and inter coding, where every frame is processed by a single model that is trained to perform intra/inter coding adaptively. Moreover, we propose a simultaneous two-frame compression design to exploit interframe redundancy not only forwardly but also backwardly. Experimental results show that our scheme outperforms DCVC-RT by an average of 12.1% BD-rate reduction, delivers more stable bitrate and quality per frame, and retains real-time encoding/decoding performances. Code and models will be released.

cs.CV

PromptAL: Sample-Aware Dynamic Soft Prompts for Few-Shot Active Learning

Active learning (AL) aims to optimize model training and reduce annotation costs by selecting the most informative samples for labeling. Typically, AL methods rely on the empirical distribution of labeled data to define the decision boundary and perform uncertainty or diversity estimation, subsequently identifying potential high-quality samples. In few-shot scenarios, the empirical distribution often diverges significantly from the target distribution, causing the decision boundary to shift away from its optimal position. However, existing methods overlook the role of unlabeled samples in enhancing the empirical distribution to better align with the target distribution, resulting in a suboptimal decision boundary and the selection of samples that inadequately represent the target distribution. To address this, we propose a hybrid AL framework, termed \textbf{PromptAL} (Sample-Aware Dynamic Soft \textbf{Prompts} for Few-Shot \textbf{A}ctive \textbf{L}earning). This framework accounts for the contribution of each unlabeled data point in aligning the current empirical distribution with the target distribution, thereby optimizing the decision boundary. Specifically, PromptAL first leverages unlabeled data to construct sample-aware dynamic soft prompts that adjust the model's predictive distribution and decision boundary. Subsequently, based on the adjusted decision boundary, it integrates uncertainty estimation with both global and local diversity to select high-quality samples that more accurately represent the target distribution. Experimental results on six in-domain and three out-of-domain datasets show that PromptAL achieves superior performance over nine baselines. Our codebase is openly accessible.

cs.CL

Physics-Informed Neural Networks for the Korteweg-de Vries Equation for Internal Solitary Wave Problem: Forward Simulation and Inverse Parameter Estimation

Physics-informed neural networks (PINNs) have emerged as a transformative framework for addressing operator learning and inverse problems involving the Korteweg-de Vries (KdV) equation for internal solitary waves. By integrating physical constraints with data-driven optimization, PINNs overcome the critical challenges of parameter unmeasurability in the KdV equation for internal solitary waves in two-layer fluid systems. This work addresses two problems: (1) Operator learning constructs a mapping from parameters to solutions, enabling wave evolution predictions from unknown parameters. Comparative studies demonstrate prediction errors as low as $10^{-4}$ when using 1000 training points. (2) Inverse problem solving leverages sparse and potentially noisy observational data with physics-regularized constraints to invert nonlinear coefficients successfully. Compared to conventional approaches, this end-to-end differentiable paradigm unifies operator learning and inverse problem-solving while overcoming mesh discretization errors and high-dimensional parameter space iteration costs. The method shows effectiveness for internal wave problems in stratified fluids, providing both accurate forward modeling and robust parameter inversion capabilities, even under noise.

physics.flu-dyn

AeroDiT: Diffusion Transformers for Reynolds-Averaged Navier-Stokes Simulations of Airfoil Flows

Real-time and accurate prediction of aerodynamic flow fields around airfoils is crucial for flow control and aerodynamic optimization. However, achieving this remains challenging due to the high computational costs and the non-linear nature of flow physics. Traditional Computational Fluid Dynamics (CFD) methods face limitations in balancing computational efficiency and accuracy, hindering their application in real-time scenarios. To address these challenges, this study presents AeroDiT, a novel surrogate model that integrates scalable diffusion models with transformer architectures to address these challenges. Trained on Reynolds-Averaged Navier-Stokes (RANS) simulation data for high Reynolds-number airfoil flows, AeroDiT accurately captures complex flow patterns while enabling real-time predictions. The model demonstrates impressive performance, with mean relative $L_2$ errors of 0.1, 0.025, and 0.050 for pressure $p$ and velocity components $u_x, u_y$, confirming its reliability. To further enhance physical consistency, we incorporate explicit physics-informed losses based on RANS residuals, including mass and momentum conservation constraints. The transformer-based structure allows for real-time predictions within seconds, enabling efficient aerodynamic simulations. This work underscores the potential of generative machine learning techniques to advance computational fluid dynamics, offering potential solutions to challenges in simulating high-fidelity aerodynamic flows.

physics.flu-dyn

Solution multiplicity and effects of data and eddy viscosity on Navier-Stokes solutions inferred by physics-informed neural networks

Physics-informed neural networks (PINNs) have emerged as a new simulation paradigm for fluid flows and are especially effective for inverse and hybrid problems. However, vanilla PINNs often fail in forward problems, especially at high Reynolds (Re) number flows. Herein, we study systematically the classical lid-driven cavity flow at $Re=2,000$, $3,000$ and $5,000$. We observe that vanilla PINNs obtain two classes of solutions, one class that agrees with direct numerical simulations (DNS), and another that is an unstable solution to the Navier-Stokes equations and not physically realizable. We attribute this solution multiplicity to singularities and unbounded vorticity, and we propose regularization methods that restore a unique solution within 1\% difference from the DNS solution. In particular, we introduce a parameterized entropy-viscosity method as artificial eddy viscosity and identify suitable parameters that drive the PINNs solution towards the DNS solution. Furthermore, we solve the inverse problem by subsampling the DNS solution, and identify a new eddy viscosity distribution that leads to velocity and pressure fields almost identical to their DNS counterparts. Surprisingly, a single measurement at a random point suffices to obtain a unique PINNs DNS-like solution even without artificial viscosity, which suggests possible pathways in simulating high Reynolds number turbulent flows using vanilla PINNs.

physics.flu-dyn

How to Control Hydrodynamic Force on Fluidic Pinball via Deep Reinforcement Learning

Deep reinforcement learning (DRL) for fluidic pinball, three individually rotating cylinders in the uniform flow arranged in an equilaterally triangular configuration, can learn the efficient flow control strategies due to the validity of self-learning and data-driven state estimation for complex fluid dynamic problems. In this work, we present a DRL-based real-time feedback strategy to control the hydrodynamic force on fluidic pinball, i.e., force extremum and tracking, from cylinders' rotation. By adequately designing reward functions and encoding historical observations, and after automatic learning of thousands of iterations, the DRL-based control was shown to make reasonable and valid control decisions in nonparametric control parameter space, which is comparable to and even better than the optimal policy found through lengthy brute-force searching. Subsequently, one of these results was analyzed by a machine learning model that enabled us to shed light on the basis of decision-making and physical mechanisms of the force tracking process. The finding from this work can control hydrodynamic force on the operation of fluidic pinball system and potentially pave the way for exploring efficient active flow control strategies in other complex fluid dynamic problems.

eess.SY