Searcharxiv⌕ Search

arXiv subjects

Khoa Pham-Dinh

Publications and source records attributed to Khoa Pham-Dinh.

2 recordsLinked to original sources

CASR: Content-Adaptive Neural Super-Resolution Post-Filter for Versatile Video Coding via Low-Rank Overfitting

The use of super-resolution as a post-processing step following a video codec allows videos to be encoded at reduced spatial resolution at the encoder side and to be reconstructed and upsampled to the original resolution at the decoder side. In this way, the required bitrate is reduced and the quality of the reconstructed frames is improved without modifying the core coding architecture. However, generic SR models are typically trained offline and lack sufficient adaptability to the diverse content characteristics and compression artifacts produced by video codecs, which limits their effectiveness during test time. To address the limitation, this paper proposes CASR (Content-Adaptive Super-Resolution), a content-adaptive SR post-filter framework for Versatile Video Coding (VVC), based on encoder-side overfitting on each input sequence. In order to limit the bitrate overhead required for signalling the content adaptation signal, i.e. the weight-update, Low-Rank Adaptation (LoRA) is leveraged. The method freezes the convolution kernels of a pretrained SR network and fine-tunes only lightweight rank-r matrices attached to selected convolution layers, using VVC decoded frames and quantization-parameter (QP) maps of test sequences as supervision. The resulting low-rank update is compressed with the MPEG Neural Network Compression and Representation (NNR) standard. Experiments on the JVET common test conditions (CTC) class A1 and A2 sequences indicate that LoRA-based content adaptation provides bitrate savings over a non-adapted SR post-filter at a small signalling cost. Compared with the VVC Test Model (VTM21), the proposed method achieves BD-rate savings of -10.93% (Y), -15.39% (U), -24.41% (V) under random access and -13.43% (Y), -5.75% (U), -22.94% (V) under all-intra. An ablation of the LoRA rank r further shows that r=4 provides the best trade-off between coding gain and signalling cost.

eess.IV↗

You've Seen Enough: Quality-Constrained Image Coding for Machines

Visual data is increasingly consumed by machine-vision systems rather than by human observers. Image Coding for Machines (ICM) compresses images assuming the main observer is a computer vision application and that the human observer needs to inspect or validate the decisions. Inspired by just-noticeable distortion, we cap human-observed quality at a desired level and devote the remaining bits to machine performance. Specifically, joint compression-segmentation training is recast as a constrained optimization problem in which the codec must meet a predefined acceptable target visual quality while a task term consumes the remaining coding capacity. This paper proposes two variants of a penalty function that guides the quality toward the target: an absolute function and a bilinear function, the latter applying a steeper slope once the target visual quality is exceeded. Experimental results show that, under the quality constraint, the proposed method achieves BD-rates of $-22.82\%$ and $-29.81\%$ relative to an unconstrained joint rate--distortion--task optimization and a simple rate--distortion baseline, respectively, showcasing bitrate reduction with the same task performance. This is achieved while the codec also meets the target visual quality with a reasonable error and without adding any complexity overhead.

cs.CV↗