SearcharxivSearch

arXiv subjects

Dongxu Wu

Publications and source records attributed to Dongxu Wu.

5 recordsLinked to original sources

Beyond Cosine Similarity: Magnitude-Aware CLIP for No-Reference Image Quality Assessment

Recent efforts have repurposed the Contrastive Language-Image Pre-training (CLIP) model for No-Reference Image Quality Assessment (NR-IQA) by measuring the cosine similarity between the image embedding and textual prompts such as "a good photo" or "a bad photo." However, this semantic similarity overlooks a critical yet underexplored cue: the magnitude of the CLIP image features, which we empirically find to exhibit a strong correlation with perceptual quality. In this work, we introduce a novel adaptive fusion framework that complements cosine similarity with a magnitude-aware quality cue. Specifically, we first extract the absolute CLIP image features and apply a Box-Cox transformation to statistically normalize the feature distribution and mitigate semantic sensitivity. The resulting scalar summary serves as a semantically-normalized auxiliary cue that complements cosine-based prompt matching. To integrate both cues effectively, we further design a confidence-guided fusion scheme that adaptively weighs each term according to its relative strength. Extensive experiments on multiple benchmark IQA datasets demonstrate that our method consistently outperforms standard CLIP-based IQA and state-of-the-art baselines, without any task-specific training.

cs.CV

Q-Doc: Benchmarking Document Image Quality Assessment Capabilities in Multi-modal Large Language Models

The rapid advancement of Multi-modal Large Language Models (MLLMs) has expanded their capabilities beyond high-level vision tasks. Nevertheless, their potential for Document Image Quality Assessment (DIQA) remains underexplored. To bridge this gap, we propose Q-Doc, a three-tiered evaluation framework for systematically probing DIQA capabilities of MLLMs at coarse, middle, and fine granularity levels. a) At the coarse level, we instruct MLLMs to assign quality scores to document images and analyze their correlation with Quality Annotations. b) At the middle level, we design distortion-type identification tasks, including single-choice and multi-choice tests for multi-distortion scenarios. c) At the fine level, we introduce distortion-severity assessment where MLLMs classify distortion intensity against human-annotated references. Our evaluation demonstrates that while MLLMs possess nascent DIQA abilities, they exhibit critical limitations: inconsistent scoring, distortion misidentification, and severity misjudgment. Significantly, we show that Chain-of-Thought (CoT) prompting substantially enhances performance across all levels. Our work provides a benchmark for DIQA capabilities in MLLMs, revealing pronounced deficiencies in their quality perception and promising pathways for enhancement. The benchmark and code are publicly available at: https://github.com/cydxf/Q-Doc.

cs.CV

Mitigating Perception Bias: A Training-Free Approach to Enhance LMM for Image Quality Assessment

Despite the impressive performance of large multimodal models (LMMs) in high-level visual tasks, their capacity for image quality assessment (IQA) remains limited. One main reason is that LMMs are primarily trained for high-level tasks (e.g., image captioning), emphasizing unified image semantics extraction under varied quality. Such semantic-aware yet quality-insensitive perception bias inevitably leads to a heavy reliance on image semantics when those LMMs are forced for quality rating. In this paper, instead of retraining or tuning an LMM costly, we propose a training-free debiasing framework, in which the image quality prediction is rectified by mitigating the bias caused by image semantics. Specifically, we first explore several semantic-preserving distortions that can significantly degrade image quality while maintaining identifiable semantics. By applying these specific distortions to the query or test images, we ensure that the degraded images are recognized as poor quality while their semantics mainly remain. During quality inference, both a query image and its corresponding degraded version are fed to the LMM along with a prompt indicating that the query image quality should be inferred under the condition that the degraded one is deemed poor quality. This prior condition effectively aligns the LMM's quality perception, as all degraded images are consistently rated as poor quality, regardless of their semantic variance. Finally, the quality scores of the query image inferred under different prior conditions (degraded versions) are aggregated using a conditional probability model. Extensive experiments on various IQA datasets show that our debiasing framework could consistently enhance the LMM performance.

cs.CV

Influence Factors of the Evaporation Rate of a Solar Steam Generation System: a Numerical Study

Many efforts have been dedicated to improve the solar steam generation by using a bi-layer structure. In this paper, a two-dimensional mathematical model describing the water evaporation in a bi-layer structure is firstly established and then the finite element method is used to simulate the effects of different influence factors on the evaporation rate. Results turn out that: besides the high solar energy absorptivity of the first-layer, an optimum porosity of the second-layer porous material should be applied and the optimum porosity is about 0.45 in this work. This optimum porosity is determined by the balance between the positive effect of the lowering effective thermal conductivity of the second layer and the negative effect of the reduced vapor diffusivity in the second layer when the porosity is decreased. The influence of the thermal conductivity of the second-layer porous material is negligible because the effective thermal conductivity of the second layer is determined by the porosity while a larger porosity means more water in the second layer. The ambient air velocity could greatly enhance the evaporation rate, and the evaporation rate will decrease linearly with the increase of the air relative humidity. This study is expected to supply some information for developing a more effective bi-layer solar steam generation system.

physics.app-ph

Surrounding Effects on the Evaporation Efficiency of a Bi-layered Structure in Solar Steam Generation: a Numerical Study

The bi-layered structure has drawn a wide interest due to its good performance in solar steam generation. In this work, we firstly develop a calculation model which could give a good prediction of experimental results. Then, this model is applied to numerically study the effects of the depth of the liquid water, the temperature of the ambient air, the temperature of the liquid water, the porosity and the thermal conductivity of the second-layer porous material on the evaporation efficiency. Results show that when the depth of the liquid water is large enough, the thermal insulation at the bottom of the liquid water is not needed. There is a linear dependence of the evaporation efficiency on the temperature of the ambient air or/and the temperature of the liquid water, and an equation has been given to describe this phenomenon in the text. Compared to the temperature of the ambient air, the temperature of the liquid water could have a much larger effect on the evaporation efficiency. The effective thermal conductivity of the second layer, which could impose important effect on the evaporation efficiency, mainly depends on the porosity rather than the thermal conductivity of the second-layer porous material. Thus, we do not need to take into consideration of the thermal conductivity when selecting second-layer materials. This study is expected to provide some information for designing a high-evaporation-performance bi-layered system.

physics.app-ph