SearcharxivSearch

arXiv subjects

Jin-Yan Hu

Publications and source records attributed to Jin-Yan Hu.

2 recordsLinked to original sources

Semantic-Aware Representation Learning via Conditional Transport for Multi-Label Image Classification

Multi-label image classification is a critical task in machine learning that aims to accurately assign multiple labels to a single image. While existing methods often utilize attention mechanisms or graph convolutional networks to model visual representations, their performance is still constrained by two critical limitations: the inability to learn discriminative semantic-aware features, and the lack of fine-grained alignment between visual representations and label embeddings. To tackle these issues in a unified framework, this paper proposes a novel approach named Semantic-aware representation learning via Conditional Transport for Multi-Label Image Classification (SCT). The proposed method introduces a semantic-related feature learning module that extracts discriminative label-specific features by emphasizing semantic relevance and interaction, along with a conditional transport-based alignment mechanism that enables precise visual-semantic alignment. Extensive experiments on two widely-used benchmark datasets, VOC2007 and MS-COCO, validate the effectiveness of SCT and demonstrate its superior performance compared to existing state-of-the-art methods.

cs.CV

Incomplete Multi-Label Image Recognition by Co-learning Semantic-Aware Features and Label Recovery

Multi-label image recognition with incomplete labels is a challenging yet vital task in computer vision, which faces two fundamental challenges: learning semantic-aware features and recovering missing labels. In this paper, we propose a Co-learning framework for Semantic-aware features and Label recovery (CSL), designed to address both challenges in a unified learning paradigm. Specifically, we develop a semantic-related feature learning module that captures robust semantic-related representations by discovering semantic information and label correlations. Furthermore, a semantic-guided feature enhancement module is introduced to generate highly discriminative semantic-aware features by effectively aligning visual and semantic spaces. Finally, we present a collaborative learning framework that integrates semantic-aware feature learning with label recovery. This framework not only dynamically enhances the discriminability of semantic-aware features but also adaptively infers and recovers missing labels, thereby forming a mutually reinforcing mechanism between the two processes. Extensive experiments on three widely used public datasets (MS-COCO, VOC2007, and NUS-WIDE) demonstrate that CSL outperforms state-of-the-art methods for incomplete multi-label image recognition.

cs.CV