SearcharxivSearch

arXiv subjects

Yongdong Huang

Publications and source records attributed to Yongdong Huang.

4 recordsLinked to original sources

Fusion in Your Way: Aligning Image Fusion with Heterogeneous Demands via Direct Preference Optimization

As a key technique in multi-modal processing, infrared and visible image fusion (IVIF) plays a crucial role in integrating complementary spectral information for visual enhancement and downstream vision tasks. Despite remarkable progress, existing methods struggle to flexibly accommodate heterogeneous demands. Achieving adaptive fusion that aligns with various preferences from both human and machine vision remains an open and challenging problem. To address this challenge, we propose DPOFusion, a direct preference optimization (DPO) framework integrating the property-aligned latent diffusion model (PALDM) and the preference-controllable latent diffusion model (PCLDM), enabling task-guided, preference-adaptive IVIF for both human and machine vision. The PALDM leverages a latent fusion prior and a joint conditional loss to generate diverse candidate fusion results with various properties. PCLDM is subsequently fine-tuned via instance direct preference optimization (IDPO), enabling direct control of the final fusion results with heterogeneous preference signals. Experimental results demonstrate that our framework not only attains precise preference alignment among humans, vision-language models, and task-driven networks, but also sets a new benchmark for adaptive fusion quality and task-oriented transferability.

cs.CV

An Efficient End-to-End 3D Voxel Reconstruction based on Neural Architecture Search

Using neural networks to represent 3D objects has become popular. However, many previous works employ neural networks with fixed architecture and size to represent different 3D objects, which lead to excessive network parameters for simple objects and limited reconstruction accuracy for complex objects. For each 3D model, it is desirable to have an end-to-end neural network with as few parameters as possible to achieve high-fidelity reconstruction. In this paper, we propose an efficient voxel reconstruction method utilizing neural architecture search (NAS) and binary classification. Taking the number of layers, the number of nodes in each layer, and the activation function of each layer as the search space, a specific network architecture can be obtained based on reinforcement learning technology. Furthermore, to get rid of the traditional surface reconstruction algorithms (e.g., marching cube) used after network inference, we complete the end-to-end network by classifying binary voxels. Compared to other signed distance field (SDF) prediction or binary classification networks, our method achieves significantly higher reconstruction accuracy using fewer network parameters.

cs.CV

Embeddings from noncompact symmetric spaces to their compact duals

Every compact symmetric space $M$ admits a dual noncompact symmetric space $\check{M}$. When $M$ is a generalized Grassmannian, we can view $\check{M}$ as a open submanifold of it consisting of space-like subspaces \cite{HL}. Motivated from this, we study the embeddings from noncompact symmetric spaces to their compact duals, including space-like embedding for generalized Grassmannians, Borel embedding for Hermitian symmetric spaces and the generalized embedding for symmetric R-spaces. We will compare these embeddings and describe their images using cut loci.

math.AG

On equivariant quantum Schubert calculus for G/P

We show a Z^2-filtered algebraic structure and a "quantum to classical" principle on the torus-equivariant quantum cohomology of a complete flag variety of general Lie type, generalizing earlier works of Leung and the second author. We also provide various applications on equivariant quantum Schubert calculus, including an equivariant quantum Pieri rule for any partial flag variety of Lie type A.

math.AG