SearcharxivSearch

arXiv subjects

Shuhui Yang

Publications and source records attributed to Shuhui Yang.

At least 19 recordsLinked to original sources

ROAR-3D: Routing Arbitrary Views for High-Fidelity 3D Generation

Single-image-to-3D generative models can now produce high-quality geometry, yet conditioning on a single view inevitably introduces ambiguity about unseen regions. Multi-view conditioning can reduce this ambiguity, but existing methods either require fixed canonical viewpoints or rely on external reconstruction modules that impose heavy training costs and limit generation quality. We observe that pretrained single-view models already possess strong 2D-to-3D grounding that can be reused for multi-view conditioning. However, a closer analysis reveals that their conditioning mechanism entangles orientation control with geometry transfer, two functions that conflict when images from different viewpoints are naively combined. Based on this analysis, we propose ROAR-3D, a lightweight method that upgrades a pretrained single-view model to accept an arbitrary number of unposed images. A token-wise view router assigns each 3D latent token to its most relevant view, implicitly establishing 2D-to-3D correspondences without explicit pose input. A dual-stream attention design preserves the pretrained primary-view behavior while routing auxiliary views through a separate path dedicated to geometric enrichment. An orientation perturbation strategy ensures the auxiliary path learns orientation-independent geometry transfer. These components introduce minimal trainable parameters and add negligible inference overhead relative to the single-view baseline. ROAR-3D achieves state-of-the-art multi-view 3D generation quality and supports test-time view scaling from 1 to 12+ views with consistent improvements.

cs.CV

Tango3D: Towards Alignment for Global and Local 2D-3D Correspondence

Existing 3D foundation models typically align point clouds to frozen vision-language spaces like CLIP, which achieve strong cross-modal retrieval by compressing 3D shape into a global vector. However, this global-only alignment cannot establish fine-grained pixel-to-point correspondence. To solve this, we present Tango3D, a foundation model that unifies dense correspondence and global retrieval. We use a geometry-aware 2D visual backbone and a pretrained 3D VAE to encode images into 2D patches and point clouds into 3D tokens. These are mapped into a single shared space to achieve both local pixel-to-point alignment and global semantic alignment. To stabilize the joint learning of dense and global objectives, we introduce a three-stage progressive training strategy. Experiments show our model successfully achieves object-level pixel-to-point alignment while maintaining competitive global retrieval, a joint capability not offered by existing 3D foundation models. By establishing a fine-grained alignment feature space, Tango3D injects rich semantics into purely geometric 3D tokens, paving the way for a wide range of dense 3D downstream tasks.

cs.CV

Limit Properties at Critical Indices of Linear Canonical Riesz Potentials and Their Applications to Security of Multi-Image Encryption

In this article we introduce the linear canonical Riesz potential (for short, LCRP) and give its symbol in terms of linear canonical transforms. Driven by image processing, we establish the convergence/divergence of these LCRPs for different kinds of functions. Concretely, for grating functions, we prove that their classical Riesz potentials diverge, whereas their LCRP converge due to the key role of chirp functions. For the characteristic function ${\mathbf 1}_P$ of a convex polygon $P$, we show that the limit of its Riesz potential at any non-boundary point $\boldsymbol{x}$ equals ${\mathbf 1}_P(\boldsymbol{x})$, but its limit at the boundaries differ from ${\mathbf 1}_P$, while it is known that, for any Schwartz function $f$, the limit of its Riesz potential at any point $\boldsymbol{x}$ always equals $f(\boldsymbol{x})$. Based on these and the inverse operator of the LCRP (namely the linear canonical Laplacian operator), we propose an asymmetric cascaded LCRP method for the multi-image encryption and create an efficient and secure cryptosystem. Systematic security evaluations, including sensitivity, statistical, noise attack, and occlusion attack analyses, demonstrate its robustness and its security. Even for a single image, the proposed method is more efficient than the known encryption approach based on the fractional Riesz potential. The novelty of these results lies in that the convergence and the divergence of LCRTs at the critical indices, respectively, for ``good" Schwartz functions and for ``bad" discrete image functions essentially affect the security of image encryption and decryption.

cs.CR

MatPedia: A Universal Generative Foundation for High-Fidelity Material Synthesis

Physically-based rendering (PBR) materials are fundamental to photorealistic graphics, yet their creation remains labor-intensive and requires specialized expertise. While generative models have advanced material synthesis, existing methods lack a unified representation bridging natural image appearance and PBR properties, leading to fragmented task-specific pipelines and inability to leverage large-scale RGB image data. We present MatPedia, a foundation model built upon a novel joint RGB-PBR representation that compactly encodes materials into two interdependent latents: one for RGB appearance and one for the four PBR maps encoding complementary physical properties. By formulating them as a 5-frame sequence and employing video diffusion architectures, MatPedia naturally captures their correlations while transferring visual priors from RGB generation models. This joint representation enables a unified framework handling multiple material tasks--text-to-material generation, image-to-material generation, and intrinsic decomposition--within a single architecture. Trained on MatHybrid-410K, a mixed corpus combining PBR datasets with large-scale RGB images, MatPedia achieves native $1024\times1024$ synthesis that substantially surpasses existing approaches in both quality and diversity.

cs.CV

Hunyuan3D Studio: End-to-End AI Pipeline for Game-Ready 3D Asset Generation

The creation of high-quality 3D assets, a cornerstone of modern game development, has long been characterized by labor-intensive and specialized workflows. This paper presents Hunyuan3D Studio, an end-to-end AI-powered content creation platform designed to revolutionize the game production pipeline by automating and streamlining the generation of game-ready 3D assets. At its core, Hunyuan3D Studio integrates a suite of advanced neural modules (such as Part-level 3D Generation, Polygon Generation, Semantic UV, etc.) into a cohesive and user-friendly system. This unified framework allows for the rapid transformation of a single concept image or textual description into a fully-realized, production-quality 3D model complete with optimized geometry and high-fidelity PBR textures. We demonstrate that assets generated by Hunyuan3D Studio are not only visually compelling but also adhere to the stringent technical requirements of contemporary game engines, significantly reducing iteration time and lowering the barrier to entry for 3D content creation. By providing a seamless bridge from creative intent to technical asset, Hunyuan3D Studio represents a significant leap forward for AI-assisted workflows in game development and interactive media.

cs.CV

Hunyuan3D 2.5: Towards High-Fidelity 3D Assets Generation with Ultimate Details

In this report, we present Hunyuan3D 2.5, a robust suite of 3D diffusion models aimed at generating high-fidelity and detailed textured 3D assets. Hunyuan3D 2.5 follows two-stages pipeline of its previous version Hunyuan3D 2.0, while demonstrating substantial advancements in both shape and texture generation. In terms of shape generation, we introduce a new shape foundation model -- LATTICE, which is trained with scaled high-quality datasets, model-size, and compute. Our largest model reaches 10B parameters and generates sharp and detailed 3D shape with precise image-3D following while keeping mesh surface clean and smooth, significantly closing the gap between generated and handcrafted 3D shapes. In terms of texture generation, it is upgraded with phyiscal-based rendering (PBR) via a novel multi-view architecture extended from Hunyuan3D 2.0 Paint model. Our extensive evaluation shows that Hunyuan3D 2.5 significantly outperforms previous methods in both shape and end-to-end texture generation.

cs.CV

Hunyuan3D 2.1: From Images to High-Fidelity 3D Assets with Production-Ready PBR Material

3D AI-generated content (AIGC) is a passionate field that has significantly accelerated the creation of 3D models in gaming, film, and design. Despite the development of several groundbreaking models that have revolutionized 3D generation, the field remains largely accessible only to researchers, developers, and designers due to the complexities involved in collecting, processing, and training 3D models. To address these challenges, we introduce Hunyuan3D 2.1 as a case study in this tutorial. This tutorial offers a comprehensive, step-by-step guide on processing 3D data, training a 3D generative model, and evaluating its performance using Hunyuan3D 2.1, an advanced system for producing high-resolution, textured 3D assets. The system comprises two core components: the Hunyuan3D-DiT for shape generation and the Hunyuan3D-Paint for texture synthesis. We will explore the entire workflow, including data preparation, model architecture, training strategies, evaluation metrics, and deployment. By the conclusion of this tutorial, you will have the knowledge to finetune or develop a robust 3D generative model suitable for applications in gaming, virtual reality, and industrial design.

cs.CV

RomanTex: Decoupling 3D-aware Rotary Positional Embedded Multi-Attention Network for Texture Synthesis

Painting textures for existing geometries is a critical yet labor-intensive process in 3D asset generation. Recent advancements in text-to-image (T2I) models have led to significant progress in texture generation. Most existing research approaches this task by first generating images in 2D spaces using image diffusion models, followed by a texture baking process to achieve UV texture. However, these methods often struggle to produce high-quality textures due to inconsistencies among the generated multi-view images, resulting in seams and ghosting artifacts. In contrast, 3D-based texture synthesis methods aim to address these inconsistencies, but they often neglect 2D diffusion model priors, making them challenging to apply to real-world objects To overcome these limitations, we propose RomanTex, a multiview-based texture generation framework that integrates a multi-attention network with an underlying 3D representation, facilitated by our novel 3D-aware Rotary Positional Embedding. Additionally, we incorporate a decoupling characteristic in the multi-attention block to enhance the model's robustness in image-to-texture task, enabling semantically-correct back-view synthesis. Furthermore, we introduce a geometry-related Classifier-Free Guidance (CFG) mechanism to further improve the alignment with both geometries and images. Quantitative and qualitative evaluations, along with comprehensive user studies, demonstrate that our method achieves state-of-the-art results in texture quality and consistency.

cs.CV

Probing Peptide Adsorption Kinetics and Regioselectivity via Multipolar Plasmonic Modes of Gold Resonators

Efficient peptide adsorption on metasurfaces is essential for advanced biosensing applications. In this study, we demonstrate how ellipsometric measurements coupled with numerical simulations allow for real-time tracking of temporin-SHa peptide adsorption on gold metasurfaces. By characterizing spectral shifts at 660 nm, 920 nm, and 1000 nm, we reveal a rapid saturation of surface coverage after 3.5 hours, with a significant preferential adsorption at the resonator ends. Our approach provides a novel methodology for monitoring peptide binding, which could be applied to a wide range of biosensor designs.

physics.optics

Refining Image Edge Detection via Linear Canonical Riesz Transforms

Combining the linear canonical transform and the Riesz transform, we introduce the linear canonical Riesz transform (for short, LCRT), which is further proved to be a linear canonical multiplier. Using this LCRT multiplier, we conduct numerical simulations on images. Notably, the LCRT multiplier significantly reduces the complexity of the algorithm. Based on these we introduce the new concept of the sharpness $R^{\rm E}_{\rm sc}$ of the edge strength and continuity of images associated with the LCRT and, using it, we propose a new LCRT image edge detection method (for short, LCRT-IED method) and provide its mathematical foundation. Our experiments indicate that this sharpness $R^{\rm E}_{\rm sc}$ characterizes the macroscopic trend of edge variations of the image under consideration, while this new LCRT-IED method not only controls the overall edge strength and continuity of the image, but also excels in feature extraction in some local regions. These highlight the fundamental differences between the LCRT and the Riesz transform, which are precisely due to the multiparameter of the former. This new LCRT-IED method might be of significant importance for image feature extraction, image matching, and image refinement.

math.FA

MaterialMVP: Illumination-Invariant Material Generation via Multi-view PBR Diffusion

Physically-based rendering (PBR) has become a cornerstone in modern computer graphics, enabling realistic material representation and lighting interactions in 3D scenes. In this paper, we present MaterialMVP, a novel end-to-end model for generating PBR textures from 3D meshes and image prompts, addressing key challenges in multi-view material synthesis. Our approach leverages Reference Attention to extract and encode informative latent from the input reference images, enabling intuitive and controllable texture generation. We also introduce a Consistency-Regularized Training strategy to enforce stability across varying viewpoints and illumination conditions, ensuring illumination-invariant and geometrically consistent results. Additionally, we propose Dual-Channel Material Generation, which separately optimizes albedo and metallic-roughness (MR) textures while maintaining precise spatial alignment with the input images through Multi-Channel Aligned Attention. Learnable material embeddings are further integrated to capture the distinct properties of albedo and MR. Experimental results demonstrate that our model generates PBR textures with realistic behavior across diverse lighting scenarios, outperforming existing methods in both consistency and quality for scalable 3D asset creation.

cs.CV

Hunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generation

We present Hunyuan3D 2.0, an advanced large-scale 3D synthesis system for generating high-resolution textured 3D assets. This system includes two foundation components: a large-scale shape generation model -- Hunyuan3D-DiT, and a large-scale texture synthesis model -- Hunyuan3D-Paint. The shape generative model, built on a scalable flow-based diffusion transformer, aims to create geometry that properly aligns with a given condition image, laying a solid foundation for downstream applications. The texture synthesis model, benefiting from strong geometric and diffusion priors, produces high-resolution and vibrant texture maps for either generated or hand-crafted meshes. Furthermore, we build Hunyuan3D-Studio -- a versatile, user-friendly production platform that simplifies the re-creation process of 3D assets. It allows both professional and amateur users to manipulate or even animate their meshes efficiently. We systematically evaluate our models, showing that Hunyuan3D 2.0 outperforms previous state-of-the-art models, including the open-source models and closed-source models in geometry details, condition alignment, texture quality, and etc. Hunyuan3D 2.0 is publicly released in order to fill the gaps in the open-source 3D community for large-scale foundation generative models. The code and pre-trained weights of our models are available at: https://github.com/Tencent/Hunyuan3D-2

cs.CV

Multilinear Strongly Singular Integral Operators with Generalized Kernels on RD-Spaces

In this article, we introduce a class of multilinear strongly singular integral operators with generalized kernels on the RD-space. The boundedness of these operators on weighted Lebesgue spaces is established. Moreover, two types of endpoint estimates and their boundedness on generalized weighted Morrey spaces are obtained. Our results further generalize the relevant conclusions on generalized kernels in Euclidean spaces. Moreover, the weak-type results on weighted Lebesgue spaces are brand new even in the situation of Euclidean spaces. In addition, when the generalized kernels degenerate into classical kernels, our research results also extend the relevant known results. It is worth mentioning that our RD spaces are more general than theirs.

math.FA

Multilinear Fractional Integral Operators with Generalized Kernels

In this article, we introduce a class of multilinear fractional integral operators with generalized kernels that are weaker than the Dini kernel condition. We establish the boundedness of multilinear fractional integral operators with generalized kernels on weighted Lebesgue spaces and variable exponent Lebesgue spaces, as well as the boundedness of multilinear commutators generated by multilinear fractional integral operators with generalized kernels and $BMO$ functions. Even when the generalized kernels condition goes back to the Dini kernel condition, the conclusions on the commutators remain new.

math.FA

New Multilinear Littlewood--Paley $g_{\lambda}^{*}$ Function and Commutator on weighted Lebesgue Spaces

Via the new weight function $A_{\vec p}^{\theta }(\varphi )$, the authors introduce a new class of multilinear Littlewood--Paley $g_{\lambda}^{*}$ functions and establish the boundedness on weighted Lebesgue spaces. In addition, the authors obtain the boundedness of the multilinear commutator and multilinear iterated commutator generated by the multilinear Littlewood--Paley $g_{\lambda}^{*}$ function and the new $BMO$ function on weighted Lebesgue spaces. The results in this article include the known results in \cite{XY2015,SXY2014}. When $m=1$, that is, in the case of one-linear, our conclusions are also new, further extending the results in \cite{S1961}.

math.FA

Multilinear Commutators of Multilinear Square Operators Associated with New $BMO$ Functions and New Weight Functions

Via the new weight $A_{\vec p}^{\infty}(\varphi)$ and the new $BMO$ function, the authors introduce a new class of multilinear square operators $T$ with generalized kernels. The boundedness of multilinear commutators and multilinear iterative commutators generated by $T$ and the new $BMO$ function on weighted Lebesgue spaces and weighted Morrey spaces is obtained, respectively. The results of this article contain some known conclusions.

math.FA

Multilinear Square Operators Meet New Weight Functions

Via the new weight $A_{\vec p}^{\theta }(\varphi )$, the authors introduce a new class of multilinear square operators. The boundedness on the weighted Lebesgue space and the weighted Morrey space is obtained, respectively. Our results include the known results of the standard multilinear square operator and the weight $A_{\vec p}$. Moreover, the results in this article seem to be new even for one-linear case.

math.FA

Fractional Fourier Transforms Meet Riesz Potentials and Image Processing

Via chirp functions from fractional Fourier transforms, the authors introduce fractional Riesz potentials related to chirp functions, establish their relations with fractional Fourier transforms, fractional Laplace operators related to chirp functions, and fractional Riesz transforms related to chirp functions, and obtain their boundedness on rotation invariant spaces related to chirp functions. Finally, the authors give the numerical image simulation of fractional Riesz potentials related to chirp functions and their applications in image processing. The main novelty of this article is to propose a new image encryption method for the double phase coding based on the fractional Riesz potential related to chirp functions. The symbol of fractional Riesz potentials related to chirp functions essentially provides greater degrees of freedom and greatly makes the information more secure.

math.FA