SearcharxivSearch

arXiv subjects

Junlin Yu

Publications and source records attributed to Junlin Yu.

10 recordsLinked to original sources

HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds

We introduce HY-World 2.0, a multi-modal world model framework that advances our prior project HY-World 1.0. HY-World 2.0 accommodates diverse input modalities, including text prompts, single-view images, multi-view images, and videos, and produces 3D world representations. With text or single-view image inputs, the model performs world generation, synthesizing high-fidelity, navigable 3D Gaussian Splatting (3DGS) scenes. This is achieved through a four-stage method: a) Panorama Generation with HY-Pano 2.0, b) Trajectory Planning with WorldNav, c) World Expansion with WorldStereo 2.0, and d) World Composition with WorldMirror 2.0. Specifically, we introduce key innovations to enhance panorama fidelity, enable 3D scene understanding and planning, and upgrade WorldStereo, our keyframe-based view generation model with consistent memory. We also upgrade WorldMirror, a feed-forward model for universal 3D prediction, by refining model architecture and learning strategy, enabling world reconstruction from multi-view images or videos. Also, we introduce WorldLens, a high-performance 3DGS rendering platform featuring a flexible engine-agnostic architecture, automatic IBL lighting, efficient collision detection, and training-rendering co-design, enabling interactive exploration of 3D worlds with character support. Extensive experiments demonstrate that HY-World 2.0 achieves state-of-the-art performance on several benchmarks among open-source approaches, delivering results comparable to the closed-source model Marble. We release all model weights, code, and technical details to facilitate reproducibility and support further research on 3D world models.

cs.CV

Hunyuan3D-Omni: A Unified Framework for Controllable Generation of 3D Assets

Recent advances in 3D-native generative models have accelerated asset creation for games, film, and design. However, most methods still rely primarily on image or text conditioning and lack fine-grained, cross-modal controls, which limits controllability and practical adoption. To address this gap, we present Hunyuan3D-Omni, a unified framework for fine-grained, controllable 3D asset generation built on Hunyuan3D 2.1. In addition to images, Hunyuan3D-Omni accepts point clouds, voxels, bounding boxes, and skeletal pose priors as conditioning signals, enabling precise control over geometry, topology, and pose. Instead of separate heads for each modality, our model unifies all signals in a single cross-modal architecture. We train with a progressive, difficulty-aware sampling strategy that selects one control modality per example and biases sampling toward harder signals (e.g., skeletal pose) while downweighting easier ones (e.g., point clouds), encouraging robust multi-modal fusion and graceful handling of missing inputs. Experiments show that these additional controls improve generation accuracy, enable geometry-aware transformations, and increase robustness for production workflows.

cs.CV

Hunyuan3D Studio: End-to-End AI Pipeline for Game-Ready 3D Asset Generation

The creation of high-quality 3D assets, a cornerstone of modern game development, has long been characterized by labor-intensive and specialized workflows. This paper presents Hunyuan3D Studio, an end-to-end AI-powered content creation platform designed to revolutionize the game production pipeline by automating and streamlining the generation of game-ready 3D assets. At its core, Hunyuan3D Studio integrates a suite of advanced neural modules (such as Part-level 3D Generation, Polygon Generation, Semantic UV, etc.) into a cohesive and user-friendly system. This unified framework allows for the rapid transformation of a single concept image or textual description into a fully-realized, production-quality 3D model complete with optimized geometry and high-fidelity PBR textures. We demonstrate that assets generated by Hunyuan3D Studio are not only visually compelling but also adhere to the stringent technical requirements of contemporary game engines, significantly reducing iteration time and lowering the barrier to entry for 3D content creation. By providing a seamless bridge from creative intent to technical asset, Hunyuan3D Studio represents a significant leap forward for AI-assisted workflows in game development and interactive media.

cs.CV

HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels

Creating immersive and playable 3D worlds from texts or images remains a fundamental challenge in computer vision and graphics. Existing world generation approaches typically fall into two categories: video-based methods that offer rich diversity but lack 3D consistency and rendering efficiency, and 3D-based methods that provide geometric consistency but struggle with limited training data and memory-inefficient representations. To address these limitations, we present HunyuanWorld 1.0, a novel framework that combines the best of both worlds for generating immersive, explorable, and interactive 3D scenes from text and image conditions. Our approach features three key advantages: 1) 360{\deg} immersive experiences via panoramic world proxies; 2) mesh export capabilities for seamless compatibility with existing computer graphics pipelines; 3) disentangled object representations for augmented interactivity. The core of our framework is a semantically layered 3D mesh representation that leverages panoramic images as 360{\deg} world proxies for semantic-aware world decomposition and reconstruction, enabling the generation of diverse 3D worlds. Extensive experiments demonstrate that our method achieves state-of-the-art performance in generating coherent, explorable, and interactive 3D worlds while enabling versatile applications in virtual reality, physical simulation, game development, and interactive content creation.

cs.CV

Hunyuan3D 2.1: From Images to High-Fidelity 3D Assets with Production-Ready PBR Material

3D AI-generated content (AIGC) is a passionate field that has significantly accelerated the creation of 3D models in gaming, film, and design. Despite the development of several groundbreaking models that have revolutionized 3D generation, the field remains largely accessible only to researchers, developers, and designers due to the complexities involved in collecting, processing, and training 3D models. To address these challenges, we introduce Hunyuan3D 2.1 as a case study in this tutorial. This tutorial offers a comprehensive, step-by-step guide on processing 3D data, training a 3D generative model, and evaluating its performance using Hunyuan3D 2.1, an advanced system for producing high-resolution, textured 3D assets. The system comprises two core components: the Hunyuan3D-DiT for shape generation and the Hunyuan3D-Paint for texture synthesis. We will explore the entire workflow, including data preparation, model architecture, training strategies, evaluation metrics, and deployment. By the conclusion of this tutorial, you will have the knowledge to finetune or develop a robust 3D generative model suitable for applications in gaming, virtual reality, and industrial design.

cs.CV

Auto-Connect: Connectivity-Preserving RigFormer with Direct Preference Optimization

We introduce Auto-Connect, a novel approach for automatic rigging that explicitly preserves skeletal connectivity through a connectivity-preserving tokenization scheme. Unlike previous methods that predict bone positions represented as two joints or first predict points before determining connectivity, our method employs special tokens to define endpoints for each joint's children and for each hierarchical layer, effectively automating connectivity relationships. This approach significantly enhances topological accuracy by integrating connectivity information directly into the prediction framework. To further guarantee high-quality topology, we implement a topology-aware reward function that quantifies topological correctness, which is then utilized in a post-training phase through reward-guided Direct Preference Optimization. Additionally, we incorporate implicit geodesic features for latent top-k bone selection, which substantially improves skinning quality. By leveraging geodesic distance information within the model's latent space, our approach intelligently determines the most influential bones for each vertex, effectively mitigating common skinning artifacts. This combination of connectivity-preserving tokenization, reward-guided fine-tuning, and geodesic-aware bone selection enables our model to consistently generate more anatomically plausible skeletal structures with superior deformation properties.

cs.CV

Economics of Mobile Data Trading Market

To exploit users' heterogeneous data demands, several mobile network operators worldwide have launched the mobile data trading markets, where users can trade mobile data quota with each other. In this paper, we aim to understand the importance of data trading market (DTM) by studying the users' operator selection and trading decisions, and analyzing the operator's profit maximizing strategy. We model the interactions between the mobile operator and the users as a three-stage Stackelberg game. In Stage I, the operator chooses the operation fee imposed on sellers to maximize its profit. In Stage II, each user chooses his operator. In Stage III, each DTM user chooses his trading decisions. We derive the closed-form expression of the unique Nash equilibrium (NE) in Stages II and III, where every user proposes the same price such that the total demand matches with the total supply. We further show that the Stage I's problem is convex and compute the optimal operation fee. Our analysis and numerical results show that an operator with a small initial market share can increase its profit by proposing a DTM, which is in line with the real-world situation in Hong Kong.

cs.GT

A Novel Mobile Data Contract Design with Time Flexibility

In conventional mobile data plans, the data is associated with a fixed period (e.g., one month) and the unused data will be cleared at the end of each period. To take advantage of consumers' heterogeneous demands across different periods and meanwhile to provide more time flexibility, some mobile data service providers (SP) have offered data plans with different lengths of period. In this paper, we consider the data plan design problem for a single SP, who provides data plans with different lengths of period for consumers with different characteristics of data demands. We propose a contract-theoretic approach, wherein the SP offers a period-price data plan contract which consists of a set of period and price combinations, indicating the prices for data with different periods. We study the optimal data plan contract designs under two different models: discrete and continuous consumer-type models, depending on whether the consumer type is discrete or continuous. In the former model, each type of consumers are assigned with a specific period-price combination. In the latter model, the consumers are first categorized into a finite number of groups, and each group of consumers (possibly with different types) are assigned with a specific period-price combination. We systematically analyze the incentive compatibility (IC) constraint and individual rationality (IR) constraint, which ensure each consumer to choose the data plan with the period-price combination intended for his type. We further derive the optimal contract that maximizes the SP's expected profit, meanwhile satisfying the IC and IR constraints of consumers. Our numerical results show that the proposed optimal contract can increase the SP's profit by 35%, comparing with the conventional fixed monthly-period data plan.

cs.GT

Mobile Data Trading: Behavioral Economics Analysis and Algorithm Design

Motivated by the recently launched mobile data trading markets (e.g., China Mobile Hong Kong's 2nd exChange Market), in this paper we study the mobile data trading problem under the future data demand uncertainty. We introduce a brokerage-based market, where sellers and buyers propose their selling and buying quantities, respectively, to the trading platform that matches the market supply and demand. To understand the users' realistic trading behaviors, a prospect theory (PT) model from behavioral economics is proposed, which includes the widely adopted expected utility theory (EUT) as a special case. Although the PT modeling leads to a challenging non-convex optimization problem, the optimal solution can be characterized by exploiting the unimodal structure of the objective function. Building upon our analysis, we design an algorithm to help estimate the user's risk preference and provide trading recommendations dynamically, considering the latest market and usage information. It is shown in our simulation that the risk preferences have a significant impact on the user's decision and outcome: a risk-averse dominant user can guarantee a higher minimum profit in the trading, while a risk-seeking dominant user can achieve a higher maximum profit. By comparing with the EUT benchmark, it is shown that a PT user with a low reference point is more willing to buy mobile data. Moreover, when the probability of high future data demand is low, a PT user is more willing to buy mobile data due to the probability distortion comparing with an EUT user.

cs.GT

Spectrum Investment under Uncertainty: A Behavioral Economics Perspective

In this paper, we study a virtual wireless operator's spectrum investment problem under spectrum supply uncertainty. To obtain enough spectrum resources to meet its customer demands, the virtual operator can either sense for the temporarily unused spectrum in a licensed band, or lease spectrum from a spectrum owner. Sensing is usually cheaper than leasing, but the amount of available spectrum obtained by sensing is uncertain due to the primary users' activities in the licensed band. Previous studies on spectrum investment problems mainly considered the expected profit maximization problem of a risk-neutral operator based on the expected utility theory (EUT). In reality, however, an operator's decision is influenced by not only the consideration of expected profit maximization, but also the level of its risk preference. To capture this tradeoff between these two considerations, we analyze the operator's optimal decision problem using the prospect theory from behavioral economics, which includes EUT as a special case. The sensing and leasing optimal problem under prospect theory is non-convex and challenging to solve. Nevertheless, by exploiting the unimodal structure of the problem, we are able to compute the unique global optimal solution. We show that comparing to an EUT operator, both the risk-averse and risk-seeking operator achieve a smaller expected profit. On the other hand, a risk-averse operator can guarantee a larger minimum possible profit, while a risk-seeking operator can achieve a larger maximum possible profit. Furthermore, the tradeoff between the expected profit and the minimum possible profit for a risk-averse operator is better when the sensing cost increases, while the tradeoff between the expected profit and the maximum possible profit for a risk-seeking operator is better when the sensing cost decreases.

cs.NI