SearcharxivSearch

arXiv subjects

Zhaoen Su

Publications and source records attributed to Zhaoen Su.

12 recordsLinked to original sources

OmniFaceRig: Fully Automatic Inner-Mouth-Aware Face Rigging Across Diverse 3D Character Topologies

Facial rigging - creating FACS-based blendshapes together with inner-mouth geometry (teeth, gums, and tongue) - remains a major bottleneck in 3D character production. Existing pipelines still require substantial designer effort, especially for manual landmark annotation, per-character template adjustment, and inner-mouth placement. We present OmniFaceRig, a fully automatic end-to-end pipeline that converts a static surface-only 3D character mesh, with no pre-modeled oral cavity, into an inner-mouth-aware FACS rig with up to 155 blendshapes, procedurally fitted teeth, gums, and tongue, and re-packed UV/texture. OmniFaceRig supports diverse topologies - humans, humanoids, long-muzzled animals (e.g., dogs, wolves, foxes), and short-muzzled animals (e.g., cats, bears, rabbits, tigers) - with no manual landmarks, no user-provided templates, and no per-asset setup. The pipeline combines hybrid VLM+CV riggability checking, multi-model face parsing, dense keypoint-driven template registration, procedural inner-mouth construction, and collision-aware blendshape transfer. For non-human characters, OmniFaceRig selects topology-specific face and inner-mouth templates and uses collision-aware inner-mouth fitting to reduce teeth-face intersections without exposing users to category-specific tuning. We also publicly release Omni-Bench, a freely available benchmark dataset of 1,000 biped 3D characters with FACS facial blendshapes and inner-mouth geometry, spanning humans, humanoids, cats, dogs, and other animals. Experiments show high final rigging success on screened Omni-Bench inputs, nearly complete face detection recall from the segmentation ensemble and reliable inner-mouth placement with low penetration. Together, OmniFaceRig provides an automatic path from static generated characters to animation-ready facial rigs across both human and non-human topologies.

cs.GR

GenLCA: 3D Diffusion for Full-Body Avatars from In-the-Wild Videos

We present GenLCA, a diffusion-based generative model for generating and editing photorealistic full-body avatars from text and image inputs. The generated avatars are faithful to the inputs, while supporting high-fidelity facial and full-body animations. The core idea is a novel paradigm that enables training a full-body 3D diffusion model from partially observable 2D data, allowing the training dataset to scale to millions of real-world videos. This scalability contributes to the superior photorealism and generalizability of GenLCA. Specifically, we scale up the dataset by repurposing a pretrained feed-forward avatar reconstruction model as an animatable 3D tokenizer, which encodes unstructured video frames into structured 3D tokens. However, most real-world videos only provide partial observations of body parts, resulting in excessive blurring or transparency artifacts in the 3D tokens. To address this, we propose a novel visibility-aware diffusion training strategy that replaces invalid regions with learnable tokens and computes losses only over valid regions. We then train a flow-based diffusion model on the token dataset, inherently maintaining the photorealism and animatability provided by the pretrained avatar reconstruction model. Our approach effectively enables the use of large-scale real-world video data to train a diffusion model natively in 3D. We demonstrate the efficacy of our method through diverse and high-fidelity generation and editing results, outperforming existing solutions by a large margin. The project page is available at https://onethousandwu.com/GenLCA-Page.

cs.CV

Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining

High-quality 3D avatar modeling faces a critical trade-off between fidelity and generalization. On the one hand, multi-view studio data enables high-fidelity modeling of humans with precise control over expressions and poses, but it struggles to generalize to real-world data due to limited scale and the domain gap between the studio environment and the real world. On the other hand, recent large-scale avatar models trained on millions of in-the-wild samples show promise for generalization across a wide range of identities, yet the resulting avatars are often of low-quality due to inherent 3D ambiguities. To address this, we present Large-Scale Codec Avatars (LCA), a high-fidelity, full-body 3D avatar model that generalizes to world-scale populations in a feedforward manner, enabling efficient inference. Inspired by the success of large language models and vision foundation models, we present, for the first time, a pre/post-training paradigm for 3D avatar modeling at scale: we pretrain on 1M in-the-wild videos to learn broad priors over appearance and geometry, then post-train on high-quality curated data to enhance expressivity and fidelity. LCA generalizes across hair styles, clothing, and demographics while providing precise, fine-grained facial expressions and finger-level articulation control, with strong identity preservation. Notably, we observe emergent generalization to relightability and loose garment support to unconstrained inputs, and zero-shot robustness to stylized imagery, despite the absence of direct supervision.

cs.CV

DuoMo: Dual Motion Diffusion for World-Space Human Reconstruction

We present DuoMo, a generative method that recovers human motion in world-space coordinates from unconstrained videos with noisy or incomplete observations. Reconstructing such motion requires solving a fundamental trade-off: generalizing from diverse and noisy video inputs while maintaining global motion consistency. Our approach addresses this problem by factorizing motion learning into two diffusion models. The camera-space model first estimates motion from videos in camera coordinates. The world-space model then lifts this initial estimate into world coordinates and refines it to be globally consistent. Together, the two models can reconstruct motion across diverse scenes and trajectories, even from highly noisy or incomplete observations. Moreover, our formulation is general, generating the motion of mesh vertices directly and bypassing parametric models. DuoMo achieves state-of-the-art performance. On EMDB, our method obtains a 16% reduction in world-space reconstruction error while maintaining low foot skating. On RICH, it obtains a 30% reduction in world-space error. Project page: https://yufu-wang.github.io/duomo/

cs.CV

Detection of Active Emergency Vehicles using Per-Frame CNNs and Output Smoothing

While inferring common actor states (such as position or velocity) is an important and well-explored task of the perception system aboard a self-driving vehicle (SDV), it may not always provide sufficient information to the SDV. This is especially true in the case of active emergency vehicles (EVs), where light-based signals also need to be captured to provide a full context. We consider this problem and propose a sequential methodology for the detection of active EVs, using an off-the-shelf CNN model operating at a frame level and a downstream smoother that accounts for the temporal aspect of flashing EV lights. We also explore model improvements through data augmentation and training with additional hard samples.

cs.CV

Convolutions for Spatial Interaction Modeling

In many different fields interactions between objects play a critical role in determining their behavior. Graph neural networks (GNNs) have emerged as a powerful tool for modeling interactions, although often at the cost of adding considerable complexity and latency. In this paper, we consider the problem of spatial interaction modeling in the context of predicting the motion of actors around autonomous vehicles, and investigate alternatives to GNNs. We revisit 2D convolutions and show that they can demonstrate comparable performance to graph networks in modeling spatial interactions with lower latency, thus providing an effective and efficient alternative in time-critical systems. Moreover, we propose a novel interaction loss to further improve the interaction modeling of the considered methods.

cs.CV

Triple Andreev dot chains in semiconductor nanowires

Kitaev chain is a theoretical model of a one-dimensional topological superconductor with Majorana zero modes at the two ends of the chain. With the goal of emulating this model, we build a chain of three quantum dots in a semiconductor nanowire. We observe Andreev bound states in each of the three dots and study their magnetic field and gate voltage dependence. Theory indicates that triple dot states acquire Majorana polarization when Andreev states in all three dots reach zero energy in a narrow range of magnetic field. In our device Andreev states in one of the dots reach zero energy at a lower field than in other two, placing the Majorana regime out of reach. Devices with greater uniformity or with independent control over superconductor-semiconductor coupling should can realize the Kitaev chain with high yield. Due to its overall tunability and design flexibility the quantum dot system remains promising for quantum simulation of interesting models and in particular for modular topological quantum devices.

cond-mat.mes-hall

MultiXNet: Multiclass Multistage Multimodal Motion Prediction

One of the critical pieces of the self-driving puzzle is understanding the surroundings of a self-driving vehicle (SDV) and predicting how these surroundings will change in the near future. To address this task we propose MultiXNet, an end-to-end approach for detection and motion prediction based directly on lidar sensor data. This approach builds on prior work by handling multiple classes of traffic actors, adding a jointly trained second-stage trajectory refinement step, and producing a multimodal probability distribution over future actor motion that includes both multiple discrete traffic behaviors and calibrated continuous position uncertainties. The method was evaluated on large-scale, real-world data collected by a fleet of SDVs in several cities, with the results indicating that it outperforms existing state-of-the-art approaches.

cs.CV

Temporally-Continuous Probabilistic Prediction using Polynomial Trajectory Parameterization

A commonly-used representation for motion prediction of actors is a sequence of waypoints (comprising positions and orientations) for each actor at discrete future time-points. While this approach is simple and flexible, it can exhibit unrealistic higher-order derivatives (such as acceleration) and approximation errors at intermediate time steps. To address this issue we propose a simple and general representation for temporally continuous probabilistic trajectory prediction that is based on polynomial trajectory parameterization. We evaluate the proposed representation on supervised trajectory prediction tasks using two large self-driving data sets. The results show realistic higher-order derivatives and better accuracy at interpolated time-points, as well as the benefits of the inferred noise distributions over the trajectories. Extensive experimental studies based on existing state-of-the-art models demonstrate the effectiveness of the proposed approach relative to other representations in predicting the future motions of vehicle, bicyclist, and pedestrian traffic actors.

cs.CV

Andreev Molecules in Semiconductor Nanowire Double Quantum Dots

Quantum simulation is a way to study unexplored Hamiltonians by mapping them onto the assemblies of well-understood quantum systems such as ultracold atoms in optical lattices, trapped ions or superconducting circuits. Semiconductor nanostructures which form the backbone of classical computing hold largely untapped potential for quantum simulation. In particular, chains of quantum dots in semiconductor nanowires can be used to emulate one-dimensional Hamiltonians such as the toy model of a topological p-wave superconductor. Here we realize a building block of this model, a double quantum dot with superconducting contacts, in an indium antimonide nanowire. In each dot, tunnel-coupling to a superconductor induces Andreev bound states. We demonstrate that these states hybridize to form the double-dot Andreev molecular states. We establish the parity and the spin structure of Andreev molecular levels by monitoring their evolution in electrostatic potential and magnetic field. Understanding Andreev molecules is a key step towards building longer chains which are predicted to generate Majorana bound states at the end sites. Two superconducting quantum dots are already sufficient to test the fusion rules of Majorana bound states, a milestone towards fault-tolerant topological quantum computing.

cond-mat.mes-hall

Magnetic Field Evolution of Spin Blockade in Ge/Si Nanowire Double Quantum Dots

We perform transport measurements on double quantum dots defined in Ge/Si core/shell nanowires and focus on Pauli spin blockade in the regime where tens of holes occupy each dot. We identify spin blockade through the magnetic field dependence of leakage current. We find both a dip and a peak in the leakage current at zero field. We analyze this behavior in terms of the quantum dot parameters such as coupling to the leads, interdot tunnel coupling as well as spin-orbit interaction. We find a lower bound for spin-orbit interaction with $l_{\rm so}=500$ nm. We also extract large and anisotropic effective Land$\rm \acute{e}$ g-factors, with larger g-factors in the direction perpendicular to the nanowire axis in agreement with previous studies and experiments but with larger values reported here.

cond-mat.mes-hall

High critical magnetic field superconducting contacts to Ge/Si core/shell nanowires

Contacts between high critical field superconductors and semiconductor nanowires are important in the context of topological quantum circuits in which superconductivity must be sustained to high magnetic fields. Here we demonstrate gate-tunable supercurrent in NbTiN-Ge/Si core/shell nanowire-NbTiN junctions. We observe supercurrents up to magnetic fields of 800 mT. The induced soft superconducting gap measured by co-tunneling through a quantum dot is 220 μeV. To improve contact transparency, we deposit an aluminum interlayer prior to NbTiN and observe a systematic change in the device pinch-off voltages with aluminum thickness. We inform future multi-step device fabrication by observing that aluminum anneals with the Ge/Si nanowire at 180° celsius, the baking temperature of common electron beam lithography resists.

cond-mat.mes-hall