SearcharxivSearch

arXiv subjects

Youjiang Xu

Publications and source records attributed to Youjiang Xu.

12 recordsLinked to original sources

DanceHMR: Hand-Aware Whole-Body Human Mesh Recovery from Monocular Videos

Monocular video human mesh recovery is essential for digital humans, avatar animation, and embodied simulation, where both temporal stability and expressive whole-body motion are required. Existing video HMR methods produce coherent body motion but often overlook detailed hand articulation, while image-based whole-body methods recover SMPL-X meshes independently per frame, often leading to jittery and inaccurate hand motion. We present a temporally coherent whole-body HMR framework for challenging in-the-wild monocular videos. Our model unifies body context and part-specific hand observations through residual body-hand fusion, enabling stable body motion and detailed hand recovery within a single temporal architecture. We further introduce close-up-aware augmentation to improve robustness under upper-body framing. Experiments on whole-body and body-only benchmarks demonstrate improved hand reconstruction and competitive body accuracy. Our method also produces temporally stable and 2D-consistent SMPL-X motion in challenging real-world videos.

cs.CV

Paramagnetic phases of strongly correlated ultracold fermions coupled to an optical cavity

We numerically study a gas of two-component fermions coupled to a transversely pumped optical cavity and confined to a two-dimensional static square optical lattice. In the dispersive regime, the steady state of the system is described by an extended Hubbard Hamiltonian with cavity-mediated long-range interactions. Using real-space dynamical mean-field theory (RDMFT), we investigate the formation of the (superradiant) checkerboard density-wave phase. Our analysis focuses on paramagnetic phases both at quarter and half filling. At quarter filling, we find a reentrant homogeneous Fermi liquid to density wave phase transition with increasing temperature, which is due to the higher entropy of the ordered phase. At half filling, in addition to the Fermi liquid to Mott insulator phase transition, marked by a vanishing quasiparticle residue at the Fermi level, we identify the transition into a density-wave phase. Due to perfect Fermi surface nesting at half filling, we find that arbitrarily small long-range interactions destabilize the system towards the density-wave phase in the absence of short-range interactions. By varying short- and long-range interactions at a fixed low temperature, we obtain the full phase diagram and identify a region of coexistence between the homogeneous Fermi liquid and Mott insulating phase with the density-wave phase. In this region, we determine the thermodynamic phase transition by comparing the energies of the different RDMFT solutions.

cond-mat.quant-gas

VITON-DiT: Learning In-the-Wild Video Try-On from Human Dance Videos via Diffusion Transformers

Video try-on stands as a promising area for its tremendous real-world potential. Prior works are limited to transferring product clothing images onto person videos with simple poses and backgrounds, while underperforming on casually captured videos. Recently, Sora revealed the scalability of Diffusion Transformer (DiT) in generating lifelike videos featuring real-world scenarios. Inspired by this, we explore and propose the first DiT-based video try-on framework for practical in-the-wild applications, named VITON-DiT. Specifically, VITON-DiT consists of a garment extractor, a Spatial-Temporal denoising DiT, and an identity preservation ControlNet. To faithfully recover the clothing details, the extracted garment features are fused with the self-attention outputs of the denoising DiT and the ControlNet. We also introduce novel random selection strategies during training and an Interpolated Auto-Regressive (IAR) technique at inference to facilitate long video generation. Unlike existing attempts that require the laborious and restrictive construction of a paired training dataset, severely limiting their scalability, VITON-DiT alleviates this by relying solely on unpaired human dance videos and a carefully designed multi-stage training strategy. Furthermore, we curate a challenging benchmark dataset to evaluate the performance of casual video try-on. Extensive experiments demonstrate the superiority of VITON-DiT in generating spatio-temporal consistent try-on results for in-the-wild videos with complicated human poses.

cs.CV

Getting topological invariants from snapshots: a protocol for defining and calculating topological invariants of systems with discrete parameter space

Topological invariants, including the Chern numbers, can topologically classify parameterized Hamiltonians. We find that topological invariants can be properly defined and calculated even if the parameter space is discrete, which is done by geodesic interpolation in the classifying space. We specifically present the interpolation protocol for the Chern numbers, which can be directly generalized to other topological invariants. The protocol generates a highly efficient algorithm for numerical calculation of the second and higher Chern numbers, by which arbitrary precision can be achieved given the values of the parameterized Hamiltonians on a coarse grid with a fixed resolution in the parameter space. Our findings also open up opportunities to study topology in finite-size systems where the parameter space can be naturally discrete.

cond-mat.mes-hall

Invisible flat bands on a topological chiral edge

We prove that invisible bands associated with zeros of the single-particle Green's function exist ubiquitously at topological interfaces of 2D Chern insulators, dual to the chiral edge/domain-wall modes. We verify this statement in a repulsive Hubbard model with a topological flat band, using real-space dynamical mean-field theory to study the domain walls of its ferromagnetic ground state. Moreover, our numerical results show that the chiral modes are split into branches due to the interaction, and that the branches are connected by invisible flat bands. Our work provides deeper insight into interacting topological systems.

cond-mat.str-el

Faster Meta Update Strategy for Noise-Robust Deep Learning

It has been shown that deep neural networks are prone to overfitting on biased training data. Towards addressing this issue, meta-learning employs a meta model for correcting the training bias. Despite the promising performances, super slow training is currently the bottleneck in the meta learning approaches. In this paper, we introduce a novel Faster Meta Update Strategy (FaMUS) to replace the most expensive step in the meta gradient computation with a faster layer-wise approximation. We empirically find that FaMUS yields not only a reasonably accurate but also a low-variance approximation of the meta gradient. We conduct extensive experiments to verify the proposed method on two tasks. We show our method is able to save two-thirds of the training time while still maintaining the comparable or achieving even better generalization performance. In particular, our method achieves the state-of-the-art performance on both synthetic and realistic noisy labels, and obtains promising performance on long-tailed recognition on standard benchmarks.

cs.LG

Multicriticality and quantum fluctuation in generalized Dicke model

We consider an important generalization of the Dicke model in which multi-level atoms, instead of two-level atoms as in conventional Dicke model, interact with a single photonic mode. We explore the phase diagram of a broad class of atom-photon coupling schemes and show that, under this generalization, the Dicke model can become multicritical. For a subclass of experimentally realizable schemes, multicritical conditions of arbitrary order can be expressed analytically in compact forms. We also calculate the atom-photon entanglement entropy for both critical and non-critical cases. We find that the order of the criticality strongly affects the critical entanglement entropy: higher order yields stronger entanglement. Our work provides deep insight into quantum phase transitions and multicriticality.

quant-ph

Building Flat-Band Lattice Models from Gram Matrices

We propose a powerful and convenient method to systematically design flat-band lattice models, which overcomes the difficulties underlying the previous method. Especially, our method requires no elaborate calculations, applies to arbitrary spatial dimensions, and guarantees to result in a completely flat ground band. We use this method to generate several classes of lattice models, including models with both short- and long-range hoppings, both topologically trivial and non-trivial flat bands. Some of these models were previously known. Our method, however, provides crucial new insights. For example, we have reproduced and generalized the Kapit-Mueller model [Kapit and Mueller, Phys. Rev. Lett. \textbf{105}, 215303 (2010)] and demonstrated a universal scaling rule between the flat band degeneracy and the magnetic flux that was not noticed in previous studies. We show that the flat band of this model results from the (over-)completeness properties of coherent states.

cond-mat.quant-gas

Geometry Normalization Networks for Accurate Scene Text Detection

Large geometry (e.g., orientation) variances are the key challenges in the scene text detection. In this work, we first conduct experiments to investigate the capacity of networks for learning geometry variances on detecting scene texts, and find that networks can handle only limited text geometry variances. Then, we put forward a novel Geometry Normalization Module (GNM) with multiple branches, each of which is composed of one Scale Normalization Unit and one Orientation Normalization Unit, to normalize each text instance to one desired canonical geometry range through at least one branch. The GNM is general and readily plugged into existing convolutional neural network based text detectors to construct end-to-end Geometry Normalization Networks (GNNets). Moreover, we propose a geometry-aware training scheme to effectively train the GNNets by sampling and augmenting text instances from a uniform geometry variance distribution. Finally, experiments on popular benchmarks of ICDAR 2015 and ICDAR 2017 MLT validate that our method outperforms all the state-of-the-art approaches remarkably by obtaining one-forward test F-scores of 88.52 and 74.54 respectively.

cs.CV

Emergent Universality in a Quantum Tricritical Dicke Model

We propose a generalized Dicke model which supports a quantum tricritical point. We map out the phase diagram and investigate the critical behaviors of the model through exact low-energy effective Hamiltonian in the thermodynamic limit. As predicted by the Landau theory of phase transition, the order parameter shows non-universality at the tricritical point. Nevertheless, as a result of the separation of the classical and the quantum degrees of freedom, we find a universal relation between the excitation gap and the entanglement entropy for the entire critical line including the tricritical point. Here the universality is carried by the emergent quantum modes, whereas the order parameter is determined classically.

quant-ph

Movie Question Answering: Remembering the Textual Cues for Layered Visual Contents

Movies provide us with a mass of visual content as well as attracting stories. Existing methods have illustrated that understanding movie stories through only visual content is still a hard problem. In this paper, for answering questions about movies, we put forward a Layered Memory Network (LMN) that represents frame-level and clip-level movie content by the Static Word Memory module and the Dynamic Subtitle Memory module, respectively. Particularly, we firstly extract words and sentences from the training movie subtitles. Then the hierarchically formed movie representations, which are learned from LMN, not only encode the correspondence between words and visual content inside frames, but also encode the temporal alignment between sentences and frames inside movie clips. We also extend our LMN model into three variant frameworks to illustrate the good extendable capabilities. We conduct extensive experiments on the MovieQA dataset. With only visual content as inputs, LMN with frame-level representation obtains a large performance improvement. When incorporating subtitles into LMN to form the clip-level representation, we achieve the state-of-the-art performance on the online evaluation task of 'Video+Subtitles'. The good performance successfully demonstrates that the proposed framework of LMN is effective and the hierarchically formed movie representations have good potential for the applications of movie question answering.

cs.CV

Number-conserving interacting fermion models with exact topological superconducting ground states

We present a method to construct number-conserving Hamiltonians whose ground states exactly reproduce an arbitrarily chosen BCS-type mean-field state. Such parent Hamiltonians can be constructed not only for the usual $s$-wave BCS state, but also for more exotic states of this form, including the ground states of Kitaev wires and 2D topological superconductors. This method leads to infinite families of locally-interacting fermion models with exact topological superconducting ground states. After explaining the general technique, we apply this method to construct two specific classes of models. The first one is a one-dimensional double wire lattice model with Majorana-like degenerate ground states. The second one is a two-dimensional $p_x+ip_y$ superconducting model, where we also obtain analytic expressions for topologically degenerate ground states in the presence of vortices. Our models may provide a deeper conceptual understanding of how Majorana zero modes could emerge in condensed matter systems, as well as inspire novel routes to realize them in experiment.

cond-mat.str-el