SearcharxivSearch

arXiv subjects

Xiao Dong

Publications and source records attributed to Xiao Dong.

At least 19 recordsLinked to original sources

SAKE: Self-aware Knowledge Exploitation-Exploration for Grounded Multimodal Named Entity Recognition

Grounded Multimodal Named Entity Recognition (GMNER) aims to extract named entities and localize their visual regions within image-text pairs, serving as a pivotal capability for various downstream applications. In open-world social media platforms, GMNER remains challenging due to the prevalence of long-tailed, rapidly evolving, and unseen entities. To tackle this, existing approaches typically rely on either external knowledge exploration through heuristic retrieval or internal knowledge exploitation via iterative refinement in Multimodal Large Language Models (MLLMs). However, heuristic retrieval often introduces noisy or conflicting evidence that degrades precision on known entities, while solely internal exploitation is constrained by the knowledge boundaries of MLLMs and prone to hallucinations. To address this, we propose SAKE, an end-to-end agentic framework that harmonizes internal knowledge exploitation and external knowledge exploration via self-aware reasoning and adaptive search tool invocation. We implement this via a two-stage training paradigm. First, we propose Difficulty-aware Search Tag Generation, which quantifies the model's entity-level uncertainty through multiple forward samplings to produce explicit knowledge-gap signals. Based on these signals, we construct SAKE-SeCoT, a high-quality Chain-of-Thought dataset that equips the model with basic self-awareness and tool-use capabilities through supervised fine-tuning. Second, we employ agentic reinforcement learning with a hybrid reward function that penalizes unnecessary retrieval, enabling the model to evolve from rigid search imitation to genuine self-aware decision-making about when retrieval is truly necessary. Extensive experiments on two widely used social media benchmarks demonstrate SAKE's effectiveness.

cs.IR

Antitrust on Aisle Five: How Well Do Divestiture Remedies Work?

Antitrust authorities frequently rely on structural divestitures to address competitive concerns raised by mergers. Using census-level establishment data and proprietary transaction records from the U.S. grocery sector, we provide systematic evidence on the long-run effects of such remedies. Divested stores experience an average 31 percent decline in employment over five years, driven by elevated exit rates and persistent contraction among surviving establishments. Sales similarly decline. Transaction-level evidence indicates that divested assets are systematically weaker and are often transferred to lower-capability buyers. These findings suggest that structural remedies may be less effective when the implementation of divestitures allows merging parties substantial discretion over the assets and buyers involved.

econ.GN

High-Pressure Structural Evolution of Na2ZrSi2O7 and Na2ZrSi2O7.H2O: Topology-Driven Compression Behaviors, Phase Stability, and Electronic Transitions

Silicate frameworks exhibit diverse structural responses under extreme conditions, which are strongly influenced by hydration. Here, we present a comparative high-pressure synchrotron X-ray diffraction study of Na2ZrSi2O7 and its hydrated analogue Na2ZrSi2O7.H2O up to 30 GPa, combined with electronic structure calculations. At ambient conditions, both phases share the same primary building units (PBUs: [ZrO6] and [SiO4]) but differ in secondary building units (SBUs, M2T4 vs. M2T6). Under compression, Na2ZrSi2O7 undergoes a phase transition near 15 GPa, while the hydrated phase remains stable throughout the pressure range. The anhydrous compound exhibits a higher bulk modulus (B0 = 77.1 GPa) and less anisotropic compression compared with the hydrated phase (B0 = 66.3 GPa). Distinct deformation mechanisms are observed: the anhydrous framework accommodates pressure through [ZrO6] octahedral distortion, whereas the hydrated framework compresses via [Si2O7] group tilting. Electronic structure calculations indicate band gap widening with pressure in both phases; notably, Na2ZrSi2O7 shows a direct-to-indirect band gap transition, whereas the hydrated phase retains a direct gap. These results reveal how hydration-driven topological modifications at the secondary building unit scale dictate the pressure-induced structural evolution, phase stability, and electronic properties of zirconosilicate frameworks.

cond-mat.mtrl-sci

Domain-Direct Band Gaps: Classification and Material Realization

The conventional classification of direct band-gap semiconductors relies on point-like extrema in momentum space. Here, we introduce the concept of domain-direct band gaps, where the conduction-band minimum (CBM) and valence-band maximum (VBM) form extended manifolds in the Brillouin zone. We demonstrate this concept through the material realization of an extreme two-dimensional-two-dimensional (2D-2D) domain-direct band gap in twisted diamond. First-principles calculations show that both the CBM and VBM exhibit nearly flat 2D manifolds in the kx-ky plane with minimal energy variation (a few meV), yielding a direct band gap of 3.264 eV. In contrast, strong dispersion along the out-of-plane kz direction induces anisotropic carrier dynamics, with strongly suppressed in-plane Fermi velocities (down to about 10$^1$-10$^3$ m/s in certain directions) and much larger out-of-plane velocities (about 10$^6$ m/s). The nearly flat CBM and VBM manifolds enhance the joint density of states, leading to a pronounced optical absorption peak at the band gap onset. This new type of domain-direct gap, coupled with strong directional anisotropy, opens up opportunities for anisotropic optoelectronic applications. Our results establish domain-direct band gaps as a new class of semiconductors, demonstrating their feasibility in real materials.

cond-mat.mtrl-sci

Enhancing Automatic Chord Recognition via Pseudo-Labeling and Knowledge Distillation

Automatic Chord Recognition (ACR) is constrained by the scarcity of aligned chord annotations, which are costly to acquire. At the same time, open-weight pre-trained models are more accessible than their proprietary training data. In this work, we present a two-stage training pipeline that leverages pre-trained models together with unlabeled audio. The proposed method decouples training into two stages. In the first stage, we use the pre-trained BTC model as a teacher to generate pseudo-labels for over 1,000 hours of diverse unlabeled audio and train a student model solely on these pseudo-labels. In the second stage, the student is continually trained on ground-truth labels as they become available. To prevent catastrophic forgetting of the representations learned in the first stage, we apply selective knowledge distillation (KD) from the teacher as a regularizer. In our experiments, two models (BTC, 2E1D) were used as students. In Stage 1, using only pseudo-labels, the BTC student achieves about 99% of the teacher's performance, while the 2E1D model achieves about 97% of the teacher's performance across seven standard mir_eval metrics. After continual training with labeled data in Stage 2, the resulting BTC student model consistently surpasses both the traditional supervised learning baseline and the original pre-trained teacher model across all metrics. The resulting 2E1D student model also outperforms the supervised baseline and approaches teacher-level performance, with both models demonstrating substantial gains on rare chord qualities.

cs.SD

Effect of superconductivity by Nb and V substitution in kagome CaPd5

Materials featuring kagome lattices have attracted significant research interest due to their unique geometric frustration, which gives rise to rich physical phenomena such as non-trivial topology, spin fluctuations, and superconductivity. In this work, using CaPd5 as the prototype structure, we discover and systematically investigate a new class of kagome superconductors, CaMxPd5-x (M = Nb and V) alloys. First-principles calculations confirm that these compounds are non-magnetic metals, among which four are dynamically stable: CaNb5, CaV5, CaNb2Pd3, and CaV2Pd3. CaNb5 is identified as a strong electron-phonon coupling (EPC) superconductor with the highest superconducting transition temperature (Tc) of 10.1 K, which can be further increased to 12.8 K under external pressure. In contrast, CaV5, CaNb2Pd3, and CaV2Pd3 exhibit weaker EPC and correspondingly lower Tc values. Furthermore, by applying the method of symmetry indicators, we systematically classify the topological and nodal characteristics of CaNb5, providing valuable insights for determining its superconducting pairing symmetry. Our findings demonstrate that Nb and V substitution in kagome CaPd5 provides an effective route for designing a new type of kagome superconductor with relatively high Tc. This study also offers new perspectives on topological superconductivity in kagome systems and establishes a useful guideline for discovering other superconducting materials with unique properties.

cond-mat.supr-con

FastFit: Accelerating Multi-Reference Virtual Try-On via Cacheable Diffusion Models

Despite its great potential, virtual try-on technology is hindered from real-world application by two major challenges: the inability of current methods to support multi-reference outfit compositions (including garments and accessories), and their significant inefficiency caused by the redundant re-computation of reference features in each denoising step. To address these challenges, we propose FastFit, a high-speed multi-reference virtual try-on framework based on a novel cacheable diffusion architecture. By employing a Semi-Attention mechanism and substituting traditional timestep embeddings with class embeddings for reference items, our model fully decouples reference feature encoding from the denoising process with negligible parameter overhead. This allows reference features to be computed only once and losslessly reused across all steps, fundamentally breaking the efficiency bottleneck and achieving an average 3.5x speedup over comparable methods. Furthermore, to facilitate research on complex, multi-reference virtual try-on, we introduce DressCode-MR, a new large-scale dataset. It comprises 28,179 sets of high-quality, paired images covering five key categories (tops, bottoms, dresses, shoes, and bags), constructed through a pipeline of expert models and human feedback refinement. Extensive experiments on the VITON-HD, DressCode, and our DressCode-MR datasets show that FastFit surpasses state-of-the-art methods on key fidelity metrics while offering its significant advantage in inference efficiency.

cs.CV

TransMamba: Fast Universal Architecture Adaption from Transformers to Mamba

Transformer-based architectures have become the backbone of both uni-modal and multi-modal foundation models, largely due to their scalability via attention mechanisms, resulting in a rich ecosystem of publicly available pre-trained models such as LLaVA, CLIP, and DeiT, etc. In parallel, emerging sub-quadratic architectures like Mamba offer promising efficiency gains by enabling global context modeling with linear complexity. However, training these architectures from scratch remains resource-intensive (e.g., in terms of data and time). Motivated by this challenge, we explore a cross-architecture knowledge transfer paradigm, termed TransMamba, that facilitates the reuse of Transformer pre-trained knowledge. We propose a two-stage framework to accelerate the training of Mamba-based models, ensuring their effectiveness across both uni-modal and multi-modal tasks. The first stage leverages pre-trained Transformer models to initialize critical components of the Mamba architecture. To bridge architectural and dimensional gaps, we develop a selective weight subcloning strategy and a layered initialization scheme that prioritizes the early $n$ layers. Building on this initialization, the second stage introduces an adaptive multi-directional knowledge distillation method. This mechanism employs layer-wise adaptive scaling factors to align Mamba representations with their Transformer counterparts, while accommodating the scanning order variations inherent to multi-modal Mamba architectures. Despite operating with a reduced training dataset and a more compact model architecture, TransMamba consistently outperforms baseline approaches across diverse mamba-based backbones (e.g., PlainMamba, Vmamba, ViM and VideoMamba) and downstream tasks (e.g., image classification, visual question answering, text-video retrieval and multimodal reasoning). All code and implementation details will be released.

cs.CV

MPd5 kagome superconductors studied by density functional calculations

Kagome materials, which are composed of hexagons tiled with a shared triangle, have inspired enormous interest due to their unique structures and rich physical properties; exploring superconducting material systems with new kagome structures is still an important research direction. Here, we predict a type of kagome superconductor, MPd5 (M is a group-IIA metal element), and identify that it exhibits coexistence of superconductivity and nontrivial topological properties. We uncover its phonon-mediated superconductivity by the density functional theory for superconductors, predicting the superconducting transition temperatures (Tc) of 2.64, 2.03, and 1.50 K for CaPd5, SrPd5, and BaPd5, respectively. These Tc can be effectively tuned through the application of external pressure and electron doping. The present results also demonstrate that MPd5 have topological properties; e.g., CaPd5 shows topological nontrivial intersection near the Fermi level (EF). Our results indicate that MPd5 materials can be an emerging material platform with rich exotic physics in their kagome structures, and render themselves excellent candidates for superconducting and advanced functional materials that could be utilized in topological quantum computing and information technology.

cond-mat.supr-con

Nudged elastic band calculations of stacking and dislocation pathways in diamond

Diamond, the hardest natural crystal, has attracted significant attention for its plasticity, which is reported to be determined by its stacking faults. Studies mainly focused on one-dimensional linear pathways in stacking transitions, neglecting its transverse freedom on the main slip plane. However, in an actual stacking procedure, stacking faults can follow curve line along the slip plane rather than constrained to straight lines. In this study, using ab initio calculations, we mapped the {\gamma}-surface, defined as the landscape of generalized stacking fault energies, along the weakest direction of the {111} orientation in diamond. We then applied the Nudged Elastic Band (NEB) method to determine the minimum energy paths, finding significantly reduced stacking energy barriers compared to previous reports (for the glide-set, our energy barrier is only one-third of that for the traditional direct path). Our calculations reveal that the glide-set can round its high-energy peak, with a lower energy barrier within the entire stacking plane than the shuffle-set. By employing the NEB method, we have constructed the minimum energy path (MEP) for both the stacking and dislocation procedures. Our results provide new insights into the plasticity and stacking faults of diamond, advancing the understanding of superhard carbon material transition, especially the diamond under shear stress.

cond-mat.mtrl-sci

WonderHuman: Hallucinating Unseen Parts in Dynamic 3D Human Reconstruction

In this paper, we present WonderHuman to reconstruct dynamic human avatars from a monocular video for high-fidelity novel view synthesis. Previous dynamic human avatar reconstruction methods typically require the input video to have full coverage of the observed human body. However, in daily practice, one typically has access to limited viewpoints, such as monocular front-view videos, making it a cumbersome task for previous methods to reconstruct the unseen parts of the human avatar. To tackle the issue, we present WonderHuman, which leverages 2D generative diffusion model priors to achieve high-quality, photorealistic reconstructions of dynamic human avatars from monocular videos, including accurate rendering of unseen body parts. Our approach introduces a Dual-Space Optimization technique, applying Score Distillation Sampling (SDS) in both canonical and observation spaces to ensure visual consistency and enhance realism in dynamic human reconstruction. Additionally, we present a View Selection strategy and Pose Feature Injection to enforce the consistency between SDS predictions and observed data, ensuring pose-dependent effects and higher fidelity in the reconstructed avatar. In the experiments, our method achieves SOTA performance in producing photorealistic renderings from the given monocular video, particularly for those challenging unseen parts. The project page and source code can be found at https://wyiguanw.github.io/WonderHuman/.

cs.CV

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions

Building on the success of diffusion models, significant advancements have been made in multimodal image generation tasks. Among these, human image generation has emerged as a promising technique, offering the potential to revolutionize the fashion design process. However, existing methods often focus solely on text-to-image or image reference-based human generation, which fails to satisfy the increasingly sophisticated demands. To address the limitations of flexibility and precision in human generation, we introduce ComposeAnyone, a controllable layout-to-human generation method with decoupled multimodal conditions. Specifically, our method allows decoupled control of any part in hand-drawn human layouts using text or reference images, seamlessly integrating them during the generation process. The hand-drawn layout, which utilizes color-blocked geometric shapes such as ellipses and rectangles, can be easily drawn, offering a more flexible and accessible way to define spatial layouts. Additionally, we introduce the ComposeHuman dataset, which provides decoupled text and reference image annotations for different components of each human image, enabling broader applications in human image generation tasks. Extensive experiments on multiple datasets demonstrate that ComposeAnyone generates human images with better alignment to given layouts, text descriptions, and reference images, showcasing its multi-task capability and controllability.

cs.CV

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation

Virtual try-on (VTON) technology has gained attention due to its potential to transform online retail by enabling realistic clothing visualization of images and videos. However, most existing methods struggle to achieve high-quality results across image and video try-on tasks, especially in long video scenarios. In this work, we introduce CatV2TON, a simple and effective vision-based virtual try-on (V2TON) method that supports both image and video try-on tasks with a single diffusion transformer model. By temporally concatenating garment and person inputs and training on a mix of image and video datasets, CatV2TON achieves robust try-on performance across static and dynamic settings. For efficient long-video generation, we propose an overlapping clip-based inference strategy that uses sequential frame guidance and Adaptive Clip Normalization (AdaCN) to maintain temporal consistency with reduced resource demands. We also present ViViD-S, a refined video try-on dataset, achieved by filtering back-facing frames and applying 3D mask smoothing for enhanced temporal consistency. Comprehensive experiments demonstrate that CatV2TON outperforms existing methods in both image and video try-on tasks, offering a versatile and reliable solution for realistic virtual try-ons across diverse scenarios.

cs.CV

RMAvatar: Photorealistic Human Avatar Reconstruction from Monocular Video Based on Rectified Mesh-embedded Gaussians

We introduce RMAvatar, a novel human avatar representation with Gaussian splatting embedded on mesh to learn clothed avatar from a monocular video. We utilize the explicit mesh geometry to represent motion and shape of a virtual human and implicit appearance rendering with Gaussian Splatting. Our method consists of two main modules: Gaussian initialization module and Gaussian rectification module. We embed Gaussians into triangular faces and control their motion through the mesh, which ensures low-frequency motion and surface deformation of the avatar. Due to the limitations of LBS formula, the human skeleton is hard to control complex non-rigid transformations. We then design a pose-related Gaussian rectification module to learn fine-detailed non-rigid deformations, further improving the realism and expressiveness of the avatar. We conduct extensive experiments on public datasets, RMAvatar shows state-of-the-art performance on both rendering quality and quantitative evaluations. Please see our project page at https://rm-avatar.github.io.

cs.CV

LiTformer: Efficient Modeling and Analysis of High-Speed Link Transmitters Using Non-Autoregressive Transformer

High-speed serial links are fundamental to energy-efficient and high-performance computing systems such as artificial intelligence, 5G mobile and automotive, enabling low-latency and high-bandwidth communication. Transmitters (TXs) within these links are key to signal quality, while their modeling presents challenges due to nonlinear behavior and dynamic interactions with links. In this paper, we propose LiTformer: a Transformer-based model for high-speed link TXs, with a non-sequential encoder and a Transformer decoder to incorporate link parameters and capture long-range dependencies of output signals. We employ a non-autoregressive mechanism in model training and inference for parallel prediction of the signal sequence. LiTformer achieves precise TX modeling considering link impacts including crosstalk from multiple links, and provides fast prediction for various long-sequence signals with high data rates. Experimental results show that LiTformer achieves 148-456$\times$ speedup for 2-link TXs and 404-944$\times$ speedup for 16-link with mean relative errors of 0.68-1.25%, supporting 4-bit signals at Gbps data rates of single-ended and differential TXs, as well as PAM4 TXs.

eess.SP

A Survey of Foundation Models for Music Understanding

Music is essential in daily life, fulfilling emotional and entertainment needs, and connecting us personally, socially, and culturally. A better understanding of music can enhance our emotions, cognitive skills, and cultural connections. The rapid advancement of artificial intelligence (AI) has introduced new ways to analyze music, aiming to replicate human understanding of music and provide related services. While the traditional models focused on audio features and simple tasks, the recent development of large language models (LLMs) and foundation models (FMs), which excel in various fields by integrating semantic information and demonstrating strong reasoning abilities, could capture complex musical features and patterns, integrate music with language and incorporate rich musical, emotional and psychological knowledge. Therefore, they have the potential in handling complex music understanding tasks from a semantic perspective, producing outputs closer to human perception. This work, to our best knowledge, is one of the early reviews of the intersection of AI techniques and music understanding. We investigated, analyzed, and tested recent large-scale music foundation models in respect of their music comprehension abilities. We also discussed their limitations and proposed possible future directions, offering insights for researchers in this field.

cs.SD

OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion

Open-vocabulary detection is a challenging task due to the requirement of detecting objects based on class names, including those not encountered during training. Existing methods have shown strong zero-shot detection capabilities through pre-training and pseudo-labeling on diverse large-scale datasets. However, these approaches encounter two main challenges: (i) how to effectively eliminate data noise from pseudo-labeling, and (ii) how to efficiently leverage the language-aware capability for region-level cross-modality fusion and alignment. To address these challenges, we propose a novel unified open-vocabulary detection method called OV-DINO, which is pre-trained on diverse large-scale datasets with language-aware selective fusion in a unified framework. Specifically, we introduce a Unified Data Integration (UniDI) pipeline to enable end-to-end training and eliminate noise from pseudo-label generation by unifying different data sources into detection-centric data format. In addition, we propose a Language-Aware Selective Fusion (LASF) module to enhance the cross-modality alignment through a language-aware query selection and fusion process. We evaluate the performance of the proposed OV-DINO on popular open-vocabulary detection benchmarks, achieving state-of-the-art results with an AP of 50.6% on the COCO benchmark and 40.1% on the LVIS benchmark in a zero-shot manner, demonstrating its strong generalization ability. Furthermore, the fine-tuned OV-DINO on COCO achieves 58.4% AP, outperforming many existing methods with the same backbone. The code for OV-DINO is available at https://github.com/wanghao9610/OV-DINO.

cs.CV

CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models

Virtual try-on methods based on diffusion models achieve realistic effects but often require additional encoding modules, a large number of training parameters, and complex preprocessing, which increases the burden on training and inference. In this work, we re-evaluate the necessity of additional modules and analyze how to improve training efficiency and reduce redundant steps in the inference process. Based on these insights, we propose CatVTON, a simple and efficient virtual try-on diffusion model that transfers in-shop or worn garments of arbitrary categories to target individuals by concatenating them along spatial dimensions as inputs of the diffusion model. The efficiency of CatVTON is reflected in three aspects: (1) Lightweight network. CatVTON consists only of a VAE and a simplified denoising UNet, removing redundant image and text encoders as well as cross-attentions, and includes just 899.06M parameters. (2) Parameter-efficient training. Through experimental analysis, we identify self-attention modules as crucial for adapting pre-trained diffusion models to the virtual try-on task, enabling high-quality results with only 49.57M training parameters. (3) Simplified inference. CatVTON eliminates unnecessary preprocessing, such as pose estimation, human parsing, and captioning, requiring only a person image and garment reference to guide the virtual try-on process, reducing over 49% memory usage compared to other diffusion-based methods. Extensive experiments demonstrate that CatVTON achieves superior qualitative and quantitative results compared to baseline methods and demonstrates strong generalization performance in in-the-wild scenarios, despite being trained solely on public datasets with 73K samples.

cs.CV