SearcharxivSearch

arXiv subjects

Xu Hou

Publications and source records attributed to Xu Hou.

6 recordsLinked to original sources

MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts for Multimodal Knowledge Graph Completion

Multimodal Knowledge Graph Completion (MKGC) requires inferring missing entities from structural, textual, and visual cues. Existing diffusion-based MKGC methods usually denoise directly on raw multimodal features. Such a design forces the denoiser to simultaneously perform relation-dependent cue selection, cross-modal semantic alignment, and structure-aware entity generation, which introduces noisy and semantically inconsistent conditions for diffusion and consequently leads to suboptimal completion performance. To address this limitation, we propose MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts (MGDT), a novel MKGC framework built on an align-then-diffuse paradigm. MGDT first employs a Relation-Adaptive Semantic Routing Mixture-of-Experts (RASR-MoE) module to select relation-relevant multimodal semantic transformation paths and suppress irrelevant modality interference. MGDT then uses a frozen Multimodal Large Language Model (MLLM) as a semantic anchor to align the routed multimodal representations into a unified latent space and reduce cross-modal semantic heterogeneity. Finally, a Knowledge Graph Diffusion Transformer (KGDT) performs graph-conditioned denoising generation in the aligned space to produce the missing entity representation. Experiments on three benchmark datasets show that MGDT consistently outperforms strong baselines.

cs.AI

ELMM: Efficient Lightweight Multimodal Large Language Models for Multimodal Knowledge Graph Completion

Multimodal Knowledge Graphs (MKGs) extend traditional knowledge graphs by incorporating visual and textual modalities, enabling richer and more expressive entity representations. However, existing MKGs often suffer from incompleteness, which hinder their effectiveness in downstream tasks. Therefore, multimodal knowledge graph completion (MKGC) task is receiving increasing attention. While large language models (LLMs) have shown promise for knowledge graph completion (KGC), their application to the multimodal setting remains underexplored. Moreover, applying Multimodal Large Language Models (MLLMs) to the task of MKGC introduces significant challenges: (1) the large number of image tokens per entity leads to semantic noise and modality conflicts, and (2) the high computational cost of processing large token inputs. To address these issues, we propose Efficient Lightweight Multimodal Large Language Models (ELMM) for MKGC. ELMM proposes a Multi-view Visual Token Compressor (MVTC) based on multi-head attention mechanism, which adaptively compresses image tokens from both textual and visual views, thereby effectively reducing redundancy while retaining necessary information and avoiding modality conflicts. Additionally, we design an attention pruning strategy to remove redundant attention layers from MLLMs, thereby significantly reducing the inference cost. We further introduce a linear projection to compensate for the performance degradation caused by pruning. Extensive experiments on four benchmark datasets demonstrate that ELMM achieves state-of-the-art performance.

cs.AI

DiffusionCom: Structure-Aware Multimodal Diffusion Model for Multimodal Knowledge Graph Completion

Most current MKGC approaches are predominantly based on discriminative models that maximize conditional likelihood. These approaches struggle to efficiently capture the complex connections in real-world knowledge graphs, thereby limiting their overall performance. To address this issue, we propose a structure-aware multimodal Diffusion model for multimodal knowledge graph Completion (DiffusionCom). DiffusionCom innovatively approaches the problem from the perspective of generative models, modeling the association between the $(head, relation)$ pair and candidate tail entities as their joint probability distribution $p((head, relation), (tail))$, and framing the MKGC task as a process of gradually generating the joint probability distribution from noise. Furthermore, to fully leverage the structural information in MKGs, we propose Structure-MKGformer, an adaptive and structure-aware multimodal knowledge representation learning method, as the encoder for DiffusionCom. Structure-MKGformer captures rich structural information through a multimodal graph attention network (MGAT) and adaptively fuses it with entity representations, thereby enhancing the structural awareness of these representations. This design effectively addresses the limitations of existing MKGC methods, particularly those based on multimodal pre-trained models, in utilizing structural information. DiffusionCom is trained using both generative and discriminative losses for the generator, while the feature extractor is optimized exclusively with discriminative loss. This dual approach allows DiffusionCom to harness the strengths of both generative and discriminative models. Extensive experiments on the FB15k-237-IMG and WN18-IMG datasets demonstrate that DiffusionCom outperforms state-of-the-art models.

cs.IR

Machine learning-based seeing estimation and prediction using multi-layer meteorological data at Dome A, Antarctica

Atmospheric seeing is one of the most important parameters for evaluating and monitoring an astronomical site. Moreover, being able to predict the seeing in advance can guide observing decisions and significantly improve the efficiency of telescopes. However, it is not always easy to obtain long-term and continuous seeing measurements from a standard instrument such as differential image motion monitor (DIMM), especially for those unattended observatories with challenging environments such as Dome A, Antarctica. In this paper, we present a novel machine learning-based framework for estimating and predicting seeing at a height of 8 m at Dome A, Antarctica, using only the data from a multi-layer automated weather station (AWS). In comparison with DIMM data, our estimate has a root mean square error (RMSE) of 0.18 arcsec, and the RMSE of predictions 20 minutes in the future is 0.12 arcsec for the seeing range from 0 to 2.2 arcsec. Compared with the persistence, where the forecast is the same as the last data point, our framework reduces the RMSE by 37 percent. Our method predicts the seeing within a second of computing time, making it suitable for real-time telescope scheduling.

astro-ph.IM

Manipulation of polar vortex chirality in oxide superlattices

Topological polar vortices that are the electric analogues of magnetic objects, present great potential in applications of future nanoelectronics due to their nanometer size, anomalous dielectric response, and chirality. To enable the functionalities, it is prerequisite to manipulate the polar states and chirality by using external stimuli. Here, we probe the evolutions of polar state and chirality of polar vortices in PbTiO3/SrTiO3 superlattices under electric field by using atomically resolved in situ scanning transmission electron microscopy and phase-field simulations. We find that the adjacent clockwise and counterclockwise vortex usually have opposite chirality. The phase-field simulations suggest that the rotation reversal or axial polarization switching can lead to the chirality change. Guided by which, we experimentally validate that the vortex rotation direction can be changed by applying and subsequently removing of electric fields, offering a potential strategy to manipulate the vortex chirality. The revealed details of dynamic behavior for individual polar vortices at atomic scale and the proposed strategy for chirality manipulation provide fundamentals for future device applications.

cond-mat.mtrl-sci

Creating topological polar structure in a nonpolar matter

Nontrivial topological structures offer rich playground in condensed matter physics including fluid dynamics, superconductivity, and ferromagnetism, and they promise alternative device configurations for post-Moore spintronics and electronics. Indeed, magnetic skyrmions are actively pursued for high-density data storage, while polar vortices with exotic negative capacitance may enable ultralow power consumption in microelectronics. Following extensive investigations on a variety of magnetic textures including vortices, domain walls and skyrmions in the past decades, studies on polar topologies have taken off in recent years, resulting in discoveries of closure domains, vortices, and skyrmions in ferroelectric materials. Nevertheless, the atomic-scale creation of topological polar structures is largely confined in a single ferroelectric system, PbTiO3 (PTO) with large polarization, casting doubt on the generality of polar topologies and limiting their potential applications. In this work, we successfully create previously unrealized atomic-scale polar antivortices in the nominally nonpolar SrTiO3 (STO), expanding the reaches of topological structures and completing an important missing link in polar topologies. The work shed considerable new insight into the formation of topological polar structures, and offers guidance in searching for new polar textures.

cond-mat.mtrl-sci