SearcharxivSearch

arXiv subjects

Xinyu Xu

Publications and source records attributed to Xinyu Xu.

At least 19 recordsLinked to original sources

A Physics-Informed Chemical Rule for Topological Materials Discovery

Topological phases of matter$\unicode{x2013}$comprising both insulators and semimetals$\unicode{x2013}$offer great potential for quantum applications, but identifying new candidates remains challenging due to expensive first-principles simulations and labor-intensive experimental workflows. Here we introduce a physics-informed chemical rule that integrates compositional, orbital and crystallographic descriptors within an interpretable linear framework. By explicitly encoding electron filling, space-group symmetry and orbital-resolved chemical environments, our method overcomes a fundamental limitation of composition-only heuristics$\unicode{x2013}$their inability to distinguish polymorphs with identical stoichiometry but different crystal structures. Using only elemental characteristics, our approach reduces a material's topological propensity to a single, physically interpretable score, enabling rapid and high-throughput assessment. The model achieves superior predictive performance while maintaining physical transparency, and identifies candidate topological materials where conventional symmetry indicators fail. Consequently, our framework enables rapid and interpretable exploration of complex materials spaces, establishing a scalable paradigm for the intelligent discovery of next-generation topological and quantum materials.

cond-mat.mtrl-sci

Quantum Telepathy: A Quantum Technology with Near-Term Applications

Quantum telepathy is the concept of using quantum entanglement to solve real-world problems involving decision coordination between parties with restricted communication. One possible reason for this restriction is a latency constraint: some pairs of parties do not have enough time to communicate with each other before they have to produce their outputs. Example scenarios include high frequency trading and distributed systems. Another reason is physical or operational isolation: for some pairs of parties, there is an obstacle to communication. Example scenarios include locating a stray traveler by a rescue team and coordination within a network where nodes are owned by competing firms. In this paper we give a concise overview of the different application areas of quantum telepathy. We find that these real-world problems can be modeled as a nonlocal game or its generalizations. We also discuss possible physical implementations. Quantum telepathy guarantees a quantum advantage via Bell's theorem and can directly solve real-world problems, such as reducing risk in high frequency trading or balancing data loads efficiently in ad hoc networks. Moreover, this quantum advantage can be physically realized with existing or near-term quantum hardware.

quant-ph

PhysSFI-Net: Physics-informed Geometric Learning of Skeletal and Facial Interactions for Orthognathic Surgical Outcome Prediction

Orthognathic surgery repositions jaw bones to restore occlusion and enhance facial aesthetics. Accurate simulation of postoperative facial morphology is essential for preoperative planning. This study aims to develop and validate a physics-informed geometric deep learning framework named PhysSFI-Net for precise prediction of soft tissue deformation following orthognathic surgery. The model integrates a hierarchical feature extraction module with attention mechanisms to capture skeletal-facial interactions, an LSTM-based sequential predictor for incremental deformation, and a biomechanics-inspired reconstruction module for high-resolution facial modeling. The model was trained on 135 patients and externally validated on an independent cohort of 33 patients. Model performance was assessed using point cloud shape error, surface deviation error and landmark error between predicted facial shapes with corresponding ground truths. Quantitative analysis demonstrated that PhysSFI-Net achieved a global shape error of 1.070 +/- 0.088 mm, a surface deviation error of 1.296 +/- 0.349 mm and a landmark error of 2.445 +/- 1.326 mm. Comparative experiments indicated that PhysSFI-Net outperformed the state-of-the-art method ACMT-Net and baseline models. External validation further confirmed its robustness with a global HD of 1.431 +/- 0.087 mm and consistently lower subregional and mesh-based errors. In conclusion, PhysSFI-Net enables interpretable, high-resolution prediction of postoperative facial morphology, showing strong potential for clinical application in orthognathic surgical planning.

cs.CV

Quantum-inspired Chemical Rule for Discovering Topological Materials

Topological materials exhibit unique electronic structures that underpin both fundamental quantum phenomena and next-generation technologies, yet their discovery remains constrained by the high computational cost of first-principles calculations and the slow, resource-intensive nature of experimental synthesis. Recent machine-learning approaches, such as the heuristic topogivity rule, offer a data-driven pre-screening tool by quantifying each element's intrinsic tendency toward topological behavior. Here, we develop a hybrid quantum-classical neural network (HQCNN) that extends this rule into a quantum-inspired formulation. Within this framework, the HQCNN maps compositional descriptors to quantum probability amplitudes, naturally introducing pairwise inter-element correlations inaccessible to classical heuristics. The physical validity of these correlations is substantiated by constructing an equivalent complex-valued neural network (CVNN), confirming both the consistency and interpretability of the formulation. Retaining the simplicity of chemical reasoning while embedding quantum-native features, our quantum-inspired rule enables efficient and generalizable topological classification. High-throughput screening combined with first-principles (DFT) validation reveals five previously unreported topological compounds, demonstrating the enhanced predictive power and physical insight afforded by quantum-inspired heuristics.

cond-mat.mtrl-sci

Phase Field Study of Exchange Coupling of Hard/Soft Ferrite on Magnetic Permeability

Effective modulation of magnetic permeability plays a vital role in the development of high-performance inductors. Here, phase-field simulations of hard/soft ferrite composites (BaM/NiZn) clarify how exchange coupling and microstructure impact magnetic permeability. We show that particle size, volume fraction, and orientation of the hard phase can effectively control the transition from collinear to non-collinear coupling, with a critical exchange size r_cr approximately 12 nm. Increasing the hard-phase fraction deepens the anisotropy energy well and monotonically suppresses permeability. In contrast, rotating the BaM easy axis to 90 degrees relative to the applied field produces a strong enhancement: at a 10 nm radius and eta = 0.1 volume fraction, the effective permeability can be more than 30 times larger than in the parallel configuration and then saturates for larger particles. This study establishes a microstructure-permeability-based physical framework for designing hard/soft magnetic composite systems.

cond-mat.mtrl-sci

Intrinsic structure of relaxor ferroelectrics from first principles

We develop FIRE-Swap, a first-principles framework for sampling intrinsic compositional structures in complex perovskites with machine-learning interatomic potentials (MLIPs). Using both dedicated and universal MLIPs, we study the relaxor lead magnesium niobate (PMN) and the solid solutions lead zirconate titanate (PZT) and lead strontium titanate (PST). Across MLIP models and exchange-correlation approximations, FIRE-Swap robustly predicts a rock-salt-like chemical order in PMN, which is absent in PZT and PST with the same mixing ratio, consistent with experiments. We further identify in PMN a distinct Nb-cluster morphology. Interconnected, non-coarsened polar nanoregions are found within Nb clusters, providing a mesoscale basis for understanding relaxor ferroelectricity.

cond-mat.mtrl-sci

Quantum Nonlocality under Latency Constraints

Bell inequalities are bounds on the correlations between different parties obeying a local hidden variable theory. Here, "local" refers to spacetime locality: the parties cannot communicate their inputs because they must produce their outputs faster than the speed-of-light delay between them. In other words, the parties must satisfy a certain latency constraint. In this work, we explicitly incorporate spacetime locality into the formulation of Bell inequalities by imposing such a latency constraint. When the latency constraint is sufficiently tight such that no parties can communicate, this becomes a standard Bell scenario. When the latency constraint is relaxed such that a subset of the parties can communicate, we no longer have a Bell scenario, but we can again find a divide between classical and quantum behaviors. Hence, we observe that the classical-quantum gap should actually be a function of time. To study these more general scenarios, we introduce the mathematical framework of latency-constrained games, which models time-evolving input and output processes for spatially separated parties subject to finite communication speeds. This framework allows us to systematically study the weirdness of quantum mechanics in the "low-latency regime" where the speed-of-light delay is non-negligible. Latency-constrained games can describe real-time decision-making in real-world settings that are latency-sensitive, such as high-frequency trading and distributed systems, and can reveal the utility of quantum correlations in these settings.

quant-ph

A combinatorial simplicial cone decomposition

This paper introduces an algebraic combinatorial approach to simplicial cone decompositions, a key step in solving inhomogeneous linear Diophantine systems and counting lattice points in polytopes. We use constant term manipulation on the system \( A\alpha = \mathbf{b} \), where \( A \) is an \( r \times n \) integral matrix and \( \mathbf{b} \) is an integral vector. We establish a relationship between special constant terms and shifted simplicial cones. This leads to the \texttt{SimpCone[S]} algorithm, which efficiently decomposes polyhedra into simplicial cones. Unlike traditional geometric triangulation methods, this algorithm is versatile for many choices of the strategy \( \texttt{S} \) and can also be applied to parametric polyhedra. The algorithm is useful for efficient volume computation of polytopes and can be applied to address various new research projects. Additionally, we apply our framework to unimodular cone decompositions. This extends the effectiveness of the newly developed \texttt{DecDenu} algorithm from denumerant cones to general simplicial cones.

math.CO

Rotation Perturbation Robustness in Point Cloud Analysis: A Perspective of Manifold Distillation

Point cloud is often regarded as a discrete sampling of Riemannian manifold and plays a pivotal role in the 3D image interpretation. Particularly, rotation perturbation, an unexpected small change in rotation caused by various factors (like equipment offset, system instability, measurement errors and so on), can easily lead to the inferior results in point cloud learning tasks. However, classical point cloud learning methods are sensitive to rotation perturbation, and the existing networks with rotation robustness also have much room for improvements in terms of performance and noise tolerance. Given these, this paper remodels the point cloud from the perspective of manifold as well as designs a manifold distillation method to achieve the robustness of rotation perturbation without any coordinate transformation. In brief, during the training phase, we introduce a teacher network to learn the rotation robustness information and transfer this information to the student network through online distillation. In the inference phase, the student network directly utilizes the original 3D coordinate information to achieve the robustness of rotation perturbation. Experiments carried out on four different datasets verify the effectiveness of our method. Averagely, on the Modelnet40 and ScanobjectNN classification datasets with random rotation perturbations, our classification accuracy has respectively improved by 4.92% and 4.41%, compared to popular rotation-robust networks; on the ShapeNet and S3DIS segmentation datasets, compared to the rotation-robust networks, the improvements of mIoU are 7.36% and 4.82%, respectively. Besides, from the experimental results, the proposed algorithm also shows excellent performance in resisting noise and outliers.

cs.CV

PreCM: The Padding-based Rotation Equivariant Convolution Mode for Semantic Segmentation

Semantic segmentation is an important branch of image processing and computer vision. With the popularity of deep learning, various convolutional neural networks have been proposed for pixel-level classification and segmentation tasks. In practical scenarios, however, imaging angles are often arbitrary, encompassing instances such as water body images from remote sensing and capillary and polyp images in the medical domain, where prior orientation information is typically unavailable to guide these networks to extract more effective features. In this case, learning features from objects with diverse orientation information poses a significant challenge, as the majority of CNN-based semantic segmentation networks lack rotation equivariance to resist the disturbance from orientation information. To address this challenge, this paper first constructs a universal convolution-group framework aimed at more fully utilizing orientation information and equipping the network with rotation equivariance. Subsequently, we mathematically design a padding-based rotation equivariant convolution mode (PreCM), which is not only applicable to multi-scale images and convolutional kernels but can also serve as a replacement component for various types of convolutions, such as dilated convolutions, transposed convolutions, and asymmetric convolution. To quantitatively assess the impact of image rotation in semantic segmentation tasks, we also propose a new evaluation metric, Rotation Difference (RD). The replacement experiments related to six existing semantic segmentation networks on three datasets show that, the average Intersection Over Union (IOU) of their PreCM-based versions respectively improve 6.91%, 10.63%, 4.53%, 5.93%, 7.48%, 8.33% compared to their original versions in terms of random angle rotation. And the average RD values are decreased by 3.58%, 4.56%, 3.47%, 3.66%, 3.47%, 3.43% respectively.

cs.CV

DISCO: Embodied Navigation and Interaction via Differentiable Scene Semantics and Dual-level Control

Building a general-purpose intelligent home-assistant agent skilled in diverse tasks by human commands is a long-term blueprint of embodied AI research, which poses requirements on task planning, environment modeling, and object interaction. In this work, we study primitive mobile manipulations for embodied agents, i.e. how to navigate and interact based on an instructed verb-noun pair. We propose DISCO, which features non-trivial advancements in contextualized scene modeling and efficient controls. In particular, DISCO incorporates differentiable scene representations of rich semantics in object and affordance, which is dynamically learned on the fly and facilitates navigation planning. Besides, we propose dual-level coarse-to-fine action controls leveraging both global and local cues to accomplish mobile manipulation tasks efficiently. DISCO easily integrates into embodied tasks such as embodied instruction following. To validate our approach, we take the ALFRED benchmark of large-scale long-horizon vision-language navigation and interaction tasks as a test bed. In extensive experiments, we make comprehensive evaluations and demonstrate that DISCO outperforms the art by a sizable +8.6% success rate margin in unseen scenes, even without step-by-step instructions. Our code is publicly released at https://github.com/AllenXuuu/DISCO.

cs.CV

A Secure and Efficient Distributed Semantic Communication System for Heterogeneous Internet of Things

Semantic communications are expected to improve the transmission efficiency in Internet of Things (IoT) networks. However, the distributed nature of networks and heterogeneity of devices challenge the secure utilization of semantic communication systems. In this paper, we develop a distributed semantic communication system that achieves the security and efficiency during update and usage phases. A blockchain-based trust scheme for update is designed to continuously train and synchronize the system in dynamic IoT environments. To improve the updating efficiency, we propose a flexible semantic coding method base on compressive semantic knowledge bases. It greatly reduces the amount of data shared among devices for system update, and realizes the flexible adjustment of the size of knowledge bases and the number of transmitted signal symbols in model training and inference stages. In the usage phase, a signature mechanism for lossy semantics is introduced to guarantee the integrity and authenticity of the transmitted semantics in lossy semantic communications. We further design a noise-aware differential privacy mechanism, which introduces optimized noise based on the different channel information available to heterogeneous devices. Experiments on text transmission tasks show that the proposed system achieves the protection of the integrity and privacy for exchanged semantics, and reduces the data to be transmitted in the update phase by about $35\%$ to $88\%$, and in the usage phase by $60\%$ compared with related works.

eess.SP

HumanVLA: Towards Vision-Language Directed Object Rearrangement by Physical Humanoid

Physical Human-Scene Interaction (HSI) plays a crucial role in numerous applications. However, existing HSI techniques are limited to specific object dynamics and privileged information, which prevents the development of more comprehensive applications. To address this limitation, we introduce HumanVLA for general object rearrangement directed by practical vision and language. A teacher-student framework is utilized to develop HumanVLA. A state-based teacher policy is trained first using goal-conditioned reinforcement learning and adversarial motion prior. Then, it is distilled into a vision-language-action model via behavior cloning. We propose several key insights to facilitate the large-scale learning process. To support general object rearrangement by physical humanoid, we introduce a novel Human-in-the-Room dataset encompassing various rearrangement tasks. Through extensive experiments and analysis, we demonstrate the effectiveness of the proposed approach.

cs.RO

Algebraic Volume for Polytope Arise from Ehrhart Theory

Volume computation for $d$-polytopes $\mathcal{P}$ is fundamental in mathematics. There are known volume computation algorithms, mostly based on triangulation or signed-decomposition of $\mathcal{P}$. We consider $ \mathrm{cone}(\mathcal{P})$ as a lift of $\mathcal{P}$ in view of Ehrhart theory. By using technique from algebraic combinatorics, we obtain a volume algorithm using only signed simplicial cone decompositions of $ \mathrm{cone}(¶)$. Each cone is associated with a simple algebraic volume formula. Summing them gives the volume of the polytope. Our volume formula applies to various kind of cases. In particular, we use it to explain the traditional triangulation method and Lawrence's signed decomposition method. Moreover, we give a completely new primal-dual method for volume computation. This solves the traditional problem in this area: All existing methods are hopelessly impractical for either the class of simple polytopes or the class of simplicial polytopes. Our method has a good performance in computer experiments.

math.CO

Beyond Object Recognition: A New Benchmark towards Object Concept Learning

Understanding objects is a central building block of artificial intelligence, especially for embodied AI. Even though object recognition excels with deep learning, current machines still struggle to learn higher-level knowledge, e.g., what attributes an object has, and what can we do with an object. In this work, we propose a challenging Object Concept Learning (OCL) task to push the envelope of object understanding. It requires machines to reason out object affordances and simultaneously give the reason: what attributes make an object possesses these affordances. To support OCL, we build a densely annotated knowledge base including extensive labels for three levels of object concept (category, attribute, affordance), and the causal relations of three levels. By analyzing the causal structure of OCL, we present a baseline, Object Concept Reasoning Network (OCRN). It leverages causal intervention and concept instantiation to infer the three levels following their causal relations. In experiments, OCRN effectively infers the object knowledge while following the causalities well. Our data and code are available at https://mvig-rhos.com/ocl.

cs.CV

A polynomial time algorithm for calculating Fourier-Dedekind sums

We solve an open problem proposed in the book ``Computing the continuous discretely" written by Matthias Beck and Sinai Robins. That is, we proposed a polynomial time algorithm for calculating Fourier-Dedekind sums. The algorithm is simple modular Barvinok's simplicial cone decomposition. It can be easily adapted into De Leora et. al.'s LattE package, which gives a nice implimentation of Barvinok's polynomial time algorithm.

math.CO

PAI3D: Painting Adaptive Instance-Prior for 3D Object Detection

3D object detection is a critical task in autonomous driving. Recently multi-modal fusion-based 3D object detection methods, which combine the complementary advantages of LiDAR and camera, have shown great performance improvements over mono-modal methods. However, so far, no methods have attempted to utilize the instance-level contextual image semantics to guide the 3D object detection. In this paper, we propose a simple and effective Painting Adaptive Instance-prior for 3D object detection (PAI3D) to fuse instance-level image semantics flexibly with point cloud features. PAI3D is a multi-modal sequential instance-level fusion framework. It first extracts instance-level semantic information from images, the extracted information, including objects categorical label, point-to-object membership and object position, are then used to augment each LiDAR point in the subsequent 3D detection network to guide and improve detection performance. PAI3D outperforms the state-of-the-art with a large margin on the nuScenes dataset, achieving 71.4 in mAP and 74.2 in NDS on the test split. Our comprehensive experiments show that instance-level image semantics contribute the most to the performance gain, and PAI3D works well with any good-quality instance segmentation models and any modern point cloud 3D encoders, making it a strong candidate for deployment on autonomous vehicles.

cs.CV

Learning to Anticipate Future with Dynamic Context Removal

Anticipating future events is an essential feature for intelligent systems and embodied AI. However, compared to the traditional recognition task, the uncertainty of future and reasoning ability requirement make the anticipation task very challenging and far beyond solved. In this filed, previous methods usually care more about the model architecture design or but few attention has been put on how to train an anticipation model with a proper learning policy. To this end, in this work, we propose a novel training scheme called Dynamic Context Removal (DCR), which dynamically schedules the visibility of observed future in the learning procedure. It follows the human-like curriculum learning process, i.e., gradually removing the event context to increase the anticipation difficulty till satisfying the final anticipation target. Our learning scheme is plug-and-play and easy to integrate any reasoning model including transformer and LSTM, with advantages in both effectiveness and efficiency. In extensive experiments, the proposed method achieves state-of-the-art on four widely-used benchmarks. Our code and models are publicly released at https://github.com/AllenXuuu/DCR.

cs.CV