SearcharxivSearch

arXiv subjects

Xin Nie

Publications and source records attributed to Xin Nie.

At least 19 recordsLinked to original sources

iFLYTEK-Embodied-Omni Technical Report

General-purpose embodied agents must understand multimodal instructions, anticipate how their environment will evolve, and produce precise control actions over extended horizons. Existing approaches typically specialize in visual-language reasoning, video-based world modeling, or action generation, while cascaded pipelines that first synthesize future observations and then infer actions can introduce interface bottlenecks and compound prediction errors. We present iFLYTEK-Embodied-Omni, a unified multimodal foundation model that jointly models vision(videos and images), language, and action within a single Omni framework. Its modality-specific visual-language, video-generation, and action-generation components communicate through shared multimodal self-attention. This design establishes brain-cerebellum collaboration: the vision-language modeland video generation model form a high-level brain for instruction understanding, task planning, progress tracking, and future visual-state prediction, whereas the action generation modelserves as a low-level cerebellum that directly converts planned subgoals and shared multimodal context into executable action chunks. To develop these capabilities, we combine action-annotated and action-free embodied videos from human demonstrations and robot interactions with embodied reasoning, embodied perception, and general-purpose image-text data to construct a comprehensive dataset. We further adopt a four-stage strategy that progressively trains the VLM, VGM, and AGM before jointly fine-tuning the complete model.

cs.AI

GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation

Vision-Language-Action (VLA) models achieve strong benchmark performance but still struggle in real-world deployment with unseen objects, background shifts, and different robot embodiments. We argue that this stems from the lack of a unified geometry-aware manipulation representation, leaving existing VLAs vulnerable to low-level trajectory supervision, misaligned 3D features, and embodiment differences. To address this, we propose GEAR-VLA, a VLA framework for learning unified geometry-aware action representations for generalizable robotic manipulation. GEAR-VLA adopts coarse-to-fine action learning, where multi-source embodied pretraining equips the VLM with embodied reasoning and discrete action understanding before latent action tokens connect action semantics to a gradient-decoupled DiT continuous action expert. It further performs semantic-aligned 3D integration by aligning a trainable 3D spatial backbone with the VLA representation while freezing the original VLM-aligned visual pathway. To share this representation across robots, GEAR-VLA uses embodiment canonicalization, where embodiment-aware states and embodiment-invariant actions confine robot differences to the low-level interface. Extensive simulation and real-world experiments demonstrate strong generalization: GEAR-VLA achieves state-of-the-art performance on LIBERO, zero-shot LIBERO-Plus, and RoboTwin 2.0, reaches 85.9% success on AgileX and 81.0% on the pretraining-unseen LDT-01 embodiment, and obtains 90.1% success on a 6,360-trial universal grasping benchmark with 212 unseen objects. Code and models will be released at https://github.com/babynabeauty/GEAR-VLA.

cs.RO

SFMP: Fine-Grained, Hardware-Friendly and Search-Free Mixed-Precision Quantization for Large Language Models

Mixed-precision quantization is a promising approach for compressing large language models under tight memory budgets. However, existing mixed-precision methods typically suffer from one of two limitations: they either rely on expensive discrete optimization to determine precision allocation, or introduce hardware inefficiencies due to irregular memory layouts. We propose SFMP, a search-free and hardware-friendly mixed-precision quantization framework for large language models. The framework is built upon four novel ideas: Fractional bit-width, which extends integer bit-width for weight matrix to fractional value and transforms discrete precision allocation as a continuous problem; 2)Block-wise mixed-precision, enabling fine-grained precision within weight matrices while remaining hardware-friendly; 3)Row-column weight reordering, which aggregates salient weights via row and column reordering, incurring only a small activation reordering overhead during inference; 4)Unified GEMM kernel, which supports mixed-precision GEMM at arbitrary average bit-width. Extensive experiments demonstrate that SFMP outperforms state-of-the-art layer-wise mixed-precision methods under the same memory constraints, while significantly reducing quantization cost and improving inference efficiency. Code is available at https://github.com/Nkniexin/SFMP

cs.LG

Towards Automatic Evaluation and High-Quality Pseudo-Parallel Dataset Construction for Audio Editing: A Human-in-the-Loop Method

Audio editing aims to manipulate audio content based on textual descriptions, supporting tasks such as adding, removing, or replacing audio events. Despite recent progress, the lack of high-quality benchmark datasets and comprehensive evaluation metrics remains a major challenge for both assessing audio editing quality and improving the task itself. In this work, we propose a novel approach for audio editing task by incorporating expert knowledge into both the evaluation and dataset construction processes: 1) First, we establish AuditScore, the first comprehensive dataset for subjective evaluation of audio editing, consisting of over 6,300 edited samples generated from 7 representative audio editing frameworks and 23 system configurations. Each sample is annotated by professional raters on three key aspects of audio editing quality: overall Quality, Relevance to editing intent, and Faithfulness to original features. 2) Based on this dataset, we systematically propose AuditEval, a family of automatic MOS-style evaluators tailored for audio editing, covering both SSL-based and LLM-based approaches. It addresses the lack of effective objective metrics and the prohibitive cost of subjective evaluation in this field. 3) We further leverage AuditEval to evaluate and filter a large amount of synthetically mixed editing pairs, mining a high-quality pseudo-parallel subset by selecting the most plausible samples. Comprehensive experiments validate that our expert-informed filtering strategy effectively yields higher-quality data, while also exposing the limitations of traditional objective metrics and the advantages of AuditEval. The dataset, codes and tools can be found at: https://github.com/NKU-HLT/AuditEval.

cs.SD

ELUTQ: Optimizing Quantization Accuracy under LUT-Based Computation for Edge LLMs

Weight quantization effectively reduces memory consumption and enable the deployment of Large Language Models on edge devices, yet existing hardware-friendly methods often rely on uniform quantization, which suffers from poor weight-distribution fitting and high dequantization overhead under low-bit settings. In this paper, we propose ELUTQ, an efficient quantization framework featuring a novel quantization format termed Hierarchical Linear Quantization (HLQ). HLQ is designed to better capture the statistical characteristics of weights and eliminate dequantization overhead using Bit-serial LUT-based GEMM operations. HLQ significantly improves model accuracy under low-bit settings and achieves performance comparable to QAT methods without any retraining of the weights. Moreover, an optimized quantization pipeline is integrated into ELUTQ, enabling it to complete the quantization of LLaMA 3.1-70B using only 64 GB of CPU memory and 48 GB of VRAM, reducing the hardware requirements for large-scale model quantization. To enable efficient deployment on edge devices, ELUTQ designs high-performance kernels to support end-to-end inference. Our 2-bit LLaMA3.1-8B achieves 1.5x speedup over AWQ on RTX 3090. Code is available at https://github.com/Nkniexin/ELUTQ.

cs.LG

Spin angular momentum transfer in the Einstein-de Haas effect

We investigate spin angular momentum transfer in the Einstein-de Haas effect within prototypical magnetic crystals, focusing on its partition between phonons and rigid-body rotation. Using the Eckart frame to decouple local vibrations (phonons) from rigid-body rotation, we demonstrate that spin angular momentum is simultaneously transferred into both phonons and rigid-body rotation in an asymmetric way: rigid-body rotation acquires the dominant share of angular momentum, while phonons absorb most of the resulting kinetic energy. This divergent transfer of angular momentum and energy identifies phonons as direct and indispensable participants in the Einstein-de Haas dynamics. Furthermore, we find that pseudo-dipolar anisotropy and Dzyaloshinskii-Moriya interaction exert distinct control over the angular momentum transfer. Stronger pseudo-dipolar anisotropy increases the total amount of transferred angular momentum, whereas stronger Dzyaloshinskii-Moriya interaction accelerates the transfer rate and increases the proportion of phonon angular momentum. Our work clarifies the microscopic picture of the Einstein-de Haas effect and enables targeted angular-momentum control in magneto-mechanical devices.

cond-mat.mes-hall

Three-dimensional Damage Visualization of Civil Structures via Gaussian Splatting-enabled Digital Twins

Recent advancements in civil infrastructure inspections underscore the need for precise three-dimensional (3D) damage visualization on digital twins, transcending traditional 2D image-based damage identifications. Compared to conventional photogrammetric 3D reconstruction techniques, modern approaches such as Neural Radiance Field (NeRF) and Gaussian Splatting (GS) excel in scene representation, rendering quality, and handling featureless regions. Among them, GS stands out for its efficiency, leveraging discrete anisotropic 3D Gaussians to represent radiance fields, unlike NeRF's continuous implicit model. This study introduces a GS-enabled digital twin method tailored for effective 3D damage visualization. The method's key contributions include: 1) utilizing GS-based 3D reconstruction to visualize 2D damage segmentation results while reducing segmentation errors; 2) developing a multi-scale reconstruction strategy to balance efficiency and damage detail; 3) enabling digital twin updates as damage evolves over time. Demonstrated on an open-source synthetic dataset for post-earthquake inspections, the proposed approach offers a promising solution for comprehensive 3D damage visualization in civil infrastructure digital twins.

cs.CV

iFlyBot-VLM Technical Report

We introduce iFlyBot-VLM, a general-purpose Vision-Language Model (VLM) used to improve the domain of Embodied Intelligence. The central objective of iFlyBot-VLM is to bridge the cross-modal semantic gap between high-dimensional environmental perception and low-level robotic motion control. To this end, the model abstracts complex visual and spatial information into a body-agnostic and transferable Operational Language, thereby enabling seamless perception-action closed-loop coordination across diverse robotic platforms. The architecture of iFlyBot-VLM is systematically designed to realize four key functional capabilities essential for embodied intelligence: 1) Spatial Understanding and Metric Reasoning; 2) Interactive Target Grounding; 3) Action Abstraction and Control Parameter Generation; 4) Task Planning and Skill Sequencing. We envision iFlyBot-VLM as a scalable and generalizable foundation model for embodied AI, facilitating the progression from specialized task-oriented systems toward generalist, cognitively capable agents. We conducted evaluations on 10 current mainstream embodied intelligence-related VLM benchmark datasets, such as Blink and Where2Place, and achieved optimal performance while preserving the model's general capabilities. We will publicly release both the training data and model weights to foster further research and development in the field of Embodied Intelligence.

cs.RO

Bayesian Fully-Connected Tensor Network for Hyperspectral-Multispectral Image Fusion

Tensor decomposition is a powerful tool for data analysis and has been extensively employed in the field of hyperspectral-multispectral image fusion (HMF). Existing tensor decomposition-based fusion methods typically rely on disruptive data vectorization/reshaping or impose rigid constraints on the arrangement of factor tensors, hindering the preservation of spatial-spectral structures and the modeling of cross-dimensional correlations. Although recent advances utilizing the Fully-Connected Tensor Network (FCTN) decomposition have partially alleviated these limitations, the process of reorganizing data into higher-order tensors still disrupts the intrinsic spatial-spectral structure. Furthermore, these methods necessitate extensive manual parameter tuning and exhibit limited robustness against noise and spatial degradation. To alleviate these issues, we propose the Bayesian FCTN (BFCTN) method. Within this probabilistic framework, a hierarchical sparse prior that characterizing the sparsity of physical elements, establishes connections between the factor tensors. This framework explicitly models the intrinsic physical coupling among spatial structures, spectral signatures, and local scene homogeneity. For model learning, we develop a parameter estimation method based on Variational Bayesian inference (VB) and the Expectation-Maximization (EM) algorithm, which significantly reduces the need for manual parameter tuning. Extensive experiments demonstrate that BFCTN not only achieves state-of-the-art fusion accuracy and strong robustness but also exhibits practical applicability in complex real-world scenarios.

cs.CV

Einstein-de Haas effect: a bridge linking mechanics, magnetism, and topology

The Einstein-de Haas (EdH) effect is a fascinating phenomenon that links mechanics and magnetism. Despite being discovered over a century ago, it remains significant in contemporary science, particularly within the fields of spintronics and ultrafast magnetism. Recent predictions suggest that the EdH effect may be realized in topological magnon systems, potentially leading to even richer properties. In this perspective, we introduce recent advancements in the EdH effect and discuss its developments in three key aspects: the microscopic mechanism, its manifestation in topological systems, and chirality-selective magnon-phonon coupling. Our discussions aim to inspire further explorations of the EdH effect and highlight its promising applications in different areas.

cond-mat.mes-hall

A spin-rotation mechanism of Einstein-de Haas effect based on a ferromagnetic disk

Spin-rotation coupling (SRC) is a fundamental phenomenon that connects electronic spins with the rotational motion of a medium. We elucidate the Einstein-de Haas (EdH) effect and its inverse with SRC as the microscopic mechanism using the dynamic spin-lattice equations derived by elasticity theory and Lagrangian formalism. By applying the coupling equations to an iron disk in a magnetic field, we exhibit the transfer of angular momentum and energy between spins and lattice, with or without damping. The timescale of the angular momentum transfer from spins to the entire lattice is estimated by our theory to be on the order of 0.01 ns, for the disk with a radius of 100 nm. Moreover, we discover a linear relationship between the magnetic field strength and the rotation frequency, which is also enhanced by a higher ratio of Young's modulus to Poisson's coefficient. In the presence of damping, we notice that the spin-lattice relaxation time is nearly inversely proportional to the magnetic field. Our explorations will contribute to a better understanding of the EdH effect and provide valuable insights for magneto-mechanical manufacturing.

cond-mat.mes-hall

MLIP: Enhancing Medical Visual Representation with Divergence Encoder and Knowledge-guided Contrastive Learning

The scarcity of annotated data has sparked significant interest in unsupervised pre-training methods that leverage medical reports as auxiliary signals for medical visual representation learning. However, existing research overlooks the multi-granularity nature of medical visual representation and lacks suitable contrastive learning techniques to improve the models' generalizability across different granularities, leading to the underutilization of image-text information. To address this, we propose MLIP, a novel framework leveraging domain-specific medical knowledge as guiding signals to integrate language information into the visual domain through image-text contrastive learning. Our model includes global contrastive learning with our designed divergence encoder, local token-knowledge-patch alignment contrastive learning, and knowledge-guided category-level contrastive learning with expert knowledge. Experimental evaluations reveal the efficacy of our model in enhancing transfer performance for tasks such as image classification, object detection, and semantic segmentation. Notably, MLIP surpasses state-of-the-art methods even with limited annotated data, highlighting the potential of multimodal pre-training in advancing medical representation learning.

cs.CV

Boundary metric of Epstein-Penner convex hull and discrete conformality

The Epstein-Penner convex hull construction associates to every decorated punctured hyperbolic surface a polyhedral convex body in the Minkowski space. It works in the de Sitter and anti-de Sitter spaces as well. In these three spaces, the quotient of the spacelike boundary part of the convex body has an induced Euclidean, spherical and hyperbolic metric, respectively, with conical singularities. We show that this gives a bijection from the decorated Teichmüller space to a moduli space of such metrics in the Euclidean and hyperbolic cases, as well as a bijection between specific subspaces of them in the spherical case. Moreover, varying the decoration of a fixed hyperbolic surface corresponds to a discrete conformal change of the metric. This gives a new $3$-dimensional interpretation of discrete conformality which is in a sense inverse to the Bobenko-Pinkall-Springborn interpretation.

math.GT

On circle patterns and spherical conical metrics

The Koebe-Andreev-Thurston circle packing theorem, as well as its generalization to circle patterns due to Bobenko and Springborn, holds for Euclidean and hyperbolic metrics possibly with conical singularities, but fails for spherical metrics because of the non-uniqueness coming from Möbius transformations. In this paper, we show that a unique existence result for circle pattern with spherical conical metric holds if one prescribes the geodesic total curvature of each circle instead of the cone angles.

math.DG

Large genus asymptotics for lengths of separating closed geodesics on random surfaces

In this paper, we investigate basic geometric quantities of a random hyperbolic surface of genus $g$ with respect to the Weil-Petersson measure on the moduli space $\mathcal{M}_g$. We show that as $g$ goes to infinity, a generic surface $X\in \mathcal{M}_g$ satisfies asymptotically: (1) the separating systole of $X$ is about $2\log g$; (2) there is a half-collar of width about $\frac{\log g}{2}$ around a separating systolic curve of $X$; (3) the length of shortest separating closed multi-geodesics of $X$ is about $2\log g$. As applications, we also discuss the asymptotic behavior of the extremal separating systole, the non-simple systole and the expectation value of lengths of shortest separating closed multi-geodesics as $g$ goes to infinity.

math.GT

Cyclic Higgs bundles and minimal surfaces in pseudo-hyperbolic spaces

We introduce a type of minimal surface in the pseudo-hyperbolic space $\mathbb{H}^{n,n}$ (with $n$ even) or $\mathbb{H}^{n+1,n-1}$ (with $n$ odd) associated to cyclic $\mathrm{SO}_0(n,n+1)$-Higg bundles. By establishing the infinitesimal rigidity of these surfaces, we get a new proof, for $\mathrm{SO}_0(n,n+1)$, of Labourie's theorem that the holonomy map restricts to an immersion on the cyclic locus of Hitchin base, and extend it to Collier's components. This implies Labourie's former conjecture in the case of the exceptional group $G_2'$, for which we also show that these minimal surfaces are $\boldsymbol{J}$-holomorphic curves of a particular type in the almost complex $\mathbb{H}^{4,2}$.

math.DG

Hypersurfaces of constant Gauss-Kronecker curvature with Li-normalization in affine space

For convex hypersurfaces in the affine space $\mathbb{A}^{n+1}$ ($n\geq2$), A.-M.\ Li introduced the notion of $α$-normal field as a generalization of the affine normal field. By studying a Monge-Ampère equation with gradient blowup boundary condition, we show that regular domains in $\mathbb{A}^{n+1}$, defined with respect to a proper convex cone and satisfying some regularity assumption if $n\geq3$, are foliated by complete convex hypersurfaces with constant Gauss-Kronecker curvature relative to the Li-normalization. When $n=2$, a key feature is that no regularity assumption is required, and the result extends our recent work about the $α=1$ case.

math.DG

Affine deformations of quasi-divisible convex cones

For any subgroup of $\mathrm{SL}(3,\mathbb{R})\ltimes\mathbb{R}^3$ obtained by adding a translation part to a subgroup of $\mathrm{SL}(3,\mathbb{R})$ which is the fundamental group of a finite-volume convex projective surface, we first show that under a natural condition on the translation parts of parabolic elements, the affine action of the group on $\mathbb{R}^3$ has convex domains of discontinuity that are regular in a certain sense, generalizing a result of Mess for globally hyperbolic flat spacetimes. We then classify all these domains and show that the quotient of each of them is an affine manifold foliated by convex surfaces with constant affine Gaussian curvature. The proof is based on a correspondence between the geometry of an affine space endowed with a convex cone and the geometry of a convex tube domain. As an independent result, we show that the moduli space of such groups is a vector bundle over the moduli space of finite-volume convex projective structures, with rank equaling the dimension of the Teichmüller space.

math.DG