SearcharxivSearch

arXiv subjects

Xiaokang Wang

Publications and source records attributed to Xiaokang Wang.

9 recordsLinked to original sources

Collapsing constant scalar curvature metrics

We prove that a sequence of constant scalar curvature (CSC) metrics which is collapsing with bounded curvature to a manifold $(X,g_{\infty})$ can be perturbed to a sequence of $\mathcal{N}$-invariant collapsing CSC metrics, under a natural assumption involving the eigenvalues of the drift Laplacian on $(X,g_{\infty})$. This answers a special case of a question of Cheeger-Fukaya-Gromov. We also give some natural conditions on the limiting metric-measure space under which the eigenvalue assumption is automatically satisfied.

math.DG

VLT: A Vision-Language-Time Series Multimodal Foundation Model for Industrial Intelligence

Industrial time series serve as the foundation for Prognostics and Health Management (PHM) to ensure the reliability and safety of industrial equipment such as aero-engines. However, existing approaches are typically limited to single-modality modeling, which restricts their generalization in complex scenarios. Although recent advances in large language models (LLMs) provide new opportunities for multimodal learning, bridging continuous time-series signals and discrete textual semantics remains an open challenge. To this end, we propose VLT, a multimodal foundation model that jointly models time-series, frequency-spectrum visual representations, and textual knowledge. A key insight is to utilize the frequency spectrum as a visual bridge to connect continuous temporal signals with discrete semantics. Specifically, a Time-aware Mixture-of-Experts (Time-MoE) is designed to capture heterogeneous temporal dynamics, while a Frequency-Text Augmented Learner enables joint modeling of spectral and semantic features within a shared representation space. Furthermore, a time-centric gradient alignment mechanism is introduced to mitigate cross-modal optimization conflicts via gradient normalization and reliability-aware dynamic reweighting. Extensive experiments on multiple industrial datasets demonstrate that VLT outperforms state-of-the-art methods, achieving superior robustness and generalization under few-shot, noisy, and incomplete-modality settings.

cs.AI

DCP-Prune: Ultra-Low Token Pruning with Distribution Consistency Preservation

Recent vision token pruning methods effectively preserve model performance under moderate token budgets but become unstable under ultra-low token budget. Our analysis shows that as the pruning budget decreases, accuracy degradation is often accompanied by larger feature distribution shifts. Critically, the degree of this distribution shift strongly correlates with performance degradation. To better characterize this phenomenon, we introduce a lightweight distribution consistency metric to estimate the distribution shift between retained and full tokens. Motivated by these observations, we propose a two-stage pruning framework consisting of Anchor-Context Graph Recovery (ACGR) and Text-Aware Token Cluster Selection (TATCS). Specifically, ACGR transfers contextual information before token removal, while TATCS dynamically re-selects representative tokens when severe distribution shift is detected. Extensive experiments demonstrate that our method achieves superior and more stable performance under ultra-low token budget. Notably, it retains 92.1% of the upper-bound average performance on LLaVA-1.5-7B with only 16 visual tokens.

cs.CV

Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k

Video generation models have achieved remarkable progress in the past year. The quality of AI video continues to improve, but at the cost of larger model size, increased data quantity, and greater demand for training compute. In this report, we present Open-Sora 2.0, a commercial-level video generation model trained for only $200k. With this model, we demonstrate that the cost of training a top-performing video generation model is highly controllable. We detail all techniques that contribute to this efficiency breakthrough, including data curation, model architecture, training strategy, and system optimization. According to human evaluation results and VBench scores, Open-Sora 2.0 is comparable to global leading video generation models including the open-source HunyuanVideo and the closed-source Runway Gen-3 Alpha. By making Open-Sora 2.0 fully open-source, we aim to democratize access to advanced video generation technology, fostering broader innovation and creativity in content creation. All resources are publicly available at: https://github.com/hpcaitech/Open-Sora.

cs.GR

Pluriclosed flow on Oeljeklaus-Toma manifolds

We establish global existence of the pluriclosed flow with arbitrary initial data on Oeljeklaus-Toma manifolds, and Gromov-Hausdorff convergence of blowdown limits to a torus under natural conjectural bounds on the flow at infinity. In the case of generalized Kähler-Ricci flow we prove refined a priori estimates in support of these conjectural bounds.

math.DG

On the structure of locally conformally flat orbifolds and ALE manifolds

In this paper, we prove several structure theorems for locally conformally flat, positive Yamabe orbifolds and nonnegative scalar curvature, ALE manifolds. These two kinds of spaces can be related by conformal blow-up and conformal compactification. For the orbifolds, we prove that such orbifolds admit a manifold cover. For the ALE manifolds, the homomorphism of the fundamental group for the ALE space induced by the embedding of the ALE end is always injective. Using these properties, several classifications of such ALE manifolds and orbifolds are given in low dimensions. As an application to the moduli space, we prove that the football orbifold $\mathbb{S}^4/Γ$ cannot be realized as the Gromov-Hausdorff limit. In addition, we prove the positive mass theorem of these ALE ends and give a simple proof for the optimal decay rate. Using the positive mass theorem, we can solve the orbifold Yamabe problem in the locally conformally flat case.

math.DG

Size-Dependent Lattice Pseudosymmetry for Frustrated Decahedral Nanoparticles

Geometric frustration is a widespread phenomenon in physics, materials science, and biology, occurring when the geometry of a system prevents local interactions from being all accommodated. The resulting manifold of nearly degenerate configurations can lead to complex collective behaviors and emergent pseudosymmetry in diverse systems such as frustrated magnets, mechanical metamaterials, and protein assemblies. In synthetic multi-twinned nanomaterials, similar pseudosymmetric features have also been observed and manifest as intrinsic lattice strain. Despite extensive interest in the stability of these nanostructures, a fundamental understanding remains limited due to the lack of detailed structural characterization across varying sizes and geometries. In this work, we apply four-dimensional scanning transmission electron microscopy strain mapping over a total of 23 decahedral nanoparticles with edge lengths, d, between 20 and 55 nm. From maps of full 2D strain tensor at nanometer spatial resolution, we reveal the prevalence of heterogeneity in different modes of lattice distortions, which homogenizes and restores symmetry with increasing size. Knowing the particle crystallography, we reveal distinctive spatial patterns of local lattice phase transformation between face-centered cubic and body-centered tetragonal symmetries, with a contrast between particles below and above d of 35 nm. The results suggest a cross-over size of the internal structure occurs, as particles shape transition from modified-Wulff shape favored at nanoscale to faceted, pentagonal bipyramidal shape. Ultimately, our 4D-STEM mapping provides new insight to long-standing mysteries of this historic system and can be widely applicable to study nanocrystalline solids and material phase transformation that are important in catalysis, metallurgy, electronic devices, and energy storage materials.

cond-mat.mtrl-sci

RGB-D Indiscernible Object Counting in Underwater Scenes

Recently, indiscernible/camouflaged scene understanding has attracted lots of research attention in the vision community. We further advance the frontier of this field by systematically studying a new challenge named indiscernible object counting (IOC), the goal of which is to count objects that are blended with respect to their surroundings. Due to a lack of appropriate IOC datasets, we present a large-scale dataset IOCfish5K which contains a total of 5,637 high-resolution images and 659,024 annotated center points. Our dataset consists of a large number of indiscernible objects (mainly fish) in underwater scenes, making the annotation process all the more challenging. IOCfish5K is superior to existing datasets with indiscernible scenes because of its larger scale, higher image resolutions, more annotations, and denser scenes. All these aspects make it the most challenging dataset for IOC so far, supporting progress in this area. Benefiting from the recent advancements of depth estimation foundation models, we construct high-quality depth maps for IOCfish5K by generating pseudo labels using the Depth Anything V2 model. The RGB-D version of IOCfish5K is named IOCfish5K-D. For benchmarking purposes on IOCfish5K, we select 14 mainstream methods for object counting and carefully evaluate them. For multimodal IOCfish5K-D, we evaluate other 4 popular multimodal counting methods. Furthermore, we propose IOCFormer, a new strong baseline that combines density and regression branches in a unified framework and can effectively tackle object counting under concealed scenes. We also propose IOCFormer-D to enable the effective usage of depth modality in helping detect and count objects hidden in their environments. Experiments show that IOCFormer and IOCFormer-D achieve state-of-the-art scores on IOCfish5K and IOCfish5K-D, respectively.

cs.CV

Diversified and Compatible Web APIs Recommendation in IoT

With the ever-increasing popularity of Service-oriented Architecture (SoA) and Internet of Things (IoT), a considerable number of enterprises or organizations are attempting to encapsulate their provided complex business services into various lightweight and accessible web APIs (application programming interfaces) with diverse functions. In this situation, a software developer can select a group of preferred web APIs from a massive number of candidates to create a complex mashup economically and quickly based on the keywords typed by the developer. However, traditional keyword-based web API search approaches often suffer from the following difficulties and challenges. First, they often focus more on the functional matching between the candidate web APIs and the mashup to be developed while neglecting the compatibility among different APIs, which probably returns a group of incompatible web APIs and further leads to a mashup development failure. Second, existing approaches often return a web API composition solution to the mashup developer for reference, which narrows the developer's API selection scope considerably and may reduce developer satisfaction heavily. In view of the above challenges and successful application of game theory in the IoT, based on the idea of game theory, we propose a compatible and diverse web APIs recommendation approach for mashup creations, named MCCOMP+DIV, to return multiple sets of diverse and compatible web APIs with higher success rate. Finally, we validate the effectiveness and efficiency of MCCOMP+DIV through a set of experiments based on a real-world web API dataset, i.e., the PW dataset crawled from ProgrammableWeb.com.

cs.SI