SearcharxivSearch

arXiv subjects

Jiaxuan Liu

Publications and source records attributed to Jiaxuan Liu.

At least 19 recordsLinked to original sources

A Catalogue of Topological Moiré Bands in Twisted Semiconductors

Twisted two-dimensional semiconductors provide a route to flat and topological moiré minibands, but systematic principles for organizing their material dependence have remained unclear. Here, we establish a high-throughput framework that integrates structural relaxation, first-principles electronic structure calculations, and moiré band topology. We apply this framework to 43 experimentally realized monolayers and 91 symmetry-inequivalent bilayer prototypes, yielding over 1,000 angle-resolved moiré electronic band structures. This database reveals that the low-energy moiré electronic structure is organized primarily by the valley character of the parent band edge together with stacking symmetry. In $Γ$-valley systems, the miniband width usually follows a nearly quadratic twist-angle scaling, consistent with a folding-dominated kinetic-energy scale. In $K$-valley systems, stacking-controlled interlayer hybridization governs whether parent Berry curvature is redistributed into isolated valley Chern minibands. By contrast, $M$-valley systems form a more material-specific class associated with anisotropic and symmetry-constrained band folding. The same valley-and-stacking hierarchy rationalizes the emergence or suppression of $\mathbb{Z}_2$ minibands, and surface termination in Janus bilayers provides a microscopic knob for changing the relevant valley character. These results establish a materials-level organizing principle for designing flat and topological moiré bands in twisted semiconductors.

cond-mat.mtrl-sci

AppAgent: Multimodal Agents as Smartphone Users

Recent advancements in large language models (LLMs) have led to the creation of intelligent agents capable of performing complex tasks. This paper introduces a novel LLM-based multimodal agent framework designed to operate smartphone applications. Our framework enables the agent to operate smartphone applications through a simplified action space, mimicking human-like interactions such as tapping and swiping. This novel approach bypasses the need for system back-end access, thereby broadening its applicability across diverse apps. Central to our agent's functionality is its innovative learning method. The agent learns to navigate and use new apps either through autonomous exploration or by observing human demonstrations. This process generates a knowledge base that the agent refers to for executing complex tasks across different applications. To demonstrate the practicality of our agent, we conducted extensive testing over 50 tasks in 10 different applications, including social media, email, maps, shopping, and sophisticated image editing tools. The results affirm our agent's proficiency in handling a diverse array of high-level tasks.

cs.CV

A Geometric Design Principle for $\mathbb{Z}_2$ Topological Phases in Twisted Triangular-Lattice Bilayers

Twisted van der Waals bilayers provide a versatile platform for moiré electronic states, yet a transferable symmetry-based principle for time-reversal-invariant $\mathbb{Z}_2$ moiré bands has remained largely missing. Here we show that triangular-lattice bilayers with symmetry-related stacking minima provide a geometric route to an emergent honeycomb moiré lattice. Band-edge states derived from the untwisted $Γ$ valley are trapped by the reconstructed stacking landscape, forming A/B moiré orbitals whose inter-domain coupling generates Dirac crossings. Spin--orbit coupling opens a topological gap, yielding an effective Kane--Mele description and a quantum spin Hall phase characterized by a nontrivial $\mathbb{Z}_2$ invariant. First-principles calculations for Janus BiTeBr confirm the robustness of this phase over a broad twist-angle range and demonstrate an electric-field-driven topological transition. Representative triangular-lattice bilayers further establish this symmetry-based design principle as a broadly applicable route to tunable moiré quantum spin Hall materials.

cond-mat.mtrl-sci

RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild

Accurate document layout analysis remains a critical bottleneck for document parsing systems, due to the intricate coupling among heterogeneous document layout elements, geometric distortions (\eg, paper warping and bending, perspective variations), and reading order within diverse layout structures. Existing approaches typically rely on fragmented multi-stage pipelines or computationally heavy generative Transformer architectures, leading to error propagation and limited efficiency. In this paper, we present RT-DocLayout, a highly efficient end-to-end framework for document layout analysis, designed as a front-end for document parsing tasks. The proposed model unifies classification, detection, pixel-level segmentation, and reading order prediction for layout elements within a single 33M-parameter architecture. Built upon the RT-DETR, our key contribution is a unified multi-task formulation within a single query-based decoder that simultaneously classifies, regresses bounding box, generates masks, and constructs relationship to reason reading order. By jointly learning geometric and structural representations, RT-DocLayout introduces multi-task optimization that substantially improves robustness under real-world document distortions. Extensive experiments on public benchmarks demonstrate state-of-the-art performance in document layout analysis while maintaining real-time inference speed(132.1 FPS). When coupled with downstream OCR engines, RT-DocLayout significantly improves full-document reconstruction quality, providing a scalable and practical foundation for real-world document intelligence systems.

cs.CV

PP-OCRv6: From 1.5M to 34.5M Parameters, Surpassing Billion-Scale VLMs on OCR Tasks

Vision-Language Models (VLMs) have achieved impressive results on general vision-language tasks, yet they suffer from hallucination, imprecise localization, and prohibitive computational cost when applied to dedicated OCR scenarios. This paper presents PP-OCRv6, a lightweight OCR system that combines architectural innovation with data-centric optimization. PP-OCRv6 redesigns the backbone, detection neck, and recognition neck around a unified MetaFormer-style building block with structural reparameterization, decoupling spatial token mixing from channel mixing and supporting both tasks through task-specific stride configurations. Three model tiers (medium, small, tiny) share the same block primitives, covering deployment scenarios from server to edge. On our in-house benchmarks, PP-OCRv6_medium achieves 83.2% recognition accuracy and 86.2% detection Hmean, outperforming PP-OCRv5_server by +5.1% and +4.6% respectively while surpassing Qwen3-VL-235B, GPT-5.5, and Gemini-3.1-Pro with orders of magnitude fewer parameters. The tiny tier achieves 3.9$\times$ faster inference than PP-OCRv5_mobile on Intel Xeon CPU while maintaining comparable accuracy.

cs.CV

PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training

We introduce PaddleOCR-VL-1.6, an upgraded compact document parsing model built upon PaddleOCR-VL-1.5. Although PaddleOCR-VL-1.5 establishes a strong 0.9B baseline, its remaining errors concentrate in under-optimized regions where model behavior is unstable, data coverage is sparse, or supervision is unreliable. Rather than expanding the training corpus indiscriminately, PaddleOCR-VL-1.6 introduces a region-aware data optimization framework that identifies weak regions from the previous model, applies targeted enhancement to these regions, and improves the reliability of supervision signals. It further adopts a progressive post-training recipe based on curated data selection and reinforcement learning, pushing model performance to a higher level through staged optimization. PaddleOCR-VL-1.6 achieves a new state-of-the-art score of 96.33% on OmniDocBench v1.6, demonstrates strong competitiveness against top-tier VLMs, and provides a practical post-training recipe for the PaddleOCR-VL series.

cs.CV

Building and Road Recognition in Dense Urban Informal Settlements: A Dataset and Benchmark

As a widespread form of informal settlements, urban villages present significant challenges for sustainable urban development and governance. Precise mapping of their infrastructure is essential, however, existing remote sensing datasets primarily focus on formal urban environments, lacking fine-grained annotated data for the high-density building patterns and narrow road networks typical of urban villages. To address this gap, we introduce the \textit{DenseUIS} dataset, the first high-resolution remote sensing dataset specifically designed for building and road extraction in extremely dense urban informal settlements, covering 126 urban villages across Shenzhen and Guangzhou in China. Furthermore, we conduct a comprehensive evaluation of state-of-the-art deep learning models on this dataset. Experimental results reveal the limitations of existing methods in handling the unique morphological patterns of dense informal settlements, underscoring the need for specialized approaches. \textit{DenseUIS} therefore provides a robust benchmark for advancing fine-grained urban mapping in complex and high-density informal environments. The dataset is publicly available at https://github.com/rui-research/DenseUIS.

cs.CV

Towards Active Real-to-Twin Inspection: A New Paradigm for Zero-Shot Anomaly Detection

The deployment of zero-shot anomaly detection (AD) in embodied industrial inspection is severely bottlenecked by its reliance on passive, fixed-viewpoint 2D imagery. Such formulations inherently fail to accommodate the active, dynamic observations required in real-world environments. To break this limitation, we introduce Real-to-Twin Anomaly Detection, a novel task that evaluates physical observations directly against geometrically matched CAD Digital Twins. To tackle this new task, we propose AVATAR, a framework designed to learn robust semantic alignment between Real and Digital Twins. By bridging benign Sim2Real domain gaps using only defect-free pairs, AVATAR effectively transforms CAD priors into dynamic, anomaly-free references. This elegant formulation enables the model to localize diverse anomalies in a zero-shot manner as unalignable deviations, eliminating the need for defect annotations. Extensive experiments demonstrate that AVATAR substantially outperforms adapted state-of-the-art baselines, exhibiting exceptional robustness to severe viewpoint variations. The code and dataset will be made publicly available.

cs.CV

A superconvergent hybridizable discontinuous Galerkin method for the convective Cahn--Hilliard equation

We propose a hybridizable discontinuous Galerkin (HDG) method combined with convex-concave splitting for the temporal discretization of the convective Cahn-Hilliard equation. The convection term is discretized explicitly without stabilization, yielding three key advantages: (1) unconditional stability, (2) preservation of the optimal convergence rate for piecewise constant approximations, and (3) a symmetric system after local elimination, enabling efficient solver via minimal residual methods. We establish optimal convergence rates in the $L^2$ norm for both the scalar and flux variables for any polynomial degree $k \geq 0$. To achieve optimal $L^2$-norm estimates, we introduce a specialized HDG elliptic projection operator and analyze its approximation properties. Within the HDG framework, local elimination is employed to reduce the degrees of freedom associated with the globally coupled unknowns, and the scalar variables exhibit superconvergence. Finally, numerical experiments validate the theoretical convergence rates and demonstrate the effectiveness of the proposed method.

math.NA

PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing

We introduce PaddleOCR-VL-1.5, an upgraded model achieving a new state-of-the-art (SOTA) accuracy of 94.5% on OmniDocBench v1.5. To rigorously evaluate robustness against real-world physical distortions, including scanning, skew, warping, screen-photography, and illumination, we propose the Real5-OmniDocBench benchmark. Experimental results demonstrate that this enhanced model attains SOTA performance on the newly curated benchmark. Furthermore, we extend the model's capabilities by incorporating seal recognition and text spotting tasks, while remaining a 0.9B ultra-compact VLM with high efficiency. Code: https://github.com/PaddlePaddle/PaddleOCR

cs.CV

Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing

Document parsing is a fine-grained task where image resolution significantly impacts performance. While advanced research leveraging vision-language models benefits from high-resolution input to boost model performance, this often leads to a quadratic increase in the number of vision tokens and significantly raises computational costs. We attribute this inefficiency to substantial visual regions redundancy in document images, like background. To tackle this, we propose PaddleOCR-VL, a novel coarse-to-fine architecture that focuses on semantically relevant regions while suppressing redundant ones, thereby improving both efficiency and performance. Specifically, we introduce a lightweight Valid Region Focus Module (VRFM) which leverages localization and contextual relationship prediction capabilities to identify valid vision tokens. Subsequently, we design and train a compact yet powerful 0.9B vision-language model (PaddleOCR-VL-0.9B) to perform detailed recognition, guided by VRFM outputs to avoid direct processing of the entire large image. Extensive experiments demonstrate that PaddleOCR-VL achieves state-of-the-art performance in both page-level parsing and element-level recognition. It significantly outperforms existing solutions, exhibits strong competitiveness against top-tier VLMs, and delivers fast inference while utilizing substantially fewer vision tokens and parameters, highlighting the effectiveness of targeted coarse-to-fine parsing for accurate and efficient document understanding. The source code and models are publicly available at https://github.com/PaddlePaddle/PaddleOCR.

cs.CV

PP-OCRv5: A Specialized 5M-Parameter Model Rivaling Billion-Parameter Vision-Language Models on OCR Tasks

The advent of "OCR 2.0" and large-scale vision-language models (VLMs) has set new benchmarks in text recognition. However, these unified architectures often come with significant computational demands, challenges in precise text localization within complex layouts, and a propensity for textual hallucinations. Revisiting the prevailing notion that model scale is the sole path to high accuracy, this paper introduces PP-OCRv5, a meticulously optimized, lightweight OCR system with merely 5 million parameters. We demonstrate that PP-OCRv5 achieves performance competitive with many billion-parameter VLMs on standard OCR benchmarks, while offering superior localization precision and reduced hallucinations. The cornerstone of our success lies not in architectural expansion but in a data-centric investigation. We systematically dissect the role of training data by quantifying three critical dimensions: data difficulty, data accuracy, and data diversity. Our extensive experiments reveal that with a sufficient volume of high-quality, accurately labeled, and diverse data, the performance ceiling for traditional, efficient two-stage OCR pipelines is far higher than commonly assumed. This work provides compelling evidence for the viability of lightweight, specialized models in the large-model era and offers practical insights into data curation for OCR. The source code and models are publicly available at https://github.com/PaddlePaddle/PaddleOCR.

cs.CV

FunCineForge: A Unified Dataset Toolkit and Model for Zero-Shot Movie Dubbing in Diverse Cinematic Scenes

Movie dubbing is the task of synthesizing speech from scripts conditioned on video scenes, requiring accurate lip sync, faithful timbre transfer, and proper modeling of character identity and emotion. However, existing methods face two major limitations: (1) high-quality multimodal dubbing datasets are limited in scale, suffer from high word error rates, contain sparse annotations, rely on costly manual labeling, and are restricted to monologue scenes, all of which hinder effective model training; (2) existing dubbing models rely solely on the lip region to learn audio-visual alignment, which limits their applicability to complex live-action cinematic scenes, and exhibit suboptimal performance in lip sync, speech quality, and emotional expressiveness. To address these issues, we propose FunCineForge, which comprises an end-to-end production pipeline for large-scale dubbing datasets and an MLLM-based dubbing model designed for diverse cinematic scenes. Using the pipeline, we construct the first Chinese television dubbing dataset with rich annotations, and demonstrate the high quality of these data. Experiments across monologue, narration, dialogue, and multi-speaker scenes show that our dubbing model consistently outperforms SOTA methods in audio quality, lip sync, timbre transfer, and instruction following. Code and demos are available at https://anonymous.4open.science/w/FunCineForge.

cs.CV

PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model

In this report, we propose PaddleOCR-VL, a SOTA and resource-efficient model tailored for document parsing. Its core component is PaddleOCR-VL-0.9B, a compact yet powerful vision-language model (VLM) that integrates a NaViT-style dynamic resolution visual encoder with the ERNIE-4.5-0.3B language model to enable accurate element recognition. This innovative model efficiently supports 109 languages and excels in recognizing complex elements (e.g., text, tables, formulas, and charts), while maintaining minimal resource consumption. Through comprehensive evaluations on widely used public benchmarks and in-house benchmarks, PaddleOCR-VL achieves SOTA performance in both page-level document parsing and element-level recognition. It significantly outperforms existing solutions, exhibits strong competitiveness against top-tier VLMs, and delivers fast inference speeds. These strengths make it highly suitable for practical deployment in real-world scenarios. Code is available at https://github.com/PaddlePaddle/PaddleOCR .

cs.CV

Chern-Selective multi-valley Flat Bands in Twisted Mono-Bilayer and Mono-Trilayer MoTe$_2$

The interplay between moiré flat bands originating from different valleys can give rise to a variety of exotic quantum phases. In this work, we investigate the electronic properties of twisted mono-bilayer (A-AB) and mono-trilayer (A-ABA) MoTe$_2$ using first-principles calculations and continuum models. Unlike previous studies on twisted bilayer systems, in which low-energy flat bands originate solely from the $K/K'$ valleys, in A-AB and A-ABA twisted MoTe$_2$ (\tmt) the moiré bands at low energies arise from both the $Γ$ and $K/K'$ valleys, with spin Chern numbers $C_s=0$ (for $Γ$) and $C_{\uparrow/\downarrow}=\pm1$ (for $K/K'$), respectively. We show that the multi-valley moiré flat bands are governed by interlayer-hybridization effects, and that different stacking configurations and thicknesses tune the relative energy alignment between the $Γ$ and $K$ valley moiré flat bands. By constructing valley-resolved continuum models and performing Wannierization for the low-energy moiré bands, we further uncover that the Berry curvature and quantum metric distributions can be effectively tuned by the layer number and stacking configuration. Unlike other moiré systems, where only one kind of valley influenced the low energy physics, the simultaneous appearance of two distinct types of valleys, with different symmetries, establish A-AB and A-ABA \tmt\ as ideal platforms for studying layer-controlled multi-valley physics.

cond-mat.mtrl-sci

UDDETTS: Unifying Discrete and Dimensional Emotions for Controllable Emotional Text-to-Speech

Recent large language models (LLMs) have made great progress in the field of text-to-speech (TTS), but they still face major challenges in synthesizing fine-grained emotional speech in an interpretable manner. Traditional methods rely on discrete emotion labels to control emotion categories and intensities, which cannot capture the complexity and continuity of human emotional perception and expression. The lack of large-scale emotional speech datasets with balanced emotion distributions and fine-grained emotional annotations often causes overfitting in synthesis models and impedes effective emotion control. To address these issues, we propose UDDETTS, a universal LLM framework unifying discrete and dimensional emotions for controllable emotional TTS. This model introduces the interpretable Arousal-Dominance-Valence (ADV) space for dimensional emotion description and supports emotion control driven by either discrete emotion labels or nonlinearly quantified ADV values. Furthermore, a semi-supervised training strategy is designed to comprehensively utilize diverse speech datasets with different types of emotional annotations to train the UDDETTS. Experiments show that UDDETTS achieves linear emotion control along three interpretable dimensions, and exhibits superior end-to-end emotional speech synthesis capabilities. Code and demos are available at: https://anonymous.4open.science/w/UDDETTS.

cs.LG

VASPilot: MCP-Facilitated Multi-Agent Intelligence for Autonomous VASP Simulations

Density-functional-theory (DFT) simulations with the Vienna Ab initio Simulation Package (VASP) are indispensable in computational materials science but often require extensive manual setup, monitoring, and postprocessing. Here, we introduce VASPilot, an open-source platform that fully automates VASP workflows via a multi-agent architecture built on the CrewAI framework and a standardized Model Context Protocol (MCP). VASPilot's agent suite handles every stage of a VASP study-from retrieving crystal structures and generating input files to submitting Slurm jobs, parsing error messages, and dynamically adjusting parameters for seamless restarts. A lightweight Flask-based web interface provides intuitive task submission, real-time progress tracking, and drill-down access to execution logs, structure visualizations, and plots. We validate VASPilot on both routine and advanced benchmarks: automated band-structure and density-of-states calculations (including on-the-fly symmetry corrections), plane-wave cutoff convergence tests, lattice-constant optimizations with various van der Waals corrections, and cross-material band-gap comparisons for transition-metal dichalcogenides. In all cases, VASPilot completed the missions reliably and without manual intervention. Moreover, its modular design allows easy extension to other DFT codes simply by deploying the appropriate MCP server. By offloading technical overhead, VASPilot enables researchers to focus on scientific discovery and accelerates high-throughput computational materials research.

cond-mat.mtrl-sci

Now and Future of Artificial Intelligence-based Signet Ring Cell Diagnosis: A Survey

Signet ring cells (SRCs), associated with a high propensity for peripheral metastasis and poor prognosis, critically influence surgical decision-making and outcome prediction. However, their detection remains challenging even for experienced pathologists. While artificial intelligence (AI)-based automated SRC diagnosis has gained increasing attention for its potential to enhance diagnostic efficiency and accuracy, existing methodologies lack systematic review. This gap impedes the assessment of disparities between algorithmic capabilities and clinical applicability. This paper presents a comprehensive survey of AI-driven SRC analysis from 2008 through June 2025. We systematically summarize the biological characteristics of SRCs and challenges in their automated identification. Representative algorithms are analyzed and categorized as unimodal or multi-modal approaches. Unimodal algorithms, encompassing image, omics, and text data, are reviewed; image-based ones are further subdivided into classification, detection, segmentation, and foundation model tasks. Multi-modal algorithms integrate two or more data modalities (images, omics, and text). Finally, by evaluating current methodological performance against clinical assistance requirements, we discuss unresolved challenges and future research directions in SRC analysis. This survey aims to assist researchers, particularly those without medical backgrounds, in understanding the landscape of SRC analysis and the prospects for intelligent diagnosis, thereby accelerating the translation of computational algorithms into clinical practice.

eess.IV