SearcharxivSearch

arXiv subjects

Xiaoyan Liu

Publications and source records attributed to Xiaoyan Liu.

At least 19 recordsLinked to original sources

Dream4D: Lifting Camera-Controlled I2V towards Spatiotemporally Consistent 4D Generation

The synthesis of spatiotemporally coherent 4D content presents fundamental challenges in computer vision, requiring simultaneous modeling of high-fidelity spatial representations and physically plausible temporal dynamics. Current approaches often struggle to maintain view consistency while handling complex scene dynamics, particularly in large-scale environments with multiple interacting elements. This work introduces Dream4D, a novel framework that bridges this gap through a synergy of controllable video generation and neural 4D reconstruction. Our approach seamlessly combines a two-stage architecture: it first predicts optimal camera trajectories from a single image using few-shot learning, then generates geometrically consistent multi-view sequences via a specialized pose-conditioned diffusion process, which are finally converted into a persistent 4D representation. This framework is the first to leverage both rich temporal priors from video diffusion models and geometric awareness of the reconstruction models, which significantly facilitates 4D generation and shows higher quality (e.g., mPSNR, mSSIM) over existing methods.

cs.CV

Streaming4D: Accelerate 4D World Models via Block-wise Video Generation and Incremental Reconstruction

Current 4D generation paradigms are often bottlenecked by a sequential decoupling design: video is generated first, followed by 3D reconstruction, leading to high interaction latency. This limits applications in interactive real-time scenarios. To this end, we propose \textbf{Streaming4D}, a tightly coupled synchronous pipeline that integrates block-wise autoregressive video generation with incremental 3D reconstruction. Unlike traditional frame-by-frame emission and delayed geometry recovery, Streaming4D generates temporal video blocks and immediately triggers reconstruction for each completed block, enabling parallel execution between synthesis and geometric updates. This approach allows the world representation to evolve online with the video stream, reducing feedback latency while preserving geometric fidelity. We instantiate \textbf{Streaming4D} using a Self-Forcing-style autoregressive generator and an incremental reconstruction backend. Experiments show consistent runtime improvements across resolutions on a single RTX 4090 (1.24$\times$ speedup), while maintaining high-quality 4D geometry and multi-view consistency.

cs.CV

A substrate booster for P-type 2D ferromagnetic semiconductor

Spin transistors with its both charge and spin properties tuned via electrostatic gating are believed capable for widespread use, which however have proven challenging due to the extreme rareness of their physical base -- magnetic semiconductors. The latter are limited within very few systems including diluted magnetic semiconductors (DMS) and two-dimensional ferromagnetic semiconductors (2D-FMS), and known to suffer from inadequate gate-tunability of their electric and/or magnetic properties. Here, we show a substrate engineering paradigm by interfacing few-layered Cr$_{2}$Ge$_{2}$Te$_{6}$ (FL-CGT) with an antiferromagnetic insulator CrOCl. Owing to the subtle interfacial charge transfer couplings, CGT can be drastically turned from an ambipolar semiconductor into a high performance P-type semiconductor. When cooled below the Curie temperature, the ON-OFF ratio in such substrate-boosted FMS field-effect transistor (FET) reaches 10$^{5}$ with its coercive field $H_{c}$ of magnetic hysteresis loop tunable by a factor of more than 200$\%$, enabling {gate-assisted magnetic switching in the prototype semiconducting spin transistor architecture}. A crossover from critical power-law scaling to a dual power-law behaviour under heavy hole doping was further observed. Our findings {signify} an efficient interfacial charge transfer and electrically modulated magnetic anisotropy energy supported by calculations. This high performance P-type FMS-FET system suggests that active substrate-boosting paradigm might be a powerful path for the investigation of future gate-tunable spintronic devices.

cond-mat.mes-hall

High-bandwidth photodetector enabled by frequency-domain equalization

High-bandwidth germanium (Ge) photodetectors are crucial for silicon photonic integrated circuits. However, their bandwidth is restricted by carrier transit time and parasitic parameters. In this work, we propose an equalization photodetector (EqPD) utilizing the frequency response of a high-bandwidth photodetector PDA to subtract the frequency response of a low-bandwidth photodetector PDB. With the response of PDB attenuating more severely than PDA at high frequency, the differential frequency response (the response of EqPD) can get higher values at high frequency than at low frequency, flattening the overall frequency response and expanding the bandwidth. Experimental results show that the EqPD has a bandwidth exceeding 110 GHz with a responsivity of 94 mA/W. A 100 Gbaud non-return-to-zero (NRZ) operation without digital signal processing is also demonstrated. To the best of our knowledge, this represents the highest bandwidth in a vertical Ge photodetector, providing a promising solution for high-speed photodetection in next-generation optical communication.

physics.optics

TISC: A Text-Driven Image Semantic Communication System for Faithful Reconstruction

Generative image semantic communication converts an image into a text description and then performs text-to-image reconstruction at the receiver via diffusion-based generative models. This paradigm has attracted broad attention due to its extremely low bandwidth cost. However, existing methods still face two critical bottlenecks across image-to-text (I2T) semantic extraction at the transmitter and text-to-image (T2I) semantic reconstruction at the receiver: (i) semantic loss and distortion in I2T, where holistic image descriptions may omit fine-grained object attributes and spatial-position information, causing the generated text to deviate from the original image semantics; and (ii) insufficient semantic faithfulness in T2I, where even with the same semantically faithful text description, different initial noise settings may lead diffusion-based reconstruction to produce images with different levels of semantic consistency with the original image. These issues jointly limit the semantic faithfulness of image reconstruction. To address them, we propose TISC, a text-driven image semantic communication framework tailored for faithful reconstruction. TISC incorporates two key designs: (1) Tree-Structured Attribute Semantic Extraction (TSASE), which decomposes semantic extraction into global scene, background, and object-level attribute descriptions, covering spatial position, shape/pose, color, material, and other physical attributes for each detected object; and (2) an Initial Noise Optimization (INO) mechanism, which selects an initial noise seed at the transmitter according to a comprehensive similarity score that jointly considers visual and semantic consistency. Experiments on multiple datasets show that TSASE improves object-position recovery and semantic description faithfulness, while the INO parameter study supports the adopted configuration for noise selection.

cs.CV

Symmetry-Selective Strain Control of Anisotropic Magnetic Response in a Silicon FinFET Double Quantum Dot

Strain naturally develops in three-dimensional quantum-dot structures such as silicon FinFETs during fabrication and cooling. Such strain becomes especially important in a double quantum dot, because the two dots can experience different local strain and therefore acquire different magnetic responses. To understand how this dot-to-dot strain difference affects coupled hole spins, we theoretically study the local \(g\) tensors of a silicon FinFET double quantum dot by combining a three-dimensional Poisson--Schrödinger calculation based on a six-band \(k\!\cdot\!p\) model with configuration interaction. We find that the effect of strain depends on both its tensor component and its spatial symmetry. For the diagonal components \(ε_{yy}\) and \(ε_{zz}\), strain mainly changes the principal \(g\) values, with only a small opening of the maximum-response axes. In contrast, the shear component \(ε_{yz}\) can also change the orientation of the local magnetic response. When the strain profile preserves the transverse mirror symmetry, the shear-induced rotation is strongly suppressed. Breaking this local constraint permits a pronounced off-diagonal response and rotates the principal magnetic axes. The same component- and symmetry-selected trends appear in a Zeeman-only calculation, showing that the valence-band Zeeman coupling is sufficient to generate them, while the full Hamiltonian determines their quantitative expression. Together, these results show how the tensor component and spatial symmetry of strain can be used to control both the magnitude and orientation of the magnetic response in coupled hole-spin qubits.

cond-mat.mes-hall

Domain Walls Stabilized by Intrinsic Phonon Modes and Engineered Defects Enable Robust Ferroelectricity in HfO2

Ferroelectric $\mathrm{HfO}_2$ has attracted extensive research interest for its applications in AI era. The domain walls play a crucial role in phase structure stabilization and polarization switching of ferroelectric $\mathrm{HfO}_2$, however, a thorough understanding is still lacking. Here, we developed a unified framework based on phonon mode expansion to systematically study the effects of phonon modes and defects on domain wall structures. Using this approach combined with first-principle calculations, we revealed that the interface phonon modes play a key role in stability of domain walls; defects pin and stabilize ferroelectric domains, which in turn stabilizes the metastable orthorhombic phase and facilitates polarization switching. This provides an insight from the microscopic physics origin into the enhanced ferroelectricity in $\mathrm{HfO}_2$ by doping and defect engineering. Furthermore, the theoretically predicted domain structures and defect distributions were observed in La-doped $\mathrm{HfO}_2$ ferroelectric films by EELS and STEM experiments, which confirms the validity of our findings.

cond-mat.mtrl-sci

Observation of magnetically coupled electro-optic effect in LiNbO3/LiTaO3 at room temperature

The magnetoelectric coupling effect serves as a crucial bridge between electrical and magnetic order parameters in condensed matter physics, forming the physical basis for the development of next-generation low-power information storage and sensing technologies. However, material systems exhibiting this effect at room temperature are extremely rare, and the coupling strength is typically very weak, which has long hindered the practical application of such phenomena despite their rich tunability. Here, we break this by reporting a pronounced magnetoelectric phenomenon,the magnetically coupled EO effect in the classic ferroelectric optical materials LiNbO3 and LiTaO3. We trace its origin to an unexpected source,the problematic direct current drift,a major reliability issue in photonic integrated circuits. We unambiguously demonstrate that this drift stems not from mobile ions but from defect-bound unpaired electrons, whose slow polarization relaxation is quenched upon the magnetic-field-induced formation of a room-temperature skyrmion states, as directly visualized by Lorentz transmission electron microscopy. This collective spin ordering not only solves the decades-old drift problem but also transforms the defect states into a magnetically responsive platform, exhibiting an efficiency up to 34 pm/V in LiNbO3 and 15 pm/V in LiTaO3-an orders-of-magnitude enhancement over conventional room-temperature magnetoelectric responses and even surpassing the materials' intrinsic Pockels coefficients. We also measured the magnetic response of this additional electro-optical effect and found that under the condition of applying a magnetic field of 0.1 T, an electro-optical coefficient adjustment of about 8 pm/V can be achieved (0.008 pm/V/Oe), which is equivalent to 30% of the electro-optical coefficient of LiNbO3 itself.

cond-mat.mtrl-sci

Deep Learning Accelerated First-Principles Quantum Transport Simulations at Nonequilibrium State

The non-equilibrium Green's function method combined with density functional theory (NEGF-DFT) provides a rigorous framework for simulating nanoscale electronic transport, but its computational cost scales steeply with system size. Recent artificial intelligence (AI) approaches have sought to accelerate such simulations, yet most rely on conventional machine learning, lack atomic resolution, struggle to extrapolate to larger systems, and cannot predict multiple properties simultaneously. Here we introduce DeepQT, a deep-learning framework that integrates graph neural networks with transformer architectures to enable multi-property predictions of electronic structure and transport without manual feature engineering. By learning key intermediate quantities of NEGF-DFT, the equilibrium Hamiltonian and the non-equilibrium total potential difference, DeepQT reconstructs Hamiltonians under both equilibrium and bias conditions, yielding accurate transport predictions. Leveraging the principle of electronic nearsightedness, DeepQT generalizes from small training systems to much larger ones with high fidelity. Benchmarks on graphene, MoS2, and silicon diodes with varied defects and dopants show that DeepQT achieves first-principles accuracy while reducing computational cost by orders of magnitude. This scalable, transferable framework advances AI-assisted quantum transport, offering a powerful tool for next-generation nanoelectronic device design.

cond-mat.mes-hall

SA-OOSC: A Multimodal LLM-Distilled Semantic Communication Framework for Enhanced Coding Efficiency with Scenario Understanding

This paper introduces SA-OOSC, a multimodal large language models (MLLM)-distilled semantic communication framework that achieves efficient semantic coding with scenario-aware importance allocations. This approach addresses a critical limitation of existing object-oriented semantic communication (OOSC) systems - assigning static importance values to specific classes of objects regardless of their contextual relevance. Our framework utilizes MLLMs to identify the scenario-augmented (SA) semantic importance for objects within the image. Through knowledge distillation with the MLLM-annotated data, our vectorization/de-vectorization networks and JSCC encoder/decoder learn to dynamically allocate coding resources based on contextual significance, i.e., distinguishing between high-importance objects and low-importance according to the SA scenario information of the task. The framework features three core innovations: a MLLM-guided knowledge distillation pipeline, an importance-weighted variable-length JSCC framework, and novel loss function designs that facilitate the knowledge distillation within the JSCC framework. Experimental validation demonstrates our framework's superior coding efficiency over conventional semantic communication systems, with open-sourced MLLM-annotated and human-verified datasets established as new benchmarks for future research in semantic communications.

eess.SP

IndusGCC: A Data Benchmark and Evaluation Framework for GUI-Based General Computer Control in Industrial Automation

As Industry 4.0 progresses, flexible manufacturing has become a cornerstone of modern industrial systems, with equipment automation playing a pivotal role. However, existing control software for industrial equipment, typically reliant on graphical user interfaces (GUIs) that require human interactions such as mouse clicks or screen touches, poses significant barriers to the adoption of code-based equipment automation. Recently, Large Language Model-based General Computer Control (LLM-GCC) has emerged as a promising approach to automate GUI-based operations. However, industrial settings pose unique challenges, including visually diverse, domain-specific interfaces and mission-critical tasks demanding high precision. This paper introduces IndusGCC, the first dataset and benchmark tailored to LLM-GCC in industrial environments, encompassing 448 real-world tasks across seven domains, from robotic arm control to production line configuration. IndusGCC features multimodal human interaction data with the equipment software, providing robust supervision for GUI-level code generation. Additionally, we propose a novel evaluation framework with functional and structural metrics to assess LLM-generated control scripts. Experimental results on mainstream LLMs demonstrate both the potential of LLM-GCC and the challenges it faces, establishing a strong foundation for future research toward fully automated factories. Our data and code are publicly available at: \href{https://github.com/Golden-Arc/IndustrialLLM}{https://github.com/Golden-Arc/IndustrialLLM.

eess.SY

UMRE: A Unified Monotonic Transformation for Ranking Ensemble in Recommender Systems

Industrial recommender systems commonly rely on ensemble sorting (ES) to combine predictions from multiple behavioral objectives. Traditionally, this process depends on manually designed nonlinear transformations (e.g., polynomial or exponential functions) and hand-tuned fusion weights to balance competing goals -- an approach that is labor-intensive and frequently suboptimal in achieving Pareto efficiency. In this paper, we propose a novel Unified Monotonic Ranking Ensemble (UMRE) framework to address the limitations of traditional methods in ensemble sorting. UMRE replaces handcrafted transformations with Unconstrained Monotonic Neural Networks (UMNN), which learn expressive, strictly monotonic functions through the integration of positive neural integrals. Subsequently, a lightweight ranking model is employed to fuse the prediction scores, assigning personalized weights to each prediction objective. To balance competing goals, we further introduce a Pareto optimality strategy that adaptively coordinates task weights during training. UMRE eliminates manual tuning, maintains ranking consistency, and achieves fine-grained personalization. Experimental results on two public recommendation datasets (Kuairand and Tenrec) and online A/B tests demonstrate impressive performance and generalization capabilities.

cs.IR

An Underwater, Fault-Tolerant, Laser-Aided Robotic Multi-Modal Dense SLAM System for Continuous Underwater In-Situ Observation

Existing underwater SLAM systems are difficult to work effectively in texture-sparse and geometrically degraded underwater environments, resulting in intermittent tracking and sparse mapping. Therefore, we present Water-DSLAM, a novel laser-aided multi-sensor fusion system that can achieve uninterrupted, fault-tolerant dense SLAM capable of continuous in-situ observation in diverse complex underwater scenarios through three key innovations: Firstly, we develop Water-Scanner, a multi-sensor fusion robotic platform featuring a self-designed Underwater Binocular Structured Light (UBSL) module that enables high-precision 3D perception. Secondly, we propose a fault-tolerant triple-subsystem architecture combining: 1) DP-INS (DVL- and Pressure-aided Inertial Navigation System): fusing inertial measurement unit, doppler velocity log, and pressure sensor based Error-State Kalman Filter (ESKF) to provide high-frequency absolute odometry 2) Water-UBSL: a novel Iterated ESKF (IESKF)-based tight coupling between UBSL and DP-INS to mitigate UBSL's degeneration issues 3) Water-Stereo: a fusion of DP-INS and stereo camera for accurate initialization and tracking. Thirdly, we introduce a multi-modal factor graph back-end that dynamically fuses heterogeneous sensor data. The proposed multi-sensor factor graph maintenance strategy efficiently addresses issues caused by asynchronous sensor frequencies and partial data loss. Experimental results demonstrate Water-DSLAM achieves superior robustness (0.039 m trajectory RMSE and 100\% continuity ratio during partial sensor dropout) and dense mapping (6922.4 points/m^3 in 750 m^3 water volume, approximately 10 times denser than existing methods) in various challenging environments, including pools, dark underwater scenes, 16-meter-deep sinkholes, and field rivers. Our project is available at https://water-scanner.github.io/.

cs.RO

Out-of-Distribution Detection in Heterogeneous Graphs via Energy Propagation

Graph neural networks (GNNs) are proven effective in extracting complex node and structural information from graph data. While current GNNs perform well in node classification tasks within in-distribution (ID) settings, real-world scenarios often present distribution shifts, leading to the presence of out-of-distribution (OOD) nodes. OOD detection in graphs is a crucial and challenging task. Most existing research focuses on homogeneous graphs, but real-world graphs are often heterogeneous, consisting of diverse node and edge types. This heterogeneity adds complexity and enriches the informational content. To the best of our knowledge, OOD detection in heterogeneous graphs remains an underexplored area. In this context, we propose a novel methodology for OOD detection in heterogeneous graphs (OODHG) that aims to achieve two main objectives: 1) detecting OOD nodes and 2) classifying all ID nodes based on the first task's results. Specifically, we learn representations for each node in the heterogeneous graph, calculate energy values to determine whether nodes are OOD, and then classify ID nodes. To leverage the structural information of heterogeneous graphs, we introduce a meta-path-based energy propagation mechanism and an energy constraint to enhance the distinction between ID and OOD nodes. Extensive experimental findings substantiate the simplicity and effectiveness of OODHG, demonstrating its superiority over baseline models in OOD detection tasks and its accuracy in ID node classification.

cs.LG

Cellular-X: An LLM-empowered Cellular Agent for Efficient Base Station Operations

This paper introduces Cellular-X, an LLM-powered agent designed to automate cellular base station (BS) maintenance. Leveraging multimodal LLM and retrieval-augmented generation (RAG) techniques, Cellular-X significantly enhances field engineer efficiency by quickly interpreting user intents, retrieving relevant technical information, and configuring a BS through iterative self-correction. Key features of the demo include automatic customized BS setup, document-based query answering, and voice-controlled configuration reporting and revision. We implemented Cellular-X on a USRP X310 testbed for demonstration. Demo videos and implementation details are available at https://github.com/SeaBreezing/Cellular-X.

cs.NI

Deep Learning Based 3D Segmentation: A Survey

3D segmentation is a fundamental and challenging problem in computer vision with applications in autonomous driving and robotics. It has received significant attention from the computer vision, graphics and machine learning communities. Conventional methods for 3D segmentation, based on hand-crafted features and machine learning classifiers, lack generalization ability. Driven by their success in 2D computer vision, deep learning techniques have recently become the tool of choice for 3D segmentation tasks. This has led to an influx of many methods in the literature that have been evaluated on different benchmark datasets. Whereas survey papers on RGB-D and point cloud segmentation exist, there is a lack of a recent in-depth survey that covers all 3D data modalities and application domains. This paper fills the gap and comprehensively surveys the recent progress in deep learning-based 3D segmentation techniques. We cover over 220 works from the last six years, analyze their strengths and limitations, and discuss their competitive results on benchmark datasets. The survey provides a summary of the most commonly used pipelines and finally highlights promising research directions for the future.

cs.CV

Research on an Autonomous UAV Search and Rescue System Based on the Improved

The demand is to solve the issue of UAV (unmanned aerial vehicle) operating autonomously and implementing practical functions such as search and rescue in complex unknown environments. This paper proposes an autonomous search and rescue UAV system based on an EGO-Planner algorithm, which is improved by innovative UAV body application and takes the methods of inverse motor backstepping to enhance the overall flight efficiency of the UAV and miniaturization of the whole machine. At the same time, the system introduced the EGO-Planner planning tool, which is optimized by a bidirectional A* algorithm along with an object detection algorithm. It solves the issue of intelligent obstacle avoidance and search and rescue. Through the simulation and field verification work, and compared with traditional algorithms, this method shows more efficiency and reliability in the task. In addition, due to the existing algorithm's improved robustness, this application shows good prospection.

cs.RO

INSPIRIT: Optimizing Heterogeneous Task Scheduling through Adaptive Priority in Task-based Runtime Systems

As modern HPC computing platforms become increasingly heterogeneous, it is challenging for programmers to fully leverage the computation power of massive parallelism offered by such heterogeneity. Consequently, task-based runtime systems have been proposed as an intermediate layer to hide the complex heterogeneity from the application programmers. The core functionality of these systems is to realize efficient task-to-resource mapping in the form of Directed Acyclic Graph (DAG) scheduling. However, existing scheduling schemes face several drawbacks to determine task priorities due to the heavy reliance on domain knowledge or failure to efficiently exploit the interaction of application and hardware characteristics. In this paper, we propose INSPIRIT, an efficient and lightweight scheduling framework with adaptive priority designed for task-based runtime systems. INSPIRIT introduces two novel task attributes \textit{inspiring ability} and \textit{inspiring efficiency} for dictating scheduling, eliminating the need for application domain knowledge. In addition, INSPIRIT jointly considers runtime information such as ready tasks in worker queues to guide task scheduling. This approach exposes more performance opportunities in heterogeneous hardware at runtime while effectively reducing the overhead for adjusting task priorities. Our evaluation results demonstrate that INSPIRIT achieves superior performance compared to cutting edge scheduling schemes on both synthesized and real-world task DAGs.

cs.DC