SearcharxivSearch

arXiv subjects

Xiangyu Meng

Publications and source records attributed to Xiangyu Meng.

At least 19 recordsLinked to original sources

An Efficient Out-of-Core Tomographic Imaging Framework for Edge Devices

Computed Tomography (CT) is an essential 3D imaging technology widely used in medical diagnostics and scientific research. However, performing CT imaging on edge devices is challenging due to limitations in computational power, memory capacity, and energy budget. This paper presents an efficient CT reconstruction framework, called edgeFBP, designed for Nvidia Jetson System-on-Chip (SoC) devices. edgeFBP adopts an end-to-end pipeline design for efficient out-of-core image reconstruction under tight power and memory constraints. edgeFBP utilizes a mixed-precision strategy leveraging half-precision Tensor Cores (TCs) to accelerate the bottleneck back-projection (BP) kernel. edgeFBP achieves a 1.83x speedup over the widely used RTK library on Jetson Nano and a 2.56x speedup on Jetson AGX. Under a strict 25-Watt power budget, edgeFBP on Jetson Nano achieves up to 5-48x higher energy efficiency than an Nvidia DGX A100, enabling datacenter-scale imaging on constrained edge devices.

cs.DC

Scalable High-Fidelity Macromolecular Docking for GPU-Accelerated Supercomputers

Flexible macromolecular docking offers high-fidelity predictions of biomolecular interactions, but remains prohibitively expensive at scale. Among existing approaches, LightDock leverages Glowworm Swarm Optimization (GSO) for accuracy, yet suffers from limited parallelism, irregular computation, and severe load imbalance, preventing efficient execution on GPU supercomputers. We present SparkleDock, a scalable GSO-based docking framework enabling near-real-time flexible docking. We redesign GSO to expose massive fine-grained parallelism at the glowworm-agent level, and restructure the dominant energy scoring computation into a Tensor Core-compatible formulation, enabling efficient execution of irregular pairwise interactions through structured matrix operations. We further introduce a performance-model-driven scheduling for load balancing and out-of-core scaling across GPUs. SparkleDock achieves 9.7 $\times$ and 18.9 $\times$ speedups over LightDock on single A100 and H100 GPU, and delivers over two orders of magnitude acceleration at scale. On 512 GPUs, it reduces docking time from hours to seconds, enabling large-scale, high-fidelity virtual screening previously impractical with flexible docking.

cs.DC

Multiple Vehicles and Traction Network Interaction System Stability Analysis and Oscillation Responsibility Identification

The electrical incompatibility between vehicles and traction network in railway system can result in system instability and oscillation overvoltage issues. To analyze the system stability, impedance-based frequency-domain methods are commonly used. However, the current impedance-based modeling methods face challenges in practical implementation due to the requirement of precise analytical models and detailed internal parameters for all vehicles. Moreover, multiple vehicles operate simultaneously in railway systems, each with different operating conditions and internal parameters, thereby influencing system stability to different extents. Therefore, it is crucial to accurately identify the critical vehicles to prevent resonance accidents. To address these challenges, a component connection-based modeling approach for the railway vehicle-grid system is proposed, which only requires the measured impedance results without the internal information of vehicles. In addition, a multilevel sensitivity analysis method is introduced to quantitatively identify the critical vehicles and internal parameters that influence system stability, which outperforms traditional sensitivity analysis methods in computational complexity. Furthermore, a system-level electrical compatibility test process for the railway vehicle-grid system is provided, incorporating the proposed stability and sensitivity analysis methods. Finally, case studies based on the real-world train schedule of a multivehicle-accessed railway vehicle-grid system are designed to verify the correctness of the proposed method.

eess.SY

Neural Network-Based Impedance Identification and Stability Analysis for Double-Sided Feeding Railway Systems

The double-sided power supply railway system increases the simultaneous operation of vehicles on the grid, potentially causing system instability and oscillation overvoltage issues. As vehicles frequently switch operating points during operation, it is essential to analyze system stability across a wide range of conditions. Therefore, accurately identifying the black-box impedance of vehicle converters at multiple operating points is crucial for studying railway vehicle-grid system stability. However, traditional impedance identification methods require extensive data and lack interpretability, leading to significant computational and data burdens. This study introduces an interpretable residual feedforward neural network (ResFNN) combined with SHapley Additive exPlanations for training vehicle impedance models, reducing data requirements while maintaining accuracy. Additionally, a component connection method is proposed for deriving the impedance matrix of a multivehicle railway system under the double-sided feeding mode. This method incorporates the dynamic mobility of vehicles and their positional distribution, and it utilizes the ResFNN to identify impedance for stability analysis. Real operational data from actual railway lines is used as case study to analyze the stability of the double-sided power supply railway system. The results demonstrate that this approach accurately assesses both lowfrequency and high-frequency instability issues.

eess.SY

A Flow Model for the Electrified Railway-Power Grid Hybrid Asymmetric Coupled System and its Linearized Method

In mountainous regions where traction loads constitute a significant portion of a long-chain weak power grid (PG) with sustainable energy, the interaction between the traction power supply system and the PG becomes increasingly evident. The integrated power flow calculation (PFC) method and its linearized model are quite important for the PG - traction network (TN) joint planning. However, existing research on the port load characteristics of the EMUs and the connection angle characteristics of traction transformers is insufficient, and there is a lack of effective methods for PFC or linearized PFC in systems that couple the PG with the traction network. To fill this gap, this paper proposes an integrated PFC model for the AT TN - PG coupled system, along with a linearized method. Firstly, according to the relationship of the phases between the PG and the AT traction network, the node admittance matrix of the coupled system has been constructed. Then, the issue of power injection equations being unable to deal with the EMUs port load is resolved by merging the contact line node and the rail node. Subsequently, the integrated PFC equations for the coupling system are established. Next, a hybrid phase linear decoupled power flow model for the coupling system is developed, employing the correspondence between the phases of the PG and the TN, as well as the phase angle differences among various nodes and branches. Numerical simulations conducted in a specific region demonstrate the necessity of an integrated PFC for the coupled system and validate both the accuracy and efficiency of the linearized model.

eess.SY

Dissipativity-Based Multiport Stability Root-Cause Identification and Mitigation for Solid-State Transformers

For solid-state transformers (SSTs) in high-power grid-connected applications, improperly designed control loops can excite strong inherent AC-DC port coupling, leading to low-frequency oscillation issues, especially under weak grid conditions. To address this problem, this article establishes a multiport admittance matrix for the SST, encompassing its AC dq axes and primary DC port, to characterize its inherent dynamics. Subsequently, a multiport dissipativity analysis is conducted to evaluate the robust stability of the SST. By leveraging the decomposition of passivity conditions into distinct self- and coupling-dissipativity indices, the specific root causes of instability are diagnosed. This framework reveals that a severe coupling-dissipativity failure, induced by the internal dynamics of the synchronization loop, is the dominant instability mechanism rather than a localized self-dissipativity issue. Guided by this diagnosis, a stabilizing controller featuring dynamics-free orthogonal signal reconstruction is designed to reshape the admittance characteristics of the SST. This enhancement specifically targets the identified coupling-dissipativity deficiencies, thereby resolving the root cause of the instability. Finally, the stability analysis and the effectiveness of the enhancement strategy are validated on a down-scaled SST prototype. Experimental results demonstrate that the criterion accurately predicts the coupling-induced oscillations and that the enhanced controller guarantees stable operation under challenging weak-grid conditions.

eess.SY

A Multi-Frequency Input-Admittance Model of Locomotive Rectifier Considering PWM Sideband Harmonic Coupling in Electrical Railways

Electrical railway harmonic instability issues are common in the high-frequency range. The effective frequency of the traditional converter's small-signal averaging model is below 1/2 switching frequency since the pulse width modulation (PWM) sideband harmonic components are ignored. In this article, the dynamic propagations of perturbation frequency and the generated PWM sideband components are constructed first. Then the locomotive rectifier's multi-frequency input-admittance model is derived appropriately. Afterward, an admittance conversion approach is used to convert the multi-frequency model into the single-input-single-output (SISO) model whereas retaining the sideband frequency couplings. The proposed SISO model is more accurate than the traditional small-signal averaging model in the frequency range higher than 1 / 2 switching frequency. It is found that PWM sideband harmonics dominate the locomotive rectifier's input-admittance characteristic higher than 1 / 2 switching frequency. Finally, based on the proposed model, the influence of different switching frequencies, control bandwidths, and traction network impedance on system harmonic stability is revealed by the hardware-in-the-loop (HIL) results.

eess.SY

An Improved Deep Reinforcement Learning Control Strategy for Traction Dual Rectifiers in EMUs

Due to the use of PI-based d q current decoupling in the pulse rectifier of CRH5 high-speed trains, the PI parameters directly affect the traction system's control performance. Linearized control may have issues with reference trajectory changes or model mismatches, leading to a decrease in system performance, while nonlinear control may have problems with jitter and poor steady-state accuracy. This paper proposes a new control strategy that replaces all PI in the d q current decoupling control with a single intelligent agent. This method based on Deep Reinforcement Learning (DRL) can avoid various drawbacks of linearization and nonlinear control and ensure the stability of intermediate DC voltage. However, when EMUs are in different working conditions and switching, the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm used in traction dual rectifiers does not have a good control effect. Focusing on the issue, Reward Shaping (RS) is added to re-design a nonlinear reward function, which can be combined with Prioritized Experience Replay (PER) to increase the convergence speed of the episode reward. The simulation results show that the improved control strategy can be effectively applied to EMUs working in multiple conditions. Finally, the stability analysis is carried out using Lyapunov's second method and the verification results of the hardware-in-the-loop (HIL) simulation platform show that the DRL control has a good effect.

eess.SY

General Incomplete Multimodal Learning via Dynamic Quality Perception

Multimodal learning robust to missing modalities is essential for real-world applications. Existing methods mainly focus on inter-modality missing, where entire modalities are absent, while overlooking intra-modality degradation, where modalities are present but severely corrupted. In practice, these two types of missing often coexist, making existing approaches ineffective. To address this limitation, we propose General Incomplete Multimodal Learning (GIML), a unified framework that simultaneously handles both inter-modality missing and intra-modality degradation through dynamic quality perception. Specifically, GIML models heterogeneous missing patterns as continuous modality information degradation, enabling degradation-aware adaptive fusion. To achieve reliable quality perception, we introduce a Noise-aware Quality Estimator that learns the mapping from corrupted features to noise intensity through controlled noise injection. Furthermore, we propose a Noise-Semantic Decoupled module that separates semantic information from noise interference. This improves robustness and generalization to unseen corruption patterns. Extensive experiments across datasets with diverse modality types demonstrate the effectiveness and generality of GIML. Code is available at: https://github.com/Yu-Five/GIML.

cs.CV

A Physics-Informed Neural Network for Small-Signal Stability in Multi-Inverter Power Systems

The whole-system impedance model has proven a powerful tool for assessing the small-signal stability of multi-inverter power systems; however, its application is limited to a small range around a steady-state operating point due to the inherent assumptions of time invariance and linearisation. In this paper, a dedicated physics-informed neural network (PINN) for small-signal stability analysis in high-dimensional multi-inverter power systems is developed. The PINN is trained with step-response data produced from limited sets of system electromagnetic transient (EMT) simulations, and the trained model can predict the poles and residues of the whole-system impedance/admittance model, i.e., the transfer functions, across the full operating space. Such a PINN offers unique insights into system stability that surpass what conventional analytical methods or EMT simulations can achieve. By characterising how the impedance model evolves with power flow variations, it predicts the dynamic behaviour of the time-varying system and reveals oscillation risks that may emerge while identifying their root causes. It also provides direct visualisation of the possible range of oscillatory modes under a given power flow condition, enabling an optimal generation distribution while maintaining safe operation of the system. The proposed PINN is fully validated on a 2-IBR system and a 4-IBR system, with its application details presented.

eess.SY

Stability and Droop Characteristics Analysis of Observer-Synchronized Grid-Forming Control

This paper analyzes the stability and droop characteristics of Observer-Synchronized grid-forming control. First, a second-order nonlinear autonomous model is derived under the quasi-steady-state assumption. Based on the derived model, the equilibrium points and nonlinear stability properties are investigated using the qualitative theory of differential equations. Explicit parameter conditions are obtained to guarantee almost global asymptotic stability of the desired equilibrium. Furthermore, an analytical expression of the nonlinear droop characteristic is derived to reveal the relationship between active power and frequency. The theoretical analysis is validated through electromagnetic transient simulations and experiments.

eess.SY

HoloPathTracer: Fast and Accurate Wave Path Tracing for Holography

Holography offers unique advantages for delivering perceptual realism while preserving compact form factors in VR/AR. Its perceptual quality, however, hinges on encoding rich wavefronts of photorealistic scenes into interference patterns and then incoherently multiplexing the resulting wave fields for perception. Existing CGH paradigms decouple radiance estimation from wave propagation by pre-rendering radiance on discretized scene sectors. This separation between radiometric and wave-optical computation inherently limits the range of focus cues and visual effects that can be faithfully reproduced, including depth- and view-continuity, and physically based material behaviors such as glossy or mirror-like reflection and refraction. We present a physically accurate yet computationally efficient wave optics rendering framework leveraging path tracing to encode full 3D visual cues into phase holograms. Specifically, we employ a Monte Carlo method to solve both the rendering equation and the Rayleigh--Sommerfeld integral simultaneously. Our algorithm is fully compatible with modern graphics techniques and can generate multiple time-multiplexed random holograms with minimal additional time cost via Path Reuse. By employing a fast approximation with an ambient radiance cache, we realize an order of magnitude convergence speed improvement. The resulting coherent wave fields that inherently encode comprehensive visual effects are converted into phase-only holograms under complex-amplitude supervision. Through extensive simulations and experimental validations on a spatial light modulator-based display prototype, we demonstrate faithful holographic reconstructions of natural 3D cues and complex materials, including realistic defocus blur, view-dependent effects, as well as appearance highlights and reflections.

cs.GR

A Network Transformation Mapping Approach to Synchronization of Multi-Agent Systems With Disconnected Switching Topologies

This paper focuses on the multi-agent synchronization problem with an open-loop unstable leader and followers under the switching topologies. For this issue, the typical approach is intermittent communication (including a spanning tree intermittently) or fast switching strategy. We here consider a more general scenario where each communication link between two agents is randomly disconnected and reconnected, and the durations of both the disconnected intervals and the connected intervals follow negative exponential distributions. To handle this issue, we propose a network transformation mapping method that divides the communication network into reachable and unreachable parts at any time. A node can access the leader's information and synchronize only when it lies in the reachable part; otherwise, it cannot. For each node, the synchronization speed is designed such that its convergence amplitude in the reachable part exceeds its divergence amplitude in the unreachable part. Hence, all nodes achieve asymptotic synchronization over the entire time domain. We further develop adaptive strategies for coupling gains to reduce the computational complexity introduced by the network transformation mapping. The proposed method is also applicable to jointly connected switching topologies and distributed observers under topologies that lack a spanning tree at any time. Finally, two numerical simulations -- synchronous control of multi-one-link manipulator systems and multi-motor systems based on Fieldbus -- are provided to demonstrate the effectiveness of our approach.

math.DS

Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence

Audio-video (AV) generation has recently made strong progress in perceptual quality and multimodal coherence, yet generating content with plausible motion-sound relations remains challenging. Existing methods often produce object motions that are visually unstable and sounds that are only loosely aligned with salient motion or contact events, largely because they lack an explicit motion-aware structure shared by video and audio generation. We present Tora3, a trajectory-guided AV generation framework that improves physical coherence by using object trajectories as a shared kinematic prior. Rather than treating trajectories as a video-only control signal, Tora3 uses them to jointly guide visual motion and acoustic events. Specifically, we design a trajectory-aligned motion representation for video, a kinematic-audio alignment module driven by trajectory-derived second-order kinematic states, and a hybrid flow matching scheme that preserves trajectory fidelity in trajectory-conditioned regions while maintaining local coherence elsewhere. We further curate PAV, a large-scale AV dataset emphasizing motion-relevant patterns with automatically extracted motion annotations. Extensive experiments show that Tora3 improves motion realism, motion-sound synchronization, and overall AV generation quality over strong open-source baselines. Project page: https://ali-videoai.github.io/tora3_page.

cs.CV

Identity-GRPO: Optimizing Multi-Human Identity-preserving Video Generation via Reinforcement Learning

While advanced methods like VACE and Phantom have advanced video generation for specific subjects in diverse scenarios, they struggle with multi-human identity preservation in dynamic interactions, where consistent identities across multiple characters are critical. To address this, we propose Identity-GRPO, a human feedback-driven optimization pipeline for refining multi-human identity-preserving video generation. First, we construct a video reward model trained on a large-scale preference dataset containing human-annotated and synthetic distortion data, with pairwise annotations focused on maintaining human consistency throughout the video. We then employ a GRPO variant tailored for multi-human consistency, which greatly enhances both VACE and Phantom. Through extensive ablation studies, we evaluate the impact of annotation quality and design choices on policy optimization. Experiments show that Identity-GRPO achieves up to 18.9% improvement in human consistency metrics over baseline methods, offering actionable insights for aligning reinforcement learning with personalized video generation.

cs.CV

LaTo: Landmark-tokenized Diffusion Transformer for Fine-grained Human Face Editing

Recent multimodal models for instruction-based face editing enable semantic manipulation but still struggle with precise attribute control and identity preservation. Structural facial representations such as landmarks are effective for intermediate supervision, yet most existing methods treat them as rigid geometric constraints, which can degrade identity when conditional landmarks deviate significantly from the source (e.g., large expression or pose changes, inaccurate landmark estimates). To address these limitations, we propose LaTo, a landmark-tokenized diffusion transformer for fine-grained, identity-preserving face editing. Our key innovations include: (1) a landmark tokenizer that directly quantizes raw landmark coordinates into discrete facial tokens, obviating the need for dense pixel-wise correspondence; (2) a location-mapped positional encoding and a landmark-aware classifier-free guidance that jointly facilitate flexible yet decoupled interactions among instruction, geometry, and appearance, enabling strong identity preservation; and (3) a landmark predictor that leverages vision-language models to infer target landmarks from instructions and source images, whose structured chain-of-thought improves estimation accuracy and interactive control. To mitigate data scarcity, we curate HFL-150K, to our knowledge the largest benchmark for this task, containing over 150K real face pairs with fine-grained instructions. Extensive experiments show that LaTo outperforms state-of-the-art methods by 7.8% in identity preservation and 4.6% in semantic consistency. Code and dataset will be made publicly available upon acceptance.

cs.CV

Tora2: Motion and Appearance Customized Diffusion Transformer for Multi-Entity Video Generation

Recent advances in diffusion transformer models for motion-guided video generation, such as Tora, have shown significant progress. In this paper, we present Tora2, an enhanced version of Tora, which introduces several design improvements to expand its capabilities in both appearance and motion customization. Specifically, we introduce a decoupled personalization extractor that generates comprehensive personalization embeddings for multiple open-set entities, better preserving fine-grained visual details compared to previous methods. Building on this, we design a gated self-attention mechanism to integrate trajectory, textual description, and visual information for each entity. This innovation significantly reduces misalignment in multimodal conditioning during training. Moreover, we introduce a contrastive loss that jointly optimizes trajectory dynamics and entity consistency through explicit mapping between motion and personalization embeddings. Tora2 is, to our best knowledge, the first method to achieve simultaneous multi-entity customization of appearance and motion for video generation. Experimental results demonstrate that Tora2 achieves competitive performance with state-of-the-art customization methods while providing advanced motion control capabilities, which marks a critical advancement in multi-condition video generation. Project page: https://ali-videoai.github.io/Tora2_page/.

cs.CV

Crystal-symmetry-paired spin-valley locking in a layered room-temperature antiferromagnet

Recent theoretical efforts predicted a type of unconventional antiferromagnet characterized by the crystal symmetry C (rotation or mirror), which connects antiferromagnetic sublattices in real space and simultaneously couples spin and momentum in reciprocal space. This results in a unique C-paired spin-valley locking (SVL) and corresponding novel properties such as piezomagnetism and noncollinear spin current even without spin-orbit coupling. However, the unconventional antiferromagnets reported thus far are not layered materials, limiting their potential in spintronic applications. Additionally, they do not meet the necessary symmetry requirements for nonrelativistic spin current. Here, we report the realization of C-paired SVL in a layered room-temperature antiferromagnetic compound, Rb1-{\delta}V2Te2O. Spin resolved photoemission measurements directly demonstrate the opposite spin splitting between C-paired valleys. Quasi-particle interference patterns reveal the suppression of inter-valley scattering due to the spin selection rules, as a direct consequence of C-paired SVL. All these experiments are well consistent with the results obtained from first-principles calculations. Our observations represent the first realization of layered antiferromagnets with C-paired SVL, enabling both the advantages of layered materials and possible control through crystal symmetry manipulation. These results hold significant promise and broad implications for advancements in magnetism, electronics, and information technology.

cond-mat.str-el