arXiv · 2507.02003
Unsupervised Cardiac Video Translation Via Motion Feature Guided Diffusion Model
Abstract
This paper presents a novel motion feature guided diffusion model for unpaired video-to-video translation (MFD-V2V), designed to synthesize dynamic, high-contrast cine cardiac magnetic resonance (CMR) from lower-contrast, artifact-prone displacement encoding with stimulated echoes (DENSE) CMR sequences. To achieve this, we first introduce a Latent Temporal Multi-Attention (LTMA) registration network that effectively learns more accurate and consistent cardiac motions from cine CMR image videos. A multi-level motion feature guided diffusion model, equipped with a specialized Spatio-Temporal Motion Encoder (STME) to extract fine-grained motion conditioning, is then developed to improve synthesis quality and fidelity. We evaluate our method, MFD-V2V, on a comprehensive cardiac dataset, demonstrating superior performance over the state-of-the-art in both quantitative metrics and qualitative assessments. Furthermore, we show the benefits of our synthesized cine CMRs improving downstream clinical and analytical tasks, underscoring the broader impact of our approach. Our code is publicly available at https://github.com/SwaksharDeb/MFD-V2V.
Explore related subjects
Keep this discovery
Swakshar Deb, Nian Wu, Frederick H. Epstein, Miaomiao Zhang. 2025-07-01. Unsupervised Cardiac Video Translation Via Motion Feature Guided Diffusion Model. https://arxiv.org/abs/2507.02003
Cite the original work for its findings. Save a collection to share your selection of sources.