SearcharxivSearch

arXiv subjects

Meng Ye

Publications and source records attributed to Meng Ye.

At least 19 recordsLinked to original sources

Spin-Chirality-Driven Bulk Photovoltaic Effect in van der Waals Magnet CrSBr

The bulk photovoltaic effect (BPVE) can be greatly enriched in magnetic materials. Here, we establish vector spin chirality as a tunable knob for generating an unconventional time-reversal-even magnetic BPVE, comprising the chiral shift current (CSC) and chiral injection current (CIC). Using bilayer antiferromagnetic (AFM) CrSBr as a prototype, we theoretically demonstrate the emergence of CSC and CIC. Compared with conventional photovoltaic currents arising from noncentrosymmetric crystal structures or collinear magnetic orderings, CSC and CIC not only possess comparable magnitudes but also exhibit exceptional tunability. Specifically, they can be switched on and off by magnetic-field-induced spin canting, reversed in direction upon canting-direction reversal, and continuously modulated in intensity via canting-angle variation. Furthermore, we reveal an unusual optical transition channel governing both currents in CrSBr. Our work establishes an unconventional magnetic BPVE with remarkable controllability, paving the way for applications in optoelectronics and magnetic sensing in noncollinear magnets.

cond-mat.mtrl-sci

Bi-PT: Bidirectional Cross-Attention Point Transformers for Four-Chamber Heart Reconstruction from Sparse Cardiac MRI Data

We propose Bi-PT, a pipeline for reconstructing 3D four-chamber human heart meshes from clinical sparsely sampled cardiac magnetic resonance imaging (CMR) data. This work addresses the error-prone generation of 3D cardiac shape from a sparse point cloud (SPC) extracted from 2D long-axis and short-axis views used in routine clinical CMR protocols. Bi-PT enables accurate inference of the four-chamber heart mesh from the SPC by learning robust point features via bidirectional point cross-attention between an atlas and the SPC, together with per-point semantic labels that improve correspondence estimation. We formulate the deformation field as a Neural Ordinary Differential Equation (NODE) parameterized by a per-point affine transformation and translation to deform the atlas toward the target heart shape. By learning such a NODE, we can guarantee the deformation field to be a locally affine diffeomorphic deformation. We also integrate a semantic label loss into the Chamfer distance to encourage label-consistent correspondences and add a smoothness regularization to stabilize and improve the learning of the deformation field. Extensive experiments demonstrate that Bi-PT achieves accurate and robust performance compared to baselines.

cs.CV

Cardiac MRI Through-Plane Super-Resolution Guided by Reference and Memory

Clinical cardiac MRI is commonly acquired with high in-plane resolution but coarse through-plane resolution to reduce scan time and accommodate breath-hold and cardiac-motion constraints, which limits 3D analysis and diagnostic accuracy. We propose STRMSR, a reference- and memory-guided through-plane super-resolution (SR) framework that reconstructs high-resolution (HR) cardiac volumes by leveraging HR reference views acquired from the same subject and intermediate SR results as the memory. Our method uses coarse-to-fine contextual matching to establish robust correspondence between low-resolution target and reference/memory images under spatial misalignment. A learnable patch-wise dynamic feature aggregation module predicts content-adaptive mixture weights for each local patch, effectively fusing dynamic information while suppressing unreliable feature transfers. The intermediate SR results stored in the memory bank ensure slice-to-slice consistency for the super-resolved 3D volume. Experiments on the WHS cardiac MRI dataset under two reference protocols, orthogonal-plane views and long-axis chamber views, demonstrate consistent improvements over baselines at 4x and 8x upsampling factors.Code is available at https://github.com/030108ming/STRMSR

cs.CV

QueryPlot: Generating Geological Evidence Layers using Natural Language Queries for Mineral Exploration

Mineral prospectivity mapping requires synthesizing heterogeneous geological knowledge, including textual deposit models and geospatial datasets, to identify regions likely to host specific mineral deposit types. This process is traditionally manual and knowledge-intensive. We present QueryPlot, a semantic retrieval and mapping framework that integrates large-scale geological text corpora with geologic map data using modern Natural Language Processing techniques. We curate descriptive deposit models for over 120 deposit types and transform the State Geologic Map Compilation (SGMC) polygons into structured textual representations. Given a user-defined natural language query, the system encodes both queries and region descriptions using a pretrained embedding model and computes semantic similarity scores to rank and spatially visualize regions as continuous evidence layers. QueryPlot supports compositional querying over deposit characteristics, enabling aggregation of multiple similarity-derived layers for multi-criteria prospectivity analysis. In a case study on tungsten skarn deposits, we demonstrate that embedding-based retrieval achieves high recall of known occurrences and produces prospective regions that closely align with expert-defined permissive tracts. Furthermore, similarity scores can be incorporated as additional features in supervised learning pipelines, yielding measurable improvements in classification performance. QueryPlot is implemented as a web-based system supporting interactive querying, visualization, and export of GIS-compatible prospectivity layers.To support future research, we have made the source code and datasets used in this study publicly available.

cs.CL

K-Prism: A Knowledge-Guided and Prompt Integrated Universal Medical Image Segmentation Model

Medical image segmentation is fundamental to clinical decision-making, yet existing models remain fragmented. They are usually trained on single knowledge sources and specific to individual tasks, modalities, or organs. This fragmentation contrasts sharply with clinical practice, where experts seamlessly integrate diverse knowledge: anatomical priors from training, exemplar-based reasoning from reference cases, and iterative refinement through real-time interaction. We present $\textbf{K-Prism}$, a unified segmentation framework that mirrors this clinical flexibility by systematically integrating three knowledge paradigms: (i) $\textit{semantic priors}$ learned from annotated datasets, (ii) $\textit{in-context knowledge}$ from few-shot reference examples, and (iii) $\textit{interactive feedback}$ from user inputs like clicks or scribbles. Our key insight is that these heterogeneous knowledge sources can be encoded into a dual-prompt representation: 1-D sparse prompts defining $\textit{what}$ to segment and 2-D dense prompts indicating $\textit{where}$ to attend, which are then dynamically routed through a Mixture-of-Experts (MoE) decoder. This design enables flexible switching between paradigms and joint training across diverse tasks without architectural modifications. Comprehensive experiments on 18 public datasets spanning diverse modalities (CT, MRI, X-ray, pathology, ultrasound, etc.) demonstrate that K-Prism achieves state-of-the-art performance across semantic, in-context, and interactive segmentation settings.

cs.CV

The Influence of Text Variation on User Engagement in Cross-Platform Content Sharing

In today's cross-platform social media landscape, understanding factors that drive engagement for multimodal content, especially text paired with visuals, remains complex. This study investigates how rewriting Reddit post titles adapted from YouTube video titles affects user engagement. First, we build and analyze a large dataset of Reddit posts sharing YouTube videos, revealing that 21% of post titles are minimally modified. Statistical analysis demonstrates that title rewrites measurably improve engagement. Second, we design a controlled, multi-phase experiment to rigorously isolate the effects of textual variations by neutralizing confounding factors like video popularity, timing, and community norms. Comprehensive statistical tests reveal that effective title rewrites tend to feature emotional resonance, lexical richness, and alignment with community-specific norms. Lastly, pairwise ranking prediction experiments using a fine-tuned BERT classifier achieves 74% accuracy, significantly outperforming near-random baselines, including GPT-4o. These results validate that our controlled dataset effectively minimizes confounding effects, allowing advanced models to both learn and demonstrate the impact of textual features on engagement. By bridging quantitative rigor with qualitative insights, this study uncovers engagement dynamics and offers a robust framework for future cross-platform, multimodal content strategies.

cs.SI

Are there type-III multiferroics?

Multiferroics are known to be classified into two types. However, type-I lacks sufficient magnetoelectric coupling and type-II lacks sufficient electric polarization, making both practically difficult. In this work, we explore the possibility of type-III multiferroics, where the origins of ferroelectricity and magnetism are highly intertwined but not causally related, with a combination of strong magnetoelectric coupling and large polarization. Our first-principles calculations predict that monolayer TiCdO$_{4}$ is such a type-III ferroelectric-ferromagnetic multiferroics with both electronic and magnetic orders originating from competing electron populations on oxygen atoms. It shows an electric polarization of 50 $\mu$C/m$^{2}$ while the maximum linear and quadratic magnetoelectric response are as high as 35000 ps/m and 1.59 $\times$ 10$^{-14}$ s/A, respectively. Our study opens up new perspectives for the discovery and design of much-anticipated multiferroics that can be used for cross-modulation.

cond-mat.mtrl-sci

Spin-chirality-driven second-harmonic generation in two-dimensional magnet CrSBr

The interplay between magnetism and light can create abundant optical phenomena. Here, we demonstrate the emergence of an unconventional magnetization-induced second-harmonic generation (MSHG) stemming from vector spin chirality, denoted as chiral second-harmonic generation (SHG). Taking the antiferromagnetic (AFM) CrSBr bilayer as a prototype, we theoretically show that, via spin canting, the chiral SHG can be continuously tuned from zero to a value one order of magnitude larger than its intrinsic MSHG. Chiral SHG is found to be proportional to spin chirality and spin-canting-induced electric polarization, while intrinsic MSHG is proportional to the N\'eel vector, demonstrating their different physical mechanisms. Additionally, we reveal a unique interference effect between these two types of MSHG under the reversal of spin-canting direction, generating a giant modulation of SHG signals. Our work not only uncovers a unique SHG with exceptional tunability but also promotes the applications of AFM optical devices and magnetoelectric detection techniques.

cond-mat.mtrl-sci

Unconventional bias-dependent tunneling magnetoresistance in van der Waals ferromagnetic/semiconductor heterojunctions

Two-dimensional van der Waals (vdW) ferromagnetic/semiconductor heterojunctions represent an ideal platform for studying and exploiting tunneling magnetoresistance (TMR) effects due to the versatile band structure of semiconductors and their high-quality interfaces. In the all-vdW magnetic tunnel junction (MTJ) devices, both the magnitude and sign of the TMR can be tuned by an applied voltage. Typically, as the bias voltage increases, first the amplitude of the TMR decreases, then the sign of the TMR reverses and/or oscillates. Here, we report on an unconventional bias-dependent TMR in the all-vdW Fe3GaTe2/GaSe/Fe3GaTe2 MTJs, where the TMR first increases, then decreases, and finally undergoes a sign reversal as the bias voltage increases. This dependence cannot be explained by traditional models of MTJs. We propose an in-plane electron momentum (k//) resolved tunneling model that considers both the coherent degree of k// and the decay of the electron wave function through the semiconductor spacer layer. This can explain well the conventional and unconventional bias-dependent TMR. Our results thus provide a deeper understanding of the bias-dependent spin-transport in semiconductor-based MTJs and offer new insights into semiconductor spintronics.

cond-mat.mtrl-sci

Rate-My-LoRA: Efficient and Adaptive Federated Model Tuning for Cardiac MRI Segmentation

Cardiovascular disease (CVD) and cardiac dyssynchrony are major public health problems in the United States. Precise cardiac image segmentation is crucial for extracting quantitative measures that help categorize cardiac dyssynchrony. However, achieving high accuracy often depends on centralizing large datasets from different hospitals, which can be challenging due to privacy concerns. To solve this problem, Federated Learning (FL) is proposed to enable decentralized model training on such data without exchanging sensitive information. However, bandwidth limitations and data heterogeneity remain as significant challenges in conventional FL algorithms. In this paper, we propose a novel efficient and adaptive federate learning method for cardiac segmentation that improves model performance while reducing the bandwidth requirement. Our method leverages the low-rank adaptation (LoRA) to regularize model weight update and reduce communication overhead. We also propose a \mymethod{} aggregation technique to address data heterogeneity among clients. This technique adaptively penalizes the aggregated weights from different clients by comparing the validation accuracy in each client, allowing better generalization performance and fast local adaptation. In-client and cross-client evaluations on public cardiac MR datasets demonstrate the superiority of our method over other LoRA-based federate learning approaches.

cs.CV

VerSe: Integrating Multiple Queries as Prompts for Versatile Cardiac MRI Segmentation

Despite the advances in learning-based image segmentation approach, the accurate segmentation of cardiac structures from magnetic resonance imaging (MRI) remains a critical challenge. While existing automatic segmentation methods have shown promise, they still require extensive manual corrections of the segmentation results by human experts, particularly in complex regions such as the basal and apical parts of the heart. Recent efforts have been made on developing interactive image segmentation methods that enable human-in-the-loop learning. However, they are semi-automatic and inefficient, due to their reliance on click-based prompts, especially for 3D cardiac MRI volumes. To address these limitations, we propose VerSe, a Versatile Segmentation framework to unify automatic and interactive segmentation through mutiple queries. Our key innovation lies in the joint learning of object and click queries as prompts for a shared segmentation backbone. VerSe supports both fully automatic segmentation, through object queries, and interactive mask refinement, by providing click queries when needed. With the proposed integrated prompting scheme, VerSe demonstrates significant improvement in performance and efficiency over existing methods, on both cardiac MRI and out-of-distribution medical imaging datasets. The code is available at https://github.com/bangwayne/Verse.

cs.CV

Learning Volumetric Neural Deformable Models to Recover 3D Regional Heart Wall Motion from Multi-Planar Tagged MRI

Multi-planar tagged MRI is the gold standard for regional heart wall motion evaluation. However, accurate recovery of the 3D true heart wall motion from a set of 2D apparent motion cues is challenging, due to incomplete sampling of the true motion and difficulty in information fusion from apparent motion cues observed on multiple imaging planes. To solve these challenges, we introduce a novel class of volumetric neural deformable models ($\upsilon$NDMs). Our $\upsilon$NDMs represent heart wall geometry and motion through a set of low-dimensional global deformation parameter functions and a diffeomorphic point flow regularized local deformation field. To learn such global and local deformation for 2D apparent motion mapping to 3D true motion, we design a hybrid point transformer, which incorporates both point cross-attention and self-attention mechanisms. While use of point cross-attention can learn to fuse 2D apparent motion cues into material point true motion hints, point self-attention hierarchically organised as an encoder-decoder structure can further learn to refine these hints and map them into 3D true motion. We have performed experiments on a large cohort of synthetic 3D regional heart wall motion dataset. The results demonstrated the high accuracy of our method for the recovery of dense 3D true motion from sparse 2D apparent motion cues. Project page is at https://github.com/DeepTag/VolumetricNeuralDeformableModels.

eess.IV

Continuous Spatio-Temporal Memory Networks for 4D Cardiac Cine MRI Segmentation

Current cardiac cine magnetic resonance image (cMR) studies focus on the end diastole (ED) and end systole (ES) phases, while ignoring the abundant temporal information in the whole image sequence. This is because whole sequence segmentation is currently a tedious process and inaccurate. Conventional whole sequence segmentation approaches first estimate the motion field between frames, which is then used to propagate the mask along the temporal axis. However, the mask propagation results could be prone to error, especially for the basal and apex slices, where through-plane motion leads to significant morphology and structural change during the cardiac cycle. Inspired by recent advances in video object segmentation (VOS), based on spatio-temporal memory (STM) networks, we propose a continuous STM (CSTM) network for semi-supervised whole heart and whole sequence cMR segmentation. Our CSTM network takes full advantage of the spatial, scale, temporal and through-plane continuity prior of the underlying heart anatomy structures, to achieve accurate and fast 4D segmentation. Results of extensive experiments across multiple cMR datasets show that our method can improve the 4D cMR segmentation performance, especially for the hard-to-segment regions.

cs.CV

Deep Band Crossings Enhanced Nonlinear Optical Effects

Nonlinear optical (NLO) effects in materials with band crossings have attracted significant research interests due to the divergent band geometric quantities around these crossings. Most current research has focused on band crossings between the valence and conduction bands. However, such crossings are absent in insulators, which are more relevant for NLO applications. In this work, we demonstrate that NLO effects can be significantly enhanced by band crossings within the valence or conduction bands, which we designate as "deep band crossings" (DBCs). As an example, in two dimensions, we show that shift conductivity can be substantially enhanced or even divergent due to a mirror-protected "deep Dirac nodal point". In three dimensions, we propose GeTe as an ideal material where shift conductivity is enhanced by "deep Dirac nodal lines". The ubiquity of this enhancement is further confirmed by high-throughput calculations. Other types of DBCs and NLO effects are also discussed. By manipulating band crossings between arbitrary bands, our work offers a simple, practical, and universal way to greatly enhance NLO effects.

cond-mat.mes-hall

Atom Cavity Encoding for NP-Complete Problems

We consider an atom-cavity system having long-range atomic interactions mediated by cavity modes. It has been shown that quantum simulations of spin models with this system can naturally be used to solve number partition problems. Here, we present encoding schemes for numerous NP-complete problems, encompassing the majority of Karp's 21 NP-complete problems. We find a number of such computation problems can be encoded by the atom-cavity system at a linear cost of atom number. There are still certain problems that cannot be encoded by the atom-cavity as efficiently, such as quadratic unconstrained binary optimization (QUBO), and the Hamiltonian cycle. For these problems, we provide encoding schemes with a quadratic or quartic cost in the atom number. We expect this work to provide important guidance to search for the practical quantum advantage of the atom-cavity system in solving NP-complete problems. Moreover, the encoding schemes we develop here may also be adopted in other optical systems for solving NP-complete problems, where a similar form of Mattis-type spin glass Hamiltonian as in the atom-cavity system can be implemented.

quant-ph

Empowering Interdisciplinary Insights with Dynamic Graph Embedding Trajectories

We developed DyGETViz, a novel framework for effectively visualizing dynamic graphs (DGs) that are ubiquitous across diverse real-world systems. This framework leverages recent advancements in discrete-time dynamic graph (DTDG) models to adeptly handle the temporal dynamics inherent in dynamic graphs. DyGETViz effectively captures both micro- and macro-level structural shifts within these graphs, offering a robust method for representing complex and massive dynamic graphs. The application of DyGETViz extends to a diverse array of domains, including ethology, epidemiology, finance, genetics, linguistics, communication studies, social studies, and international relations. Through its implementation, DyGETViz has revealed or confirmed various critical insights. These include the diversity of content sharing patterns and the degree of specialization within online communities, the chronological evolution of lexicons across decades, and the distinct trajectories exhibited by aging-related and non-related genes. Importantly, DyGETViz enhances the accessibility of scientific findings to non-domain experts by simplifying the complexities of dynamic graphs. Our framework is released as an open-source Python package for use across diverse disciplines. Our work not only addresses the ongoing challenges in visualizing and analyzing DTDG models but also establishes a foundational framework for future investigations into dynamic graph representation and analysis across various disciplines.

cs.LG

Giant and controllable nonlinear magneto-optical effects in two-dimensional magnets

The interplay of polarization and magnetism in materials with light can create rich nonlinear magneto-optical (NLMO) effects, and the recent discovery of two-dimensional (2D) van der Waals magnets provides remarkable control over NLMO effects due to their superb tunability. Here, based on first-principles calculations, we reported giant NLMO effects in CrI3-based 2D magnets, including a dramatic change of second-harmonics generation (SHG) polarization direction (90 degrees) and intensity (on/off switch) under magnetization reversal, and a 100% SHG circular dichroism effect. We further revealed that these effects could not only be used to design ultra-thin multifunctional optical devices, but also to detect subtle magnetic orderings. Remarkably, we analytically derived conditions to achieve giant NLMO effects and propose general strategies to realize them in 2D magnets. Our work not only uncovers a series of intriguing NLMO phenomena, but also paves the way for both fundamental research and device applications of ultra-thin NLMO materials.

cond-mat.mtrl-sci

A Video is Worth 10,000 Words: Training and Benchmarking with Diverse Captions for Better Long Video Retrieval

Existing long video retrieval systems are trained and tested in the paragraph-to-video retrieval regime, where every long video is described by a single long paragraph. This neglects the richness and variety of possible valid descriptions of a video, which could range anywhere from moment-by-moment detail to a single phrase summary. To provide a more thorough evaluation of the capabilities of long video retrieval systems, we propose a pipeline that leverages state-of-the-art large language models to carefully generate a diverse set of synthetic captions for long videos. We validate this pipeline's fidelity via rigorous human inspection. We use synthetic captions from this pipeline to perform a benchmark of a representative set of video language models using long video datasets, and show that the models struggle on shorter captions. We show that finetuning on this data can both mitigate these issues (+2.8% R@1 over SOTA on ActivityNet with diverse captions), and even improve performance on standard paragraph-to-video retrieval (+1.0% R@1 on ActivityNet). We also use synthetic data from our pipeline as query expansion in the zero-shot setting (+3.4% R@1 on ActivityNet). We derive insights by analyzing failure cases for retrieval with short captions. For data access and other details, please refer to our project website at https://mgwillia.github.io/10k-words.

cs.CV