SearcharxivSearch

arXiv subjects

Mengying Wang

Publications and source records attributed to Mengying Wang.

14 recordsLinked to original sources

Graph Query Generation with Constraint-guided Large Language Agents

Knowledge Graph Question Answering (KGQA) has advanced through structured query generation, yet most efforts target RDF/SPARQL, leaving Cypher and property graphs underexplored, despite increasing demand for unified KGQA in industry settings. We propose UniQGen, a novel constraint-based framework that employs LLM agents to dynamically extract and refine representative graph query clauses into executable, intent-aligned graph queries across query languages. The foundation of our method is a variant of Chase & Backchase, a family of algorithms for query optimization and reformulation. We extend Chase & Backchase with a dynamic reasoning process over query constraints that also interact with LLMs for query quality estimation. With a Cypher-supported Freebase graph deployed on Amazon Neptune, we extensively evaluate our approach on popular KGQA benchmarks (GraphQ, GrailQA, and WebQSP). We demonstrate that UniQGen outperforms state-of-the-art graph query generation techniques in both accuracy and efficiency, with F1 gains of 31.6% on GraphQ and 4.9% on GrailQA. Unlike prior methods, our framework does not require fine-tuning for schema matching, making it more extensible to schema-less graphs and semantics in query workloads, and is more suitable for enterprise-grade KGQA. We release Cypher outputs and a Neptune-ready Freebase snapshot to support reproducible, cross-language KGQA research.

cs.DB

Assessing LLMs for Serendipity Discovery in Knowledge Graphs: A Case for Drug Repurposing

Large Language Models (LLMs) have greatly advanced knowledge graph question answering (KGQA), yet existing systems are typically optimized for returning highly relevant but predictable answers. A missing yet desired capacity is to exploit LLMs to suggest surprise and novel ("serendipitious") answers. In this paper, we formally define the serendipity-aware KGQA task and propose the SerenQA framework to evaluate LLMs' ability to uncover unexpected insights in scientific KGQA tasks. SerenQA includes a rigorous serendipity metric based on relevance, novelty, and surprise, along with an expert-annotated benchmark derived from the Clinical Knowledge Graph, focused on drug repurposing. Additionally, it features a structured evaluation pipeline encompassing three subtasks: knowledge retrieval, subgraph reasoning, and serendipity exploration. Our experiments reveal that while state-of-the-art LLMs perform well on retrieval, they still struggle to identify genuinely surprising and valuable discoveries, underscoring a significant room for future improvements. Our curated resources and extended version are released at: https://cwru-db-group.github.io/serenQA.

cs.CL

ML-Asset Management: Curation, Discovery, and Utilization

Machine learning (ML) assets, such as models, datasets, and metadata, are central to modern ML workflows. Despite their explosive growth in practice, these assets are often underutilized due to fragmented documentation, siloed storage, inconsistent licensing, and lack of unified discovery mechanisms, making ML-asset management an urgent challenge. This tutorial offers a comprehensive overview of ML-asset management activities across its lifecycle, including curation, discovery, and utilization. We provide a categorization of ML assets, and major management issues, survey state-of-the-art techniques, and identify emerging opportunities at each stage. We further highlight system-level challenges related to scalability, lineage, and unified indexing. Through live demonstrations of systems, this tutorial equips both researchers and practitioners with actionable insights and practical tools for advancing ML-asset management in real-world and domain-specific settings.

cs.DB

Generating Skyline Datasets for Data Science Models

Preparing high-quality datasets required by various data-driven AI and machine learning models has become a cornerstone task in data-driven analysis. Conventional data discovery methods typically integrate datasets towards a single pre-defined quality measure that may lead to bias for downstream tasks. This paper introduces MODis, a framework that discovers datasets by optimizing multiple user-defined, model-performance measures. Given a set of data sources and a model, MODis selects and integrates data sources into a skyline dataset, over which the model is expected to have the desired performance in all the performance measures. We formulate MODis as a multi-goal finite state transducer, and derive three feasible algorithms to generate skyline datasets. Our first algorithm adopts a "reduce-from-universal" strategy, that starts with a universal schema and iteratively prunes unpromising data. Our second algorithm further reduces the cost with a bi-directional strategy that interleaves data augmentation and reduction. We also introduce a diversification algorithm to mitigate the bias in skyline datasets. We experimentally verify the efficiency and effectiveness of our skyline data discovery algorithms, and showcase their applications in optimizing data science pipelines.

cs.DB

Quantifying the Risks of Tool-assisted Rephrasing to Linguistic Diversity

Writing assistants and large language models see widespread use in the creation of text content. While their effectiveness for individual users has been evaluated in the literature, little is known about their proclivity to change language or reduce its richness when adopted by a large user base. In this paper, we take a first step towards quantifying this risk by measuring the semantic and vocabulary change enacted by the use of rephrasing tools on a multi-domain corpus of human-generated text.

cs.CL

Generating Robust Counterfactual Witnesses for Graph Neural Networks

This paper introduces a new class of explanation structures, called robust counterfactual witnesses (RCWs), to provide robust, both counterfactual and factual explanations for graph neural networks. Given a graph neural network M, a robust counterfactual witness refers to the fraction of a graph G that are counterfactual and factual explanation of the results of M over G, but also remains so for any "disturbed" G by flipping up to k of its node pairs. We establish the hardness results, from tractable results to co-NP-hardness, for verifying and generating robust counterfactual witnesses. We study such structures for GNN-based node classification, and present efficient algorithms to verify and generate RCWs. We also provide a parallel algorithm to verify and generate RCWs for large graphs with scalability guarantees. We experimentally verify our explanation generation process for benchmark datasets, and showcase their applications.

cs.LG

Model-based multi-sensor fusion for reconstructing wall-bounded turbulence

Wall-bounded turbulent flows can be challenging to measure within experiments due to the breadth of spatial and temporal scales inherent in such flows. Instrumentation capable of obtaining time-resolved data (e.g., Hot-Wire Anemometers) tends to be restricted to spatially-localized point measurements; likewise, instrumentation capable of achieving spatially-resolved field measurements (e.g., Particle Image Velocimetry) tends to lack the sampling rates needed to attain time-resolution in many such flows. In this study, we propose to fuse measurements from multi-rate and multi-fidelity sensors with predictions from a physics-based model to reconstruct the spatiotemporal evolution of a wall-bounded turbulent flow. A "fast" filter is formulated to assimilate high-rate point measurements with estimates from a linear model derived from the Navier-Stokes equations. Additionally, a "slow" filter is used to update the reconstruction every time a new field measurement becomes available. By marching through the data both forward and backward in time, we are able to reconstruct the turbulent flow with greater spatiotemporal resolution than either sensing modality alone. We demonstrate the approach using direct numerical simulations of a turbulent channel flow from the Johns Hopkins Turbulence Database. A statistical analysis of the model-based multi-sensor fusion approach is also conducted.

physics.flu-dyn

Topological Phase Transition Induced by Image Potential States in MXenes: A Theoretical Investigation

MXenes, a family of two-dimensional transition metal carbides and nitrides, have various tunable physical and chemical properties. Their diverse prospective applications in electronics and energy storage devices have triggered great interests in science and technology. MXenes can be functionalized by different surface terminations. Some O and F functionalized MXenes monolayers have been predicted to be topological insulators (TIs). However, the reported OH functionalized MXenes TIs are very few and their electronic structures need to be investigated in more detail. It has been revealed that the work functions of MXenes are reduced significantly by OH termination and the image potential (IP) states move close to the Fermi level. The wave functions of these IP states are spatially extensive outside the surfaces. By stacking the OH-functionalized MXenes, the energies of the IP states can be modulated by the interlayer distances of multilayers, because the overlap and hybridization of the wave functions between the neighboring layers are significant. Therefore, these stacking layers are interacted and coupled with IP states. Here, based on first-principles calculations, we demonstrate that the stacking of two-dimensional topologically trivial OH-functionalized MXenes, such as V$_2$HfC$_2$(OH)$_2$, possibly gives rise to the topologically nontrivial energy bands. In other words, the topological properties of V$_2$HfC$_2$(OH)$_2$ multilayers can be modulated by its interlayer distance. An energy band inversion involving IP states is proposed. We expect that these results can advance the future application of MXenes or other low work function multilayer materials as controllable TI devices.

cond-mat.mtrl-sci

Reconstructing the time evolution of wall-bounded turbulent flows from non-time resolved PIV measurements

Particle Image Velocimetry (PIV) systems are often limited in their ability to fully resolve the spatiotemporal fluctuations inherent in turbulent flows due to hardware constraints. In this study, we develop models based on Rapid Distortion Theory (RDT) and Taylor's Hypothesis (TH) to reconstruct the time evolution of a turbulent flow field in the intermediate period between consecutive PIV snapshots obtained using a non-time resolved system. The linear governing equations are evolved forwards and backwards in time using the PIV snapshots as initial conditions. The flow field in the intervening period is then reconstructed by taking a weighted sum of the forward and backward estimates. This spatiotemporal weighting function is designed to account for the advective nature of the RDT and TH equations. Reconstruction accuracy is evaluated as a function of spatial resolution and reconstruction time horizon using Direct Numerical Simulation data for turbulent channel flow from the Johns Hopkins Turbulence Database. This method reconstructs single-point turbulence statistics well and resolves velocity spectra at frequencies higher than the temporal Nyquist limit of the acquisition system. Reconstructions obtained using a characteristics-based evolution of the flow field under TH prove to be more accurate compared to reconstructions obtained from numerical integration of the discretized forms of RDT and TH. The effect of measurement noise on reconstruction error is also evaluated.

physics.flu-dyn

Cutting and shuffling with diffusion: Evidence for cut-offs in interval exchange maps

Low-dimensional dynamical systems are fruitful models for mixing in fluid and granular flows. We study a one-dimensional discontinuous dynamical system (termed "cutting and shuffling" of a line segment), and we present a comprehensive computational study of its finite-time mixing properties including the effect of diffusion. To explore a large parameter space, we introduce fit functions for two mixing metrics of choice: the number of cutting interfaces (a standard quantity in dynamical systems theory of interval exchange transformations) and a mixing norm (a more physical measure of mixing). We compute averages of the mixing metrics across different permutations (shuffling protocols), showing that the latter averages are a robust descriptor of mixing for any permutation choice. If the decay of the normalized mixing norm is plotted against the number of map iterations rescaled by the characteristic e-folding time, then universality emerges: mixing norm decay curves across all cutting and shuffling protocols collapse onto a single stretched-exponential profile. Next, we predict this critical number of iterations using the average length of unmixed subsegments of continuous color during cutting and shuffling and a Batchelor-scale-type diffusion argument. This prediction, called a "stopping time" for finite Markov chains, compares well with the e-folding time of the stretched-exponential fit. Finally, we quantify the effect of diffusion on cutting and shuffling through a Péclet number (a dimensionless inverse diffusivity), showing that the system transitions more sharply from an unmixed initial state to a mixed final state as the Péclet number becomes large. Our numerical investigation of cutting and shuffling of a line segment in the presence of diffusion thus present evidence for the latter phenomenon, known as a "cut-off" for finite Markov chains, in interval exchange maps.

math.DS

Detecting exotic wakes with hydrodynamic sensors

Wake sensing for bioinspired robotic swimmers has been the focus of much investigation owing to its relevance to locomotion control, especially in the context of schooling and target following. Many successful wake sensing strategies have been devised based on models of von Karman-type wakes; however, such wake sensing technologies are invalid in the context of exotic wake types that commonly arise in swimming locomotion. Indeed, exotic wakes can exhibit markedly different dynamics, and so must be modeled and sensed accordingly. Here, we propose a general wake detection protocol for distinguishing between wake types from measured hydrodynamic signals alone. An ideal-flow model is formulated and used to demonstrate the general wake detection framework in a proof-of-concept study. We show that wakes with different underlying dynamics impart distinct signatures on a fish-like body, which can be observed in time-series measurements at a single location on the body surface. These hydrodynamic wake signatures are used to construct a wake classification library that is then used to classify unknown wakes from hydrodynamic signal measurements. The wake detection protocol is found to have an accuracy rate of over 95% in the majority of performance studies conducted here. Thus, exotic wake detection is shown to be viable, which suggests that such technologies have the potential to become key enablers of multi-model sensing and locomotion control strategies in the future.

physics.flu-dyn

A Timed Calculus for Mobile Ad Hoc Networks

We develop a timed calculus for Mobile Ad Hoc Networks embodying the peculiarities of local broadcast, node mobility and communication interference. We present a Reduction Semantics and a Labelled Transition Semantics and prove the equivalence between them. We then apply our calculus to model and study some MAC-layer protocols with special emphasis on node mobility and communication interference. A main purpose of the semantics is to describe the various forms of interference while nodes change their locations in the network. Such interference only occurs when a node is simultaneously reached by more than one ongoing transmission over the same channel.

cs.LO

LAGE: A Java Framework to reconstruct Gene Regulatory Networks from Large-Scale Continues Expression Data

LAGE is a systematic framework developed in Java. The motivation of LAGE is to provide a scalable and parallel solution to reconstruct Gene Regulatory Networks (GRNs) from continuous gene expression data for very large amount of genes. The basic idea of our framework is motivated by the philosophy of divideand-conquer. Specifically, LAGE recursively partitions genes into multiple overlapping communities with much smaller sizes, learns intra-community GRNs respectively before merge them altogether. Besides, the complete information of overlapping communities serves as the byproduct, which could be used to mine meaningful functional modules in biological networks.

cs.LG

LSBN: A Large-Scale Bayesian Structure Learning Framework for Model Averaging

The motivation for this paper is to apply Bayesian structure learning using Model Averaging in large-scale networks. Currently, Bayesian model averaging algorithm is applicable to networks with only tens of variables, restrained by its super-exponential complexity. We present a novel framework, called LSBN(Large-Scale Bayesian Network), making it possible to handle networks with infinite size by following the principle of divide-and-conquer. The method of LSBN comprises three steps. In general, LSBN first performs the partition by using a second-order partition strategy, which achieves more robust results. LSBN conducts sampling and structure learning within each overlapping community after the community is isolated from other variables by Markov Blanket. Finally LSBN employs an efficient algorithm, to merge structures of overlapping communities into a whole. In comparison with other four state-of-art large-scale network structure learning algorithms such as ARACNE, PC, Greedy Search and MMHC, LSBN shows comparable results in five common benchmark datasets, evaluated by precision, recall and f-score. What's more, LSBN makes it possible to learn large-scale Bayesian structure by Model Averaging which used to be intractable. In summary, LSBN provides an scalable and parallel framework for the reconstruction of network structures. Besides, the complete information of overlapping communities serves as the byproduct, which could be used to mine meaningful clusters in biological networks, such as protein-protein-interaction network or gene regulatory network, as well as in social network.

cs.LG