SearcharxivSearch

arXiv subjects

Zihang Wu

Publications and source records attributed to Zihang Wu.

4 recordsLinked to original sources

Text2GraphQuery-Bench: A Text to Graph Query Benchmark

Graph models are fundamental to data analysis in domains rich with complex relationships. Unlike SQL, which benefits from a rel- atively unified standard and widespread familiarity, graph query languages are diverse (e.g., Cypher, GQL, SQL/PGQ) and far less fa- miliar to most users, making them significantly harder to learn and use. Text-to-Graph-Query systems address this barrier by trans- lating natural language into executable graph queries, enabling LLMs to serve as interfaces for Graph Database Management Systems (GDBMS). Existing benchmarks are limited in language coverage, rely on rigid synthesis, and lack comprehensive evaluation. We present Text2GraphQuery-Bench, the first benchmark covering all mainstream declarative property graph query languages (Cypher, GQL, and SQL/PGQ). It contains 267,276 (Question, Graph Query) pairs across 34 databases and 13 domains. Its construction supports adaptation from heterogeneous resources and domain-aware synthesis, while its Graph-IR-based design enables rapid extension to new languages. The evaluation protocol reports Grammar, GLEU, Similarity, and EX under graph-native difficulty, question abstraction, and schema aliasing. Experiments on 8 LLMs reveal: (i) a significant language gap exists - zero-shot GQL and SQL/PGQ Grammar is far below Cypher, yet few-shot prompting largely recovers it; (ii) fine-tuning an 8B model reaches or exceeds zero-shot large models, indicating unfamiliarity - rather than model capacity - is the primary barrier; (iii) as supervision increases, syntax errors recede, shifting bottlenecks to aggregation logic in GQL and schema linking in SQL/PGQ; (iv) higher question abstraction degrades EX due to intent-to-schema grounding issues, while schema aliasing has minimal impact; (v) EX consistently degrades from Easy to Extra Hard, with Extra Hard remaining a persistent bottleneck. *(Due to arXiv constraints, this abstract is shortened. See PDF for the full version.)*

cs.AI

Sublinear Data Structures for Nearest Neighbor in Ultra High Dimensions

Geometric data structures have been extensively studied in the regime where the dimension is much smaller than the number of input points. But in many scenarios in Machine Learning, the dimension can be much higher than the number of points and can be so high that the data structure might be unable to read and store all coordinates of the input and query points. Inspired by these scenarios and related studies in feature selection and explainable clustering, we initiate the study of geometric data structures in this ultra-high dimensional regime. Our focus is the {\em approximate nearest neighbor} problem. In this problem, we are given a set of $n$ points $C\subseteq \mathbb{R}^d$ and have to produce a {\em small} data structure that can {\em quickly} answer the following query: given $q\in \mathbb{R}^d$, return a point $c\in C$ that is approximately nearest to $q$. The main question in this paper is: {\em Is there a data structure with sublinear ($o(nd)$) space and sublinear ($o(d)$) query time when $d\gg n$?} In this paper, we answer this question affirmatively. We present $(1+\epsilon)$-approximation data structures with the following guarantees. For $\ell_1$- and $\ell_2$-norm distances: $\tilde O(n \log(d)/\mathrm{poly}(\epsilon))$ space and $\tilde O(n/\mathrm{poly}(\epsilon))$ query time. We show that these space and time bounds are tight up to $\mathrm{poly}{(\log n/\epsilon)}$ factors. For $\ell_p$-norm distances: $\tilde O(n^2 \log(d) (\log\log (n)/\epsilon)^p)$ space and $\tilde O\left(n(\log\log (n)/\epsilon)^p\right)$ query time. Via simple reductions, our data structures imply sublinear-in-$d$ data structures for some other geometric problems; e.g. approximate orthogonal range search, furthest neighbor, and give rise to a sublinear $O(1)$-approximate representation of $k$-median and $k$-means clustering.

cs.DS

Maintaining Expander Decompositions via Sparse Cuts

In this article, we show that the algorithm of maintaining expander decompositions in graphs undergoing edge deletions directly by removing sparse cuts repeatedly can be made efficient. Formally, for an $m$-edge undirected graph $G$, we say a cut $(S, \overline{S})$ is $ϕ$-sparse if $|E_G(S, \overline{S})| < ϕ\cdot \min\{vol_G(S), vol_G(\overline{S})\}$. A $ϕ$-expander decomposition of $G$ is a partition of $V$ into sets $X_1, X_2, \ldots, X_k$ such that each cluster $G[X_i]$ contains no $ϕ$-sparse cut (meaning it is a $ϕ$-expander) with $\tilde{O}(ϕm)$ edges crossing between clusters. A natural way to compute a $ϕ$-expander decomposition is to decompose clusters by $ϕ$-sparse cuts until no such cut is contained in any cluster. We show that even in graphs undergoing edge deletions, a slight relaxation of this meta-algorithm can be implemented efficiently with amortized update time $m^{o(1)}/ϕ^2$. Our approach naturally extends to maintaining directed $ϕ$-expander decompositions and $ϕ$-expander hierarchies and thus gives a unifying framework while having simpler proofs than previous state-of-the-art work. In all settings, our algorithm matches the run-times of previous algorithms up to subpolynomial factors. Moreover, our algorithm provides stronger guarantees for $ϕ$-expander decompositions. For example, for graphs undergoing edge deletions, our approach is the first to maintain a dynamic expander decomposition where each updated decomposition is a refinement of the previous decomposition, and our approach is the first to guarantee a sublinear $ϕm^{1+o(1)}$ bound on the total number of edges that cross between clusters across the entire sequence of dynamic updates.

cs.DS

Renormalizing Antiferroelectric Nanostripes in $β'-\mathrm{In}_{2}\mathrm{Se}_{3}$ via Optomechanics

Antiferroelectric (AFE) materials have received tremendous attention owing to their high energy conversion efficiency and good tunability. Recently, an exotic two-dimensional (2D) AFE material, $β'-\mathrm{In}_{2}\mathrm{Se}_{3}$ monolayer that could host atomically thin AFE nanostripe domains has been experimentally synthesized and theoretically examined. In this work, we apply first-principles calculations and theoretical estimations to predict that light irradiation can control the nanostripe width of such a system. We suggest that an intermediate near-infrared light (below bandgap) could effectively harness the thermodynamic Gibbs free energy, and the AFE nanostripe width will gradually reduce. We also propose to use an above bandgap linearly polarized light to generate AFE nanostripespecific photocurrent, providing an all-optical pump-probe setup for such AFE nanostripe width phase transitions.

cond-mat.mtrl-sci