SearcharxivSearch

arXiv subjects

Aoqian Zhang

Publications and source records attributed to Aoqian Zhang.

15 recordsLinked to original sources

TLSQL: Table Learning Structured Query Language

Table learning has recently emerged as an important paradigm at the intersection of database systems and machine learning. However, applying table learning in practice often requires exporting data from databases and building complex external machine learning pipelines, which disrupts the SQL-centric workflow commonly used by database practitioners. We present TLSQL (Table Learning Structured Query Language), a lightweight SQL-like interface for specifying table learning tasks over SQL-centric data systems. TLSQL allows users to declaratively define predictive tasks using three simple constructs: PREDICT VALUE, TRAIN WITH, and VALIDATE WITH. TLSQL specifications are compiled into standard SQL queries executed by the database engine and task descriptions consumed by downstream table learning frameworks. This design allows users to focus on modeling rather than low-level data preparation and pipeline orchestration. Our demonstration shows that TLSQL enables end-to-end multi-table learning workflows while preserving familiar SQL-based data processing environments. Our code is available at https://github.com/tlsql-project/tlsql.

cs.DB

LakeMLB: Data Lake Machine Learning Benchmark

Data lakes have become a fundamental platform for large-scale machine learning by enabling flexible management of heterogeneous data. Despite their growing importance, standardized benchmarks for evaluating machine learning performance in data lake environments remain scarce. To address this gap, we present LakeMLB (Data Lake Machine Learning Benchmark), the first benchmark designed for multi-table machine learning in data lakes. LakeMLB focuses on two representative scenarios, Union and Join, and provides six real-world datasets spanning diverse domains. It supports three representative multi-table learning paradigms: pre-training, data augmentation, and feature augmentation, together with standardized data splits and evaluation protocols. We conduct extensive experiments with state-of-the-art tabular learning methods and provide insights into their performance across different data lake scenarios. We release both datasets and code to facilitate rigorous research on machine learning in data lake ecosystems; the benchmark is available at https://github.com/zhengwang100/LakeMLB.

cs.LG

Ising Dirac fermions across a topological phase transition

Dirac fermions have attracted significant interest due to their relativistic dispersions and close connections to topological physics, yet they are generally expected to be gapped in two-dimensional systems with strong Ising spin orbit coupling, making their realization in such materials an outstanding challenge. Here we report the emergence of six fold degenerate Dirac fermions in an Ising moire system across a quantum spin Hall transition in twisted WSe2. In a 3.65 degree device, we observe a quantum spin Hall phase at high electric fields with nearly quantized resistance h/(2e2), and a Dirac semimetal phase over a broad range of electric fields near zero field. Magnetotransport measurements of the Dirac phase exhibit a half-integer Landau fan sequence, characteristic of Dirac fermions, with six-fold degeneracy on the hole-doped side and two fold degeneracy on the electron-doped side. Temperature dependence shows weakly metallic behavior consistent with a semimetallic state. Our twist-angle-dependent transport measurements map out a complete phase diagram and identify a critical twist angle of 3.3 degree, establishing the phase boundary between the quantum spin Hall and Dirac semimetal regimes. Our work establishes a new route to realizing Dirac fermions in strongly spin orbit coupled moire systems through a topological phase transition, providing a promising platform for high mobility spintronics.

cond-mat.mes-hall

Observation of a Mott quantum spin Hall insulator in twisted WSe2

Quantum spin Hall (QSH) insulators and Mott insulators are conventionally regarded as distinct insulating phases, arising from band topology and strong Coulomb interactions, respectively. Here, we report the observation of QSH edge transport in a magnetic-field-stabilized Mott insulating state at half filling of the second moire band in a 2.29 degree twisted WSe2 device. This state exhibits a resistance plateau identical to that of the single-particle QSH state at full filling of the first moire valence band, indicating the same number of helical edge channels. Electrical transport measurements reveal nearly quantized resistance that is insensitive to vertical electric field, out-of-plane magnetic field, and temperature below 5 K. Pronounced nonlocal transport and strong negative in-plane magnetoconductance further support helical edge conduction, establishing robust edge transport in the strongly correlated regime. Temperature-dependent Hall measurements reveal a characteristic temperature scale of approximately 10 K, corresponding to an energy scale of about 1 meV. Our results demonstrate that spin-conserved QSH edge states can persist in a half filled, strongly correlated insulating phase and under external magnetic field, opening a route toward interaction-resilient topological transport in moire quantum materials.

cond-mat.mes-hall

Correlation enhanced resistance hysteresis near half filling in MoS2/WSe2 heterobilayer

Ferroelectricity, typically arising from ionic displacements in noncentrosymmetric lattices, enabling applications in memory devices and sensors. Recent advances in two-dimensional materials and van der Waals heterostructures have revealed novel ferroelectric phenomena, including sliding ferroelectricity and correlation-driven ferroelectricity in moire superlattices. In this work, we fabricate and study a MoS2/WSe2 moire superlattice device exhibiting a high field-effect mobility of 17,650 $cm^2V^{-1}s^{-1}$. Electrical transport measurements reveal correlated insulating states accompanied by a prominent and reproducible resistance hysteresis near half filling. Temperature and displacement field dependence further confirms the correlation-enhanced nature of the hysteresis. Our analysis suggests that displacement field-induced metal-to-insulator transition at correlated insulating state coupled with interfacial dipoles enables the observed resistance hysteresis. These results establish correlation enhanced resistance hysteresis near half filling in a MoS2/WSe2 heterobilayer, offering opportunities for exploring emergent quantum phases and device functionalities.

cond-mat.mes-hall

SAC-Loco: Safe and Adjustable Compliant Quadrupedal Locomotion

Quadruped robots are designed to achieve agile and robust locomotion by drawing inspiration from legged animals. However, most existing control methods for quadruped robots lack a key capacity observed in animals: the ability to exhibit diverse compliance behaviors while ensuring stability when experiencing external forces. In particular, achieving adjustable compliance while maintaining robust safety under force disturbances remains a significant challenge. In this work, we propose a safety aware compliant locomotion framework that integrates adjustable disturbance compliance with robust failure prevention. We first train a force compliant policy with adjustable compliance levels using a teacher student reinforcement learning framework, allowing deployment without explicit force sensing. To handle disturbances beyond the limits of compliant control, we develop a safety oriented policy for rapid recovery and stabilization. Finally, we introduce a learned safety critic that monitors the robot's safety in real time and coordinates between compliant locomotion and recovery behaviors. Together, this framework enables quadruped robots to achieve smooth force compliance and robust safety under a wide range of external force disturbances.

cs.RO

Soft Responsive Materials Enhance Humanoid Safety

Humanoid robots are envisioned as general-purpose platforms in human-centered environments, yet their deployment is limited by vulnerability to falls and the risks posed by rigid metal-plastic structures to people and surroundings. We introduce a soft-rigid co-design framework that leverages non-Newtonian fluid-based soft responsive materials to enhance humanoid safety. The material remains compliant during normal interaction but rapidly stiffens under impact, absorbing and dissipating fall-induced forces. Physics-based simulations guide protector placement and thickness and enable learning of active fall policies. Applied to a 42 kg life-size humanoid, the protector markedly reduces peak impact and allows repeated falls without hardware damage, including drops from 3 m and tumbles down long staircases. Across diverse scenarios, the approach improves robot robustness and environmental safety. By uniting responsive materials, structural co-design, and learning-based control, this work advances interact-safe, industry-ready humanoid robots.

cs.RO

Sync Without Guesswork: Incomplete Time Series Alignment

Multivariate time series alignment is critical for ensuring coherent analysis across variables, but missing values and timestamp inconsistencies make this task highly challenging. Existing approaches often rely on prior imputation, which can introduce errors and lead to suboptimal alignments. To address these limitations, we propose a constraint-based alignment framework for incomplete multivariate time series that avoids imputation and ensures temporal and structural consistency. We further design efficient approximation algorithms to balance accuracy and scalability. Experiments on multiple real-world datasets demonstrate that our approach achieves superior alignment quality compared to existing methods under varying missing rates. Our contributions include: (1) formally defining incomplete multiple temporal data alignment problem; (2) proposing three approximation algorithms balancing accuracy and efficiency; and (3) validating our approach on diverse real-world datasets, where it consistently outperforms existing methods in alignment accuracy and the number of aligned tuples.

cs.DB

CARPO: Leveraging Listwise Learning-to-Rank for Context-Aware Query Plan Optimization

Efficient data processing is increasingly vital, with query optimizers playing a fundamental role in translating SQL queries into optimal execution plans. Traditional cost-based optimizers, however, often generate suboptimal plans due to flawed heuristics and inaccurate cost models, leading to the emergence of Learned Query Optimizers (LQOs). To address challenges in existing LQOs, such as the inconsistency and suboptimality inherent in pairwise ranking methods, we introduce CARPO, a generic framework leveraging listwise learning-to-rank for context-aware query plan optimization. CARPO distinctively employs a Transformer-based model for holistic evaluation of candidate plan sets and integrates a robust hybrid decision mechanism, featuring Out-Of-Distribution (OOD) detection with a top-k fallback strategy to ensure reliability. Furthermore, CARPO can be seamlessly integrated with existing plan embedding techniques, demonstrating strong adaptability. Comprehensive experiments on TPC-H and STATS benchmarks demonstrate that CARPO significantly outperforms both native PostgreSQL and Lero, achieving a Top-1 Rate of 74.54% on the TPC-H benchmark compared to Lero's 3.63%, and reducing the total execution time to 3719.16 ms compared to PostgreSQL's 22577.87 ms.

cs.DB

Multivariate Time Series Cleaning under Speed Constraints

Errors are common in time series due to unreliable sensor measurements. Existing methods focus on univariate data but do not utilize the correlation between dimensions. Cleaning each dimension separately may lead to a less accurate result, as some errors can only be identified in the multivariate case. We also point out that the widely used minimum change principle is not always the best choice. Instead, we try to change the smallest number of data to avoid a significant change in the data distribution. In this paper, we propose MTCSC, the constraint-based method for cleaning multivariate time series. We formalize the repair problem, propose a linear-time method to employ online computing, and improve it by exploiting data trends. We also support adaptive speed constraint capturing. We analyze the properties of our proposals and compare them with SOTA methods in terms of effectiveness, efficiency versus error rates, data sizes, and applications such as classification. Experiments on real datasets show that MTCSC can have higher repair accuracy with less time consumption. Interestingly, it can be effective even when there are only weak or no correlations between the dimensions.

cs.DB

A Survey on Sampling and Profiling over Big Data (Technical Report)

Due to the development of internet technology and computer science, data is exploding at an exponential rate. Big data brings us new opportunities and challenges. On the one hand, we can analyze and mine big data to discover hidden information and get more potential value. On the other hand, the 5V characteristic of big data, especially Volume which means large amount of data, brings challenges to storage and processing. For some traditional data mining algorithms, machine learning algorithms and data profiling tasks, it is very difficult to handle such a large amount of data. The large amount of data is highly demanding hardware resources and time consuming. Sampling methods can effectively reduce the amount of data and help speed up data processing. Hence, sampling technology has been widely studied and used in big data context, e.g., methods for determining sample size, combining sampling with big data processing frameworks. Data profiling is the activity that finds metadata of data set and has many use cases, e.g., performing data profiling tasks on relational data, graph data, and time series data for anomaly detection and data repair. However, data profiling is computationally expensive, especially for large data sets. Therefore, this paper focuses on researching sampling and profiling in big data context and investigates the application of sampling in different categories of data profiling tasks. From the experimental results of these studies, the results got from the sampled data are close to or even exceed the results of the full amount of data. Therefore, sampling technology plays an important role in the era of big data, and we also have reason to believe that sampling technology will become an indispensable step in big data processing in the future.

cs.DB

A Survey of Approximate Quantile Computation on Large-scale Data (Technical Report)

As data volume grows extensively, data profiling helps to extract metadata of large-scale data. However, one kind of metadata, order statistics, is difficult to be computed because they are not mergeable or incremental. Thus, the limitation of time and memory space does not support their computation on large-scale data. In this paper, we focus on an order statistic, quantiles, and present a comprehensive analysis of studies on approximate quantile computation. Both deterministic algorithms and randomized algorithms that compute approximate quantiles over streaming models or distributed models are covered. Then, multiple techniques for improving the efficiency and performance of approximate quantile algorithms in various scenarios, such as skewed data and high-speed data streams, are presented. Finally, we conclude with coverage of existing packages in different languages and with a brief discussion of the future direction in this area.

cs.DS

Learning Individual Models for Imputation (Technical Report)

Missing numerical values are prevalent, e.g., owing to unreliable sensor reading, collection and transmission among heterogeneous sources. Unlike categorized data imputation over a limited domain, the numerical values suffer from two issues: (1) sparsity problem, the incomplete tuple may not have sufficient complete neighbors sharing the same/similar values for imputation, owing to the (almost) infinite domain; (2) heterogeneity problem, different tuples may not fit the same (regression) model. In this study, enlightened by the conditional dependencies that hold conditionally over certain tuples rather than the whole relation, we propose to learn a regression model individually for each complete tuple together with its neighbors. Our IIM, Imputation via Individual Models, thus no longer relies on sharing similar values among the k complete neighbors for imputation, but utilizes their regression results by the aforesaid learned individual (not necessary the same) models. Remarkably, we show that some existing methods are indeed special cases of our IIM, under the extreme settings of the number l of learning neighbors considered in individual learning. In this sense, a proper number l of neighbors is essential to learn the individual models (avoid over-fitting or under-fitting). We propose to adaptively learn individual models over various number l of neighbors for different complete tuples. By devising efficient incremental computation, the time complexity of learning a model reduces from linear to constant. Experiments on real data demonstrate that our IIM with adaptive learning achieves higher imputation accuracy than the existing approaches.

cs.DB

Time Series Data Cleaning: From Anomaly Detection to Anomaly Repairing (Technical Report)

Errors are prevalent in time series data, such as GPS trajectories or sensor readings. Existing methods focus more on anomaly detection but not on repairing the detected anomalies. By simply filtering out the dirty data via anomaly detection, applications could still be unreliable over the incomplete time series. Instead of simply discarding anomalies, we propose to (iteratively) repair them in time series data, by creatively bonding the beauty of temporal nature in anomaly detection with the widely considered minimum change principle in data repairing. Our major contributions include: (1) a novel framework of iterative minimum repairing (IMR) over time series data, (2) explicit analysis on convergence of the proposed iterative minimum repairing, and (3) efficient estimation of parameters in each iteration. Remarkably, with incremental computation, we reduce the complexity of parameter estimation from O(n) to O(1). Experiments on real datasets demonstrate the superiority of our proposal compared to the state-of-the-art approaches. In particular, we show that (the proposed) repairing indeed improves the time series classification application.

cs.DB