SearcharxivSearch

arXiv subjects

Jian Tan

Publications and source records attributed to Jian Tan.

At least 19 recordsLinked to original sources

Dual spaces and T1 theorem of variable Dunkl--Hardy spaces

In this paper, we identify the dual spaces of variable Dunkl--Hardy spaces using a discrete Dunkl--Calder\'on reproducing formula and the Frazier and Jawerth's method adapted to the variable exponent setting. Furthermore, we prove a $T1$ theorem for Dunkl--Calder\'on--Zygmund operators acting on variable Dunkl--Hardy spaces and on their dual spaces via using almost orthogonal estimate, the duality result and the density argument.

math.CA

Variable Hardy spaces in the Dunkl Setting

In this paper, we aim to define a new variable Hardy spaces in the Dunkl setting by the discrete Littlewood--Paley square functions. We also consider the atomic decomposition characterization for the new variable Hardy spaces with the help of the Littlewood--Paley theory and Hardy spaces on spaces of homogeneous type in the sense of Coifman and Weiss. As applications, we obtain the variable Hardy spaces boundedness for Dunkl--Calder\'on--Zygmund singular integrals involving by the Euclidean metric and the Dunkl ``metric'' which is associated with the finite reflection groups on the Euclidean space.

math.CA

HYVE: Hybrid Views for LLM Context Engineering over Machine Data

Machine data is central to observability and diagnosis in modern computing systems, appearing in logs, metrics, telemetry traces, and configuration snapshots. When provided to large language models (LLMs), this data typically arrives as a mixture of natural language and structured payloads such as JSON or Python/AST literals. Yet LLMs remain brittle on such inputs, particularly when they are long, deeply nested, and dominated by repetitive structure. We present HYVE (HYbrid ViEw), a framework for LLM context engineering for inputs containing large machine-data payloads, inspired by database management principles. HYVE surrounds model invocation with coordinated preprocessing and postprocessing, centered on a request-scoped datastore augmented with schema information. During preprocessing, HYVE detects repetitive structure in raw inputs, materializes it in the datastore, transforms it into hybrid columnar and row-oriented views, and selectively exposes only the most relevant representation to the LLM. During postprocessing, HYVE either returns the model output directly, queries the datastore to recover omitted information, or performs a bounded additional LLM call for SQL-augmented semantic synthesis. We evaluate HYVE on diverse real-world workloads spanning knowledge QA, chart generation, anomaly detection, and multi-step network troubleshooting. Across these benchmarks, HYVE reduces token usage by 50-90% while maintaining or improving output quality. On structured generation tasks, it improves chart-generation accuracy by up to 132% and reduces latency by up to 83%. Overall, HYVE offers a practical approximation to an effectively unbounded context window for prompts dominated by large machine-data payloads.

cs.AI

The complete boundedness of singular integrals on weighted flag and product Hardy space

It is known that product singular integrals are bounded on product Hardy spaces and that flag singular integrals are bounded on flag Hardy spaces. The purpose of this paper is to obtain the complete boundedness of singular integrals on weighted flag and product Hardy spaces. In particular, we prove the boundedness of one-parameter singular integrals on both weighted flag and product Hardy spaces, as well as the boundedness of flag singular integrals on weighted product Hardy spaces.

math.CA

Hardy type spaces estimates for multilinear fractional integral operators

In this paper, we prove the boundedness of multilinear fractional integral operators from products of Hardy spaces associated with ball quasi-Banach function spaces into their corresponding ball quasi-Banach function spaces. As applications, we establish the boundedness of these operators on various function spaces, including weighted Hardy spaces, variable Hardy spaces, mixed-norm Hardy spaces, Hardy--Lorentz spaces, and Hardy--Orlicz spaces. Notably, several of these results are new, even in special cases, and extend the existing theory of multilinear operators in the context of generalized Hardy spaces.

math.FA

Multilinear operators on Hardy spaces associated with ball quasi-Banach function spaces

This paper establishes that multilinear Calder\'on--Zygmund operators and their maximal operators are bounded on Hardy spaces associated with ball quasi-Banach function spaces. Moreover, we also obtain the boundedness of multilinear pseudo-differential operators on local Hardy spaces associated with ball quasi-Banach function spaces. Since these (local) Hardy type spaces encompass a wide range of classical (local) Hardy-type spaces including weighted (local) Hardy spaces, variable (local) Hardy space, (local) Hardy--Morrey space, mixed-norm (local) Hardy space, (local) Hardy--Lorentz space and (local) Hardy--Orlicz spaces, the results presented in this paper are highly general and essentially improve the existing results.

math.FA

Automatic Database Configuration Debugging using Retrieval-Augmented Language Models

Database management system (DBMS) configuration debugging, e.g., diagnosing poorly configured DBMS knobs and generating troubleshooting recommendations, is crucial in optimizing DBMS performance. However, the configuration debugging process is tedious and, sometimes challenging, even for seasoned database administrators (DBAs) with sufficient experience in DBMS configurations and good understandings of the DBMS internals (e.g., MySQL or Oracle). To address this difficulty, we propose Andromeda, a framework that utilizes large language models (LLMs) to enable automatic DBMS configuration debugging. Andromeda serves as a natural surrogate of DBAs to answer a wide range of natural language (NL) questions on DBMS configuration issues, and to generate diagnostic suggestions to fix these issues. Nevertheless, directly prompting LLMs with these professional questions may result in overly generic and often unsatisfying answers. To this end, we propose a retrieval-augmented generation (RAG) strategy that effectively provides matched domain-specific contexts for the question from multiple sources. They come from related historical questions, troubleshooting manuals and DBMS telemetries, which significantly improve the performance of configuration debugging. To support the RAG strategy, we develop a document retrieval mechanism addressing heterogeneous documents and design an effective method for telemetry analysis. Extensive experiments on real-world DBMS configuration debugging datasets show that Andromeda significantly outperforms existing solutions.

cs.DB

Local Hardy spaces associated with ball quasi-Banach function spaces and their dual spaces

Let $X$ be a ball quasi-Banach function space on $\mathbb R^{n}$ and $h_{X}(\mathbb R^{n})$ the local Hardy space associated with $X$. In this paper, under some reasonable assumptions on $X$, the infinite and finite atomic decompositions for the local Hardy space $h_{X}(\mathbb R^{n})$ are established directly, without relying on the relation between $H_{X}(\mathbb R^{n})$ and $h_{X}(\mathbb R^{n})$. Moreover, we apply the finite atomic decomposition to obtain the dual space of the local Hardy space $h_{X}(\mathbb R^{n})$. Especially, the above results can be applied to several specific ball quasi-Banach function spaces, demonstrating their wide range of applications.

math.FA

Weak type estimates for Bochner--Riesz means on Hardy-type spaces associated with ball quasi-Banach function spaces

Let $X\left(\mathbb{R}^{n}\right)$ be a ball quasi-Banach function space on $\mathbb{R}^{n}$, $WX\left(\mathbb{R}^{n}\right)$ be the weak ball quasi-Banach function space on $\mathbb{R}^{n}$, $H_{X}\left(\mathbb{R}^{n}\right)$ be the Hardy space associated with $X\left(\mathbb{R}^{n}\right)$ and $WH_{X}\left(\mathbb{R}^{n}\right)$ be the weak Hardy space associated with $X\left(\mathbb{R}^{n}\right)$. In this paper, we obtain the boundedness of the Bochner--Riesz means and the maximal Bochner--Riesz means from $H_{X}\left(\mathbb{R}^{n}\right)$ to $WH_{X}\left(\mathbb{R}^{n}\right)$ or $WX\left(\mathbb{R}^{n}\right)$, which includes the critical case. Moreover, we apply these results to several examples of ball quasi-Banach function spaces, namely, weighted Lebesgue spaces, Herz spaces, Lorentz spaces, variable Lebesgue spaces and Morrey spaces. This shows that all the results obtained in this article are of wide applications, and more applications of these results are predictable.

math.FA

The Atomic Characterization of Weighted Local Hardy Spaces and Its Applications

The purpose of this paper is to obtain atomic decomposition characterization of the weighted local Hardy space $h_{\omega}^{p}(\mathbb {R}^{n})$ with $\omega\in A_{\infty}(\mathbb {R}^{n})$. We apply the discrete version of Calder\'on's identity and the weighted Littlewood--Paley--Stein theory to prove that $h_{\omega}^{p}(\mathbb {R}^{n})$ coincides with the weighted$\text{-}(p,q,s)$ atomic local Hardy space $h_{\omega,atom}^{p,q,s}(\mathbb {R}^{n})$ for $0<p<\infty$. The atomic decomposition theorems in our paper improve the previous atomic decomposition results of local weighted Hardy spaces in the literature. As applications, we derive the boundedness of inhomogeneous Calder\'on--Zygmund singular integrals and local fractional integrals on weighted local Hardy spaces.

math.CA

Product Hardy spaces meet ball quasi-Banach function spaces

The main purpose of this paper is to develop the theory of product Hardy spaces built on Banach lattices on $\mathbb R^n\times\mathbb R^m$. First we introduce new product Hardy spaces ${H}_X(\mathbb R^n\times\mathbb R^m)$ associated with ball quasi-Banach function spaces $X(\mathbb R^n\times\mathbb R^m)$ via applying the Littlewood-Paley-Stein theory. Then we establish a decomposition theorem for ${H}_X(\mathbb R^n\times\mathbb R^m)$ in terms of the discrete Calder\'on's identity. Moreover, we explore some useful and general extrapolation theorems of Rubio de Francia on $X(\mathbb R^n\times\mathbb R^m)$ and give some applications to boundedness of operators. Finally, we conclude that the two-parameter singular integral operators $\widetilde T$ are bounded from ${H}_X(\mathbb R^n\times\mathbb R^m)$ to itself and bounded from ${H}_X(\mathbb R^n\times\mathbb R^m)$ to $X(\mathbb R^n\times\mathbb R^m)$ via extrapolation. The main results obtained in this paper have a wide range of generality. Especially, we can apply these results to many concrete examples of ball quasi-Banach function spaces, including product Herz spaces, weighted product Morrey spaces and product Musielak--Orlicz--type spaces.

math.FA

OneShotSTL: One-Shot Seasonal-Trend Decomposition For Online Time Series Anomaly Detection And Forecasting

Seasonal-trend decomposition is one of the most fundamental concepts in time series analysis that supports various downstream tasks, including time series anomaly detection and forecasting. However, existing decomposition methods rely on batch processing with a time complexity of O(W), where W is the number of data points within a time window. Therefore, they cannot always efficiently support real-time analysis that demands low processing delay. To address this challenge, we propose OneShotSTL, an efficient and accurate algorithm that can decompose time series online with an update time complexity of O(1). OneShotSTL is more than $1,000$ times faster than the batch methods, with accuracy comparable to the best counterparts. Extensive experiments on real-world benchmark datasets for downstream time series anomaly detection and forecasting tasks demonstrate that OneShotSTL is from 10 to over 1,000 times faster than the state-of-the-art methods, while still providing comparable or even better accuracy.

cs.LG

A Unified and Efficient Coordinating Framework for Autonomous DBMS Tuning

Recently using machine learning (ML) based techniques to optimize modern database management systems has attracted intensive interest from both industry and academia. With an objective to tune a specific component of a DBMS (e.g., index selection, knobs tuning), the ML-based tuning agents have shown to be able to find better configurations than experienced database administrators. However, one critical yet challenging question remains unexplored -- how to make those ML-based tuning agents work collaboratively. Existing methods do not consider the dependencies among the multiple agents, and the model used by each agent only studies the effect of changing the configurations in a single component. To tune different components for DBMS, a coordinating mechanism is needed to make the multiple agents cognizant of each other. Also, we need to decide how to allocate the limited tuning budget among the agents to maximize the performance. Such a decision is difficult to make since the distribution of the reward for each agent is unknown and non-stationary. In this paper, we study the above question and present a unified coordinating framework to efficiently utilize existing ML-based agents. First, we propose a message propagation protocol that specifies the collaboration behaviors for agents and encapsulates the global tuning messages in each agent's model. Second, we combine Thompson Sampling, a well-studied reinforcement learning algorithm with a memory buffer so that our framework can allocate budget judiciously in a non-stationary environment. Our framework defines the interfaces adapted to a broad class of ML-based tuning agents, yet simple enough for integration with existing implementations and future extensions. We show that it can effectively utilize different ML-based agents and find better configurations with 1.4~14.1X speedups on the workload execution time compared with baselines.

cs.DB

Interactive Log Parsing via Light-weight User Feedback

Template mining is one of the foundational tasks to support log analysis, which supports the diagnosis and troubleshooting of large scale Web applications. This paper develops a human-in-the-loop template mining framework to support interactive log analysis, which is highly desirable in real-world diagnosis or troubleshooting of Web applications but yet previous template mining algorithms fails to support it. We formulate three types of light-weight user feedbacks and based on them we design three atomic human-in-the-loop template mining algorithms. We derive mild conditions under which the outputs of our proposed algorithms are provably correct. We also derive upper bounds on the computational complexity and query complexity of each algorithm. We demonstrate the versatility of our proposed algorithms by combining them to improve the template mining accuracy of five representative algorithms over sixteen widely used benchmark datasets.

cs.AI

LPC-AD: Fast and Accurate Multivariate Time Series Anomaly Detection via Latent Predictive Coding

This paper proposes LPC-AD, a fast and accurate multivariate time series (MTS) anomaly detection method. LPC-AD is motivated by the ever-increasing needs for fast and accurate MTS anomaly detection methods to support fast troubleshooting in cloud computing, micro-service systems, etc. LPC-AD is fast in the sense that its reduces the training time by as high as 38.2% compared to the state-of-the-art (SOTA) deep learning methods that focus on training speed. LPC-AD is accurate in the sense that it improves the detection accuracy by as high as 18.9% compared to SOTA sophisticated deep learning methods that focus on enhancing detection accuracy. Methodologically, LPC-AD contributes a generic architecture LPC-Reconstruct for one to attain different trade-offs between training speed and detection accuracy. More specifically, LPC-Reconstruct is built on ideas from autoencoder for reducing redundancy in time series, latent predictive coding for capturing temporal dependence in MTS, and randomized perturbation for avoiding overfitting of anomalous dependence in the training data. We present simple instantiations of LPC-Reconstruct to attain fast training speed, where we propose a simple randomized perturbation method. The superior performance of LPC-AD over SOTA methods is validated by extensive experiments on four large real-world datasets. Experiment results also show the necessity and benefit of each component of the LPC-Reconstruct architecture and that LPC-AD is robust to hyper parameters.

cs.LG

CobBO: Coordinate Backoff Bayesian Optimization with Two-Stage Kernels

Bayesian optimization is a popular method for optimizing expensive black-box functions. Yet it oftentimes struggles in high dimensions where the computation could be prohibitively heavy. To alleviate this problem, we introduce Coordinate backoff Bayesian Optimization (CobBO) with two-stage kernels. During each round, the first stage uses a simple coarse kernel that sacrifices the approximation accuracy for computational efficiency. It captures the global landscape by purposely smoothing away local fluctuations. Then, in the second stage of the same round, past observed points in the full space are projected to the selected subspace to form virtual points. These virtual points, along with the means and variances of their unknown function values estimated using the simple kernel of the first stage, are fitted to a more sophisticated kernel model in the second stage. Within the selected low dimensional subspace, the computational cost of conducting Bayesian optimization therein becomes affordable. To further enhance the performance, a sequence of consecutive observations in the same subspace are collected, which can effectively refine the approximation of the function. This refinement lasts until a stopping rule is met determining when to back off from a certain subspace and switch to another. This decoupling significantly reduces the computational burden in high dimensions, which fully leverages the observations in the whole space rather than only relying on observations in each coordinate subspace. Extensive evaluations show that CobBO finds solutions comparable to or better than other state-of-the-art methods for dimensions ranging from tens to hundreds, while reducing both the trial complexity and computational costs.

cs.LG

Towards Dynamic and Safe Configuration Tuning for Cloud Databases

Configuration knobs of database systems are essential to achieve high throughput and low latency. Recently, automatic tuning systems using machine learning methods (ML) have shown to find better configurations compared to experienced database administrators (DBAs). However, there are still gaps to apply the existing systems in production environments, especially in the cloud. First, they conduct tuning for a given workload within a limited time window and ignore the dynamicity of workloads and data. Second, they rely on a copied instance and do not consider the availability of the database when sampling configurations, making the tuning expensive, delayed, and unsafe. To fill these gaps, we propose OnlineTune, which tunes the online databases safely in changing cloud environments. To accommodate the dynamicity, OnlineTune embeds the environmental factors as context feature and adopts contextual Bayesian Optimization with context space partition to optimize the database adaptively and scalably. To pursue safety during tuning, we leverage the black-box and the white-box knowledge to evaluate the safety of configurations and propose a safe exploration strategy via subspace adaptation.%, greatly decreasing the risks of applying bad configurations. We conduct evaluations on dynamic workloads from benchmarks and real-world workloads. Compared with the state-of-the-art methods, OnlineTune achieves 14.4%~165.3% improvement on cumulative performance while reducing 91.0%~99.5% unsafe configuration recommendations.

cs.DB

Facilitating Database Tuning with Hyper-Parameter Optimization: A Comprehensive Experimental Evaluation

Recently, using automatic configuration tuning to improve the performance of modern database management systems (DBMSs) has attracted increasing interest from the database community. This is embodied with a number of systems featuring advanced tuning capabilities being developed. However, it remains a challenge to select the best solution for database configuration tuning, considering the large body of algorithm choices. In addition, beyond the applications on database systems, we could find more potential algorithms designed for configuration tuning. To this end, this paper provides a comprehensive evaluation of configuration tuning techniques from a broader perspective, hoping to better benefit the database community. In particular, we summarize three key modules of database configuration tuning systems and conduct extensive ablation studies using various challenging cases. Our evaluation demonstrates that the hyper-parameter optimization algorithms can be borrowed to further enhance the database configuration tuning. Moreover, we identify the best algorithm choices for different modules. Beyond the comprehensive evaluations, we offer an efficient and unified database configuration tuning benchmark via surrogates that reduces the evaluation cost to a minimum, allowing for extensive runs and analysis of new techniques.

cs.DB