SearcharxivSearch

arXiv subjects

Jianxing Zhao

Publications and source records attributed to Jianxing Zhao.

7 recordsLinked to original sources

Constraint-Aware Synthetic Tabular Data Generation via Inter-Column Constraint Discovery with LLM Agents

Generating structurally valid synthetic tabular data remains difficult: outputs with high statistical fidelity and downstream utility can still violate semantically meaningful domain constraints. We study the discovery and enforcement of three complementary inter-column constraint families---equations, linear inequalities, and logical dependencies. Our unified tool-grounded workflow represents all three as machine-executable hypotheses and applies a common interface for full-table validation, deterministic diagnosis, and counterexample-guided revision. A generator-agnostic postprocessor coordinates family-specific repairs on outputs from unchanged tabular generators. Across curated behavioral audits and end-to-end evaluations, the complete workflow improves held-out violation detection over one-shot direct prompting, while postprocessing yields zero measured violations for every retained, applicable constraint, improves downstream utility on most datasets, and largely preserves univariate marginals.

cs.AI

Embedding Space Selection for Detecting Memorization and Fingerprinting in Generative Models

In the rapidly evolving landscape of artificial intelligence, generative models such as Generative Adversarial Networks (GANs) and Diffusion Models have become cornerstone technologies, driving innovation in diverse fields from art creation to healthcare. Despite their potential, these models face the significant challenge of data memorization, which poses risks to privacy and the integrity of generated content. Among various metrics of memorization detection, our study delves into the memorization scores calculated from encoder layer embeddings, which involves measuring distances between samples in the embedding spaces. Particularly, we find that the memorization scores calculated from layer embeddings of Vision Transformers (ViTs) show an notable trend - the latter (deeper) the layer, the less the memorization measured. It has been found that the memorization scores from the early layers' embeddings are more sensitive to low-level memorization (e.g. colors and simple patterns for an image), while those from the latter layers are more sensitive to high-level memorization (e.g. semantic meaning of an image). We also observe that, for a specific model architecture, its degree of memorization on different levels of information is unique. It can be viewed as an inherent property of the architecture. Building upon this insight, we introduce a unique fingerprinting methodology. This method capitalizes on the unique distributions of the memorization score across different layers of ViTs, providing a novel approach to identifying models involved in generating deepfakes and malicious content. Our approach demonstrates a marked 30% enhancement in identification accuracy over existing baseline methods, offering a more effective tool for combating digital misinformation.

cs.LG

The Weighted Arithmetic Mean-Geometric Mean Inequality is Equivalent to the Hölder Inequality

In the current note, we investigate the mathematical relations among the weighted arithmetic mean-geometric mean (AM-GM) inequality, the Hölder inequality and the weighted power-mean inequality. Meanwhile, the proofs of mathematical equivalence among the weighted AM-GM inequality, the weighted power-mean inequality and the Hölder inequality are fully achieved. The new results are more generalized than those of previous studies.

math.FA

CFS: A Distributed File System for Large Scale Container Platforms

We propose CFS, a distributed file system for large scale container platforms. CFS supports both sequential and random file accesses with optimized storage for both large files and small files, and adopts different replication protocols for different write scenarios to improve the replication performance. It employs a metadata subsystem to store and distribute the file metadata across different storage nodes based on the memory usage. This metadata placement strategy avoids the need of data rebalancing during capacity expansion. CFS also provides POSIX-compliant APIs with relaxed semantics and metadata atomicity to improve the system performance. We performed a comprehensive comparison with Ceph, a widely-used distributed file system on container platforms. Our experimental results show that, in testing 7 commonly used metadata operations, CFS gives around 3 times performance boost on average. In addition, CFS exhibits better random-read/write performance in highly concurrent environments with multiple clients and processes.

cs.DC

A new $Z$-eigenvalue inclusion theorem for tensors

A new $Z$-eigenvalue inclusion theorem for tensors is given and proved to be tighter than those in [G. Wang, G.L. Zhou, L. Caccetta, $Z$-eigenvalue inclusion theorems for tensors, Discrete and Continuous Dynamical Systems Series B,22(1) (2017) 187--198]. Based on this set, a sharper upper bound for the $Z$-spectral radius of weakly symmetric nonnegative tensors is obtained. Finally, numerical examples are given to show the effectiveness of the proposed bound.

math.NA

A tighter $Z$-eigenvalue localization set for tensors and its applications

A new $Z$-eigenvalue localization set for tensors is given and proved to be tighter than those presented by Wang \emph{et al}. (Discrete and Continuous Dynamical Systems Series B 22(1): 187-198, 2017) and Zhao (J. Inequal. Appl., to appear, 2017). As an application, a sharper upper bound for the $Z$-spectral radius of weakly symmetric nonnegative tensors is obtained. Finally, numerical examples are given to verify the theoretical results.

math.NA

Lower bounds of the minimum eigenvalue for $M$-matrices

Some monotone increasing sequences of the lower bounds for the minimum eigenvalue of $M$-matrices are given. It is proved that these sequences are convergent and improve some existing results. Numerical examples show that these sequences are more accurate than some existing results and could reach the true value of the minimum eigenvalue in some cases.

math.NA