SearcharxivSearch

arXiv subjects

Yulong Wu

Publications and source records attributed to Yulong Wu.

15 recordsLinked to original sources

SynBench: A Benchmark for Differentially Private Text Generation

Synthetic text generation with Differential Privacy (DP) guarantees emerges as a principled approach that can enable the sharing of sensitive datasets across institutional and regulatory boundaries, while bounding the risks of re-identification and membership inference. LLM-based methods deliver promising results; however, comparisons are exacerbated by differing evaluation setups and "private" datasets, potential pre-training contamination is not considered and guarantees are not verified with DP audits. To advance this field, we introduce a unified evaluation framework with standardised utility and fidelity metrics and privacy audits, encompassing nine curated datasets that capture domain-specific complexities such as technical jargon, long-context dependencies, and specialised document structures. In a large-scale empirical study, we benchmark LLM-based state-of-the-art DP text generators of varying sizes (between 1--8B). Our results indicate that DP synthetic text generation remains an unsolved challenge, with quality deteriorating more as the private datasets deviate further from the generators' pre-training corpora. Our novel synthetic text membership inference attack (MIA) explains this observation: Synthetic data quality is overestimated when LLMs have been pre-trained -- without DP -- on portions of the "private" data to be generated. Finally, our work provides the first quantitative evidence that this "public pre-training and private generation" paradigm invalidates the guaranteed privacy bounds of real-world private datasets.

cs.AI

Pay Attention to Real World Perturbations! Natural Robustness Evaluation in Machine Reading Comprehension

As neural language models achieve human-comparable performance on Machine Reading Comprehension (MRC) and see widespread adoption, ensuring their robustness in real-world scenarios has become increasingly important. Current robustness evaluation research, though, primarily develops synthetic perturbation methods, leaving unclear how well they reflect real life scenarios. Considering this, we present a framework to automatically examine MRC models on naturally occurring textual perturbations, by replacing paragraph in MRC benchmarks with their counterparts based on available Wikipedia edit history. Such perturbation type is natural as its design does not stem from an arteficial generative process, inherently distinct from the previously investigated synthetic approaches. In a large-scale study encompassing SQUAD datasets and various model architectures we observe that natural perturbations result in performance degradation in pre-trained encoder language models. More worryingly, these state-of-the-art Flan-T5 and Large Language Models (LLMs) inherit these errors. Further experiments demonstrate that our findings generalise to natural perturbations found in other more challenging MRC benchmarks. In an effort to mitigate these errors, we show that it is possible to improve the robustness to natural perturbations by training on naturally or synthetically perturbed examples, though a noticeable gap still remains compared to performance on unperturbed data.

cs.CL

Natural Context Drift Undermines the Natural Language Understanding of Large Language Models

How does the natural evolution of context paragraphs affect question answering in generative Large Language Models (LLMs)? To investigate this, we propose a framework for curating naturally evolved, human-edited variants of reading passages from contemporary QA benchmarks and for analyzing LLM performance across a range of semantic similarity scores, which quantify how closely each variant aligns with content seen during pretraining. Using this framework, we evaluate six QA datasets and eight LLMs with publicly available training data. Our experiments reveal that LLM performance declines as reading passages naturally diverge from the versions encountered during pretraining-even when the question and all necessary information remains present at inference time. For instance, average model accuracy on BoolQ drops by over 30% from the highest to lowest similarity bins, with slopes exceeding 70 across several LLMs. These findings suggest that natural text evolution poses a significant challenge to the language understanding capabilities of LLMs.

cs.CL

Incompressible Euler limit from the Boltzmann equation with Maxwell reflection boundary condition in the half-space

In this paper, we rigorously justify the incompressible Euler limit of the Boltzmann equation with general Maxwell reflection boundary condition in the half-space. The accommodation coefficient $α\in (0,1]$ is assumed to be $O(1)$. Our construction of solutions includes the interior fluid part and Knudsen-Prandtl coupled boundary layers. The corresponding solutions to the nonlinear Euler and nonlinear Prandtl systems are taken to be shear flows. Due to the presence of the nonlinear Prandtl layer, the remainder equation loses one order normal derivative. The key technical novelty lies in employing the full conservation laws to convert this loss of the normal derivative into the loss of tangential spatial derivative, avoiding any loss of regularity in time. By working within an analytic $L^2 \mbox{-} L^\infty$ framework, we establish the uniform estimate on the remainder equations, thus justify the validity of the incompressible Euler limit from the Boltzmann equation for the shear flow case.

math.AP

Compressible Navier-Stokes system with slip boundary from Boltzmann equations with reflection boundary: derivations and justifications

This is the first in a series of papers connecting the boundary conditions for the compressible Navier-Stokes system from the Boltzmann equations with the Maxwell reflection boundary. The slip boundary conditions are formally derived from the Boltzmann equation with both specular and almost specular reflection boundary conditions. That is, the accommodation coefficient $α_\eps=O(\eps^β)$ with $β>0$ or $α_\eps =0$. Here, the small number $\eps>0$ denotes the Knudsen number. The systematic formal analysis is based on the Chapman-Enskog expansion and the analysis of the Knudsen layer. In particular, for the first time, we employ the appropriate ansatz for the general $β>0$. This completes the program started in \cite{aoki2017slip}. In the second part, the compressible Navier-Stokes-Fourier approximation for the Boltzmann equation with specular reflection in general bounded domains is rigorously justified. The uniform regularity for the compressible Navier-Stokes system with the derived boundary conditions is investigated. For the remainder equation, the $L^2\mbox{-}L^6\mbox{-}L^\infty$ framework is employed to obtain uniform estimates in $\eps$.

math.AP

SR-LLM: Rethinking the Structured Representation in Large Language Model

Structured representations, exemplified by Abstract Meaning Representation (AMR), have long been pivotal in computational linguistics. However, their role remains ambiguous in the Large Language Models (LLMs) era. Initial attempts to integrate structured representation into LLMs via a zero-shot setting yielded inferior performance. We hypothesize that such a decline stems from the structure information being passed into LLMs in a code format unfamiliar to LLMs' training corpora. Consequently, we propose SR-LLM, an innovative framework with two settings to explore a superior way of integrating structured representation with LLMs from training-free and training-dependent perspectives. The former integrates structural information through natural language descriptions in LLM prompts, whereas its counterpart augments the model's inference capability through fine-tuning on linguistically described structured representations. Performance improvements were observed in widely downstream datasets, with particularly notable gains of 3.17% and 12.38% in PAWS. To the best of our knowledge, this work represents the pioneering demonstration that leveraging structural representations can substantially enhance LLMs' inference capability. We hope that our work sheds light and encourages future research to enhance the reasoning and interoperability of LLMs by structure data.

cs.CL

A Method to Decipher "Genome" from Interatomic Cohesion in the Exploration for a "Central Dogma" Replacement in Material Science

In the ball-stick model, interatomic cohesions are considered "sticks". But enormous details and features of the "sticks" are usually oversimplified as indexed quantities or equivocated as geometry characteristics. These indexed quantities or geometry characteristics not only limit the explanatory capability to a few chemical/physical aspects but also eliminate generativity for expected resemblance. And these limitations can be related to the information loss during the conversion. Herein, inspired by the central dogma, a framework is introduced to compact interatomic cohesions into a detailed residue-by-residue "genome" with matched encoding/decoding tools. The framework fuses the quantum mechanical aspects, auto feature extraction, nanostructures and/or simulations, and generative models. As a proof of concept, the realization introduced in this work adopted bosonic/fermionic features, an autoencoder with image recognition processes, Density Functional Theory simulations, and a thiolate-protected gold nanocluster dataset. After repetitive modeling, validating, and analysis based on 26,528 simulated interatomic images, the interatomic cohesion can be almost losslessly encoded into an 8-value-genome, and the genome encoder-decoder pair is also obtained. The model is then automatically extended into a generative model which converts any arbitrary 8-value-genome to a bond image.

cond-mat.mes-hall

Boltzmann boundary layer equation with Maxwell reflection boundary condition and applications to fluid limits

This paper investigates the Knudsen layer equation in half-space, arising from the hydrodynamic limit of the Boltzmann equation to fluid dynamics. We consider the Maxwell reflection boundary condition with accommodation coefficient $0<α<1$. We restrict our attention to hard sphere collisions with angular cutoff, proving the existence, uniqueness, and asymptotic behavior of the solution in $L^{\infty}_{x,v}$. Additionally, we demonstrate the application of our theorem to the hydrodynamic limit through a specific example. In this expample, we derive the boundary conditions of the fluid equations using our theorem and the symmetric properties of the Knudsen layer equation for $α\in(0,1]$ and $α=O(1)$. These derivations differs significantly from the cases of specular and almost specular reflection. This explicitly characterizes the {\em vanishing sources set} defined in \cite{jiang2024knudsenboundarylayerequations}

math.AP

Kinetic-fluid boundary layers and acoustic limit for the Boltzmann equation with general Maxwell reflection boundary condition

We prove the acoustic limit from the Boltzmann equation with hard sphere collisions and the Maxwell reflection boundary condition. Our construction of solutions include the interior fluid part and Knudsen-viscous coupled boundary layers. The main novelty is that the accommodation coefficient is in the full range $0<α\leq 1$. The previous works in the context of classical solutions only considered the simplest specular reflection boundary condition, i.e. $α=0$. The mechanism of the derivation of fluid boundary conditions in the case $α=O(1)$ is quite different with the cases $α=0$ or $α=o(1)$. This rigorously justifies the corresponding formal analysis in Sone's books \cite{sone2002kinetic,sone2007molecular}. In particular, this is a smooth solution analogue of \cite{jiang2010remarks}, in which the renormalized solution was considered and the boundary layers were not visible.

math.AP

Knudsen boundary layer equations for full ranges of cutoff collision kernels: Maxwell reflection boundary with all accommodation coefficients in [0,1]

In this paper, we prove the existence and uniqueness of the Knudsen layer equation imposed on Maxwell reflection boundary condition with full ranges of cutoff collision kernels and accommodation coefficients (i.e., $- 3 < γ\leq 1$ and $0 \leq α_* \leq 1$, respectively) in the $L^\infty_{x,v}$ framework. Moreover, the solution enjoys the exponential decay $\exp \{- c x^\frac{2}{3 - γ} - c |v|^2 \}$ for some $c > 0$. In order to study the general angular cutoff collision kernel $-3 < γ\leq 1$, we should introduce a $(x,v)$-mixed weight $σ$. The biggest difficulty in this paper is the nondissipative boundary condition, hence, the boundary temperature and velocity $(T_w, u_w)$ on $\{ x = 0 \}$ and $(T, \mathfrak{u})$ on $\{ x = + \infty \}$ do not guarantee the nonnegativity of the $L^2$ boundary energy. We also do not assume that $(T_w, u_w)$ and $(T, \mathfrak{u})$ are very closed to each other. We first derive the Nondissipative boundary lemma to pull the boundary energy to the interior weighted $L^2$ norms with higher power of $x$-polynomial weights. Then a so-called spatial-velocity indices iteration approach is developed to shift the higher power $x$-polynomial weights to $|v|$-polynomial weights. Finally, we construct an interleaved iteration process such that the boundary energy is successfully dominated.

math.AP

Knudsen boundary layer equations with incoming boundary condition: full range of cutoff collision kernels and Mach numbers of the far field

This paper establishes tahe existence and uniqueness of the nonlinear Knudsen layer equation with incoming boundary conditions. It is well-known that the solvability conditions of the problem vary with the Mach number of the far Maxwellian $\mathcal{M}^\infty$. We consider full ranges of cutoff collision kernels (i.e., $- 3 < γ\leq 1$) and all the Mach numbers of the far field in the $L^\infty_{x,v}$ framework. Additionally, the solution exhibits exponential decay $\exp \{- c x^\frac{2}{3 - γ} - c |v|^2 \}$ for some $c > 0$. To address the general angular cutoff collision kernel, we introduce a $(x,v)$-mixed weight $σ$. The proof is essentially bsed on adding an artificial damping term.

math.AP

Functional Group Induced Transformations in Stacking and Electron Structure in Mo2CTx/NiS Heterostructures

The two-dimensional transition metal carbide/nitride family (MXenes) has garnered significant attention due to their highly customizable surface functional groups. Leveraging modern material science techniques, the customizability of MXenes can be enhanced further through the construction of associated heterostructures. As indicated by recent research, the Mo2CTx/NiS heterostructure has emerged as a promising candidate exhibiting superior physical and chemical application potential. The geometrical structure of Mo2CTx/NiS heterostructure is modeled and 6 possible configurations are validated by Density Functional Theory simulations. The variation in functional groups leads to structural changes in Mo2CTx/NiS interfaces, primarily attributed to the competition between van der Waals and covalent interactions. The presence of different functional groups results in significant band fluctuations near the Fermi level for Ni and Mo atoms, influencing the role of atoms and electron's ability to escape near the interface. This, in turn, modulates the strength of covalent interactions at the MXenes/NiS interface and alters the ease of dissociation of the MXenes/NiS complex. Notably, the Mo2CO2/NiS(P6_3/mmc) heterostructure exhibits polymorphism, signifying that two atomic arrangements can stabilize the structure. The transition process between these polymorphs is also simulated, further indicating the modulation of the electronic level of properties by a sliding operation.

cond-mat.mes-hall

A Possible Reason for why Data-Driven Beats Theory-Driven Computer Vision

Why do some continue to wonder about the success and dominance of deep learning methods in computer vision and AI? Is it not enough that these methods provide practical solutions to many problems? Well no, it is not enough, at least for those who feel there should be a science that underpins all of this and that we should have a clear understanding of how this success was achieved. Here, this paper proposes that the dominance we are witnessing would not have been possible by the methods of deep learning alone: the tacit change has been the evolution of empirical practice in computer vision and AI over the past decades. We demonstrate this by examining the distribution of sensor settings in vision datasets and performance of both classic and deep learning algorithms under various camera settings. This reveals a strong mismatch between optimal performance ranges of classical theory-driven algorithms and sensor setting distributions in the common vision datasets, while data-driven models were trained for those datasets. The head-to-head comparisons between data-driven and theory-driven models were therefore unknowingly biased against the theory-driven models.

cs.CV

A Unifying Hybrid Consensus Protocol

We introduce Unity, a new consensus algorithm for public blockchain settings. Unity is an eventual consistency protocol merging the Proof-of-Work (PoW) and Proof-of-Stake (PoS) into a coherent stochastic process. It encompasses hardware and economic security without sacrificing availability, unpredictability and decentralization. Empirical results indicate that the proposed protocol is fair and scalable to an arbitrary number of miners and stakers.

cs.CR

Active Control of Camera Parameters for Object Detection Algorithms

Camera parameters not only play an important role in determining the visual quality of perceived images, but also affect the performance of vision algorithms, for a vision-guided robot. By quantitatively evaluating four object detection algorithms, with respect to varying ambient illumination, shutter speed and voltage gain, it is observed that the performance of the algorithms is highly dependent on these variables. From this observation, a novel active control of camera parameters method is proposed, to make robot vision more robust under different light conditions. Experimental results demonstrate the effectiveness of our proposed approach, which improves the performance of object detection algorithms, compared with the conventional auto-exposure algorithm.

cs.CV