Searcharxiv⌕ Search

arXiv subjects

Liang Cheng

Publications and source records attributed to Liang Cheng.

At least 37 records · Page 2Linked to original sources

S2LPP: Small-to-Large Prompt Prediction across LLMs

The performance of pre-trained Large Language Models (LLMs) is often sensitive to nuances in prompt templates, requiring careful prompt engineering, adding costs in terms of computing and human effort. In this study, we present experiments encompassing multiple LLMs variants of varying sizes aimed at probing their preference with different prompts. Through experiments on Question Answering, we show prompt preference consistency across LLMs of different sizes. We also show that this consistency extends to other tasks, such as Natural Language Inference. Utilizing this consistency, we propose a method to use a smaller model to select effective prompt templates for a larger model. We show that our method substantially reduces the cost of prompt engineering while consistently matching performance with optimal prompts among candidates. More importantly, our experiment shows the efficacy of our strategy across fourteen LLMs and its applicability to a broad range of NLP tasks, highlighting its robustness

cs.CL↗

The power series expansions of logarithmic Sobolev, $\mathcal{W}$- functionals and scalar curvature rigidity

In this paper, we obtain that the logarithmic Sobolev and $\mathcal{W}$-functionals admit remarkable power series expansions when appropriate test functions are selected. Using these expansions formulas, we prove that for an open subset $V$ in an $n$-dimensional manifold $M$ with $\bar{V}\subset M$ satisfying: (a)The scalar curvature of $V$ satisfies the lower bound:$$\operatorname{Sc}(x) \geq n(n-1)K \quad \text{for all } x \in V,$$ (b) The isoperimetric profile of $V$ is no less than that of space form $M^n_K$:$$ \operatorname{I}(V,β) := \inf_{\substack{Ω\subset V \\ \mathrm{Vol}(Ω)=β}} \mathrm{Area}(\partial Ω) \geq \operatorname{I}(M^n_K,β) \quad \text{for some } β_0>0 \text{ and all } 0<β<β_0,$$\textbf{then} the sectional curvature of $V$ must satisfy $$\operatorname{Sec}(x) = K \quad \text{for all } x \in V.$$ Additionally, we derive some new scalar curvature rigidity theorems concerninglogarithmic Sobolev inequality and Perelman's $\boldsymbolμ$-functional.

math.DG↗

Advancing Antiferromagnetic Nitrides via Metal Alloy Nitridation

Nitride materials, valued for their structural stability and exceptional physical properties, have garnered significant interest in both fundamental research and technological applications. The fabrication of high-quality nitride thin films is essential for advancing their use in microelectronics and spintronics. Yet, achieving single-crystal nitride thin films with excellent structural integrity remains a challenge. Here, we introduce a straightforward yet innovative metallic alloy nitridation technique for the synthesis of stable single-crystal nitride thin films. By subjecting metal alloy thin films to a controlled nitridation process, nitrogen atoms integrate into the lattice, driving structural transformations while preserving high epitaxial quality. Combining nanoscale magnetic imaging with a diamond nitrogen-vacancy (NV) probe, X-ray magnetic linear dichroism, and comprehensive transport measurements, we confirm that the nitridated films exhibit a robust antiferromagnetic character with a zero net magnetic moment. This work not only provides a refined and reproducible strategy for the fabrication of nitride thin films but also lays a robust foundation for exploring their burgeoning device applications.

cond-mat.mtrl-sci↗

Neutralizing Bias in LLM Reasoning using Entailment Graphs

LLMs are often claimed to be capable of Natural Language Inference (NLI), which is widely regarded as a cornerstone of more complex forms of reasoning. However, recent works show that LLMs still suffer from hallucinations in NLI due to attestation bias, where LLMs overly rely on propositional memory to build shortcuts. To solve the issue, we design an unsupervised framework to construct counterfactual reasoning data and fine-tune LLMs to reduce attestation bias. To measure bias reduction, we build bias-adversarial variants of NLI datasets with randomly replaced predicates in premises while keeping hypotheses unchanged. Extensive evaluations show that our framework can significantly reduce hallucinations from attestation bias. Then, we further evaluate LLMs fine-tuned with our framework on original NLI datasets and their bias-neutralized versions, where original entities are replaced with randomly sampled ones. Extensive results show that our framework consistently improves inferential performance on both original and bias-neutralized NLI datasets.

cs.CL↗

An $ε$-regularity theorem for Perelman's reduced volume

In this article, we prove an $ε$-regularity theorem for Perelman's reduced volume. We show that on a Ricci flow, if Perelman's reduced volume is close to $1$, then the curvature radius at the base point cannot be too small.

math.DG↗

On locally conformally flat manifolds with positive pinched Ricci curvature

By using the Yamabe flow, we prove that if $(M^n,g)$, $n\geq3$, is an $n$-dimensional locally conformally flat complete Riemannian manifold $Rc\geq εRg>0$, where $ε>0$ is a uniformly constant, then $M^n$ must be compact. Our result shows that Hamilton's pinching conjecture also holds for higher dimensional case if we assume additionally the metric is locally conformally flat.

math.DG↗

Pathologist-like explainable AI for interpretable Gleason grading in prostate cancer

The aggressiveness of prostate cancer, the most common cancer in men worldwide, is primarily assessed based on histopathological data using the Gleason scoring system. While artificial intelligence (AI) has shown promise in accurately predicting Gleason scores, these predictions often lack inherent explainability, potentially leading to distrust in human-machine interactions. To address this issue, we introduce a novel dataset of 1,015 tissue microarray core images, annotated by an international group of 54 pathologists. The annotations provide detailed localized pattern descriptions for Gleason grading in line with international guidelines. Utilizing this dataset, we develop an inherently explainable AI system based on a U-Net architecture that provides predictions leveraging pathologists' terminology. This approach circumvents post-hoc explainability methods while maintaining or exceeding the performance of methods trained directly for Gleason pattern segmentation (Dice score: 0.713 $\pm$ 0.003 trained on explanations vs. 0.691 $\pm$ 0.010 trained on Gleason patterns). By employing soft labels during training, we capture the intrinsic uncertainty in the data, yielding strong results in Gleason pattern segmentation even in the context of high interobserver variability. With the release of this dataset, we aim to encourage further research into segmentation in medical tasks with high levels of subjectivity and to advance the understanding of pathologists' reasoning processes.

eess.IV↗

Work-in-Progress: Traded Control Transfer for Managing Real-Time Sensor Uncertainties in Autonomous Vehicle

At Levels 2 and 3 of autonomous driving defined by the Society of Auto-motive Engineers, drivers must take on certain driving responsibilities, and automated driving must sometimes yield to human control. This situation can occur in real time due to uncertainties in sensor measurements caused by environmental factors like fog or smoke. To address this challenge, we propose a method to manage real-time sensor uncertainties in autonomous vehicles by monitoring sensor conflicts and dynamically adjusting control authority to maintain safe operation. However, to achieve this, we have introduced a novel metric called the Degree of Conflicts (DoC), which quantifies the conflict between real-time sensor data by measuring the differences between data from multiple sensors. Our approach aims to demonstrate the importance of selecting an appropriate DoC threshold for transferring control between the automation agent and the human driver. The results have shown that choosing the correct DoC threshold can enhance safety by promptly handing over the driving control from the automation system to the human driver in challenging conditions.

eess.SY↗

Hydrodynamics of an oscillating cylinder inline with steady current

Wake and force characteristics of an oscillating cylinder in inline steady currents are investigated numerically over a wide parameter space of dimensionless oscillation amplitude ($A^* = 0.01 - 0.50$) and wavelength ($λ^* = 0.4 - 25$) at a fixed Reynolds number $Re = 500$. Fundamental issues addressed in this study are the interactions of wakes induced by steady approaching flow and cylinder oscillations and the influences of the governing parameters of $A^$ and $λ^$ on such interactions. Whilst the collinear flow is dominated by wakes induced by cylinder oscillation at $λ^* \leq 1.5$ and steady current at $λ^* \geq 10$, it exhibits characteristics of nonlinear interactions of wakes induced by the cylinder oscillation and steady current at $λ^* = 1.5 - 10$, such as the formation of multiple synchronized modes interleaved with desynchronized modes. The synchronized mode varies with both $λ^$ and $A^$, forming an inclined Arnold's tongue across $λ^-A^$ space. There is a wide variability of the vortex shedding pattern in each synchronized mode. Variations of different hydrodynamic force coefficients with $λ^$ and $A^$ are investigated with physical interpretations based on the wake characteristics. The applicability of the Morison equation in predicting inline force fluctuations is examined. We found that the Morison equation shows reasonable accuracy only for a small range of $λ^* \leq 1.5$. Beyond this range, its performance deteriorates due to the influence of steady current on wake characteristics.

physics.flu-dyn↗

Explicit Inductive Inference using Large Language Models

Large Language Models (LLMs) are reported to hold undesirable attestation bias on inference tasks: when asked to predict if a premise P entails a hypothesis H, instead of considering H's conditional truthfulness entailed by P, LLMs tend to use the out-of-context truth label of H as a fragile proxy. In this paper, we propose a pipeline that exploits this bias to do explicit inductive inference. Our pipeline uses an LLM to transform a premise into a set of attested alternatives, and then aggregate answers of the derived new entailment inquiries to support the original inference prediction. On a directional predicate entailment benchmark, we demonstrate that by applying this simple pipeline, we can improve the overall performance of LLMs on inference and substantially alleviate the impact of their attestation bias.

cs.CL↗

Variational autoencoder-based neural network model compression

Variational Autoencoders (VAEs), as a form of deep generative model, have been widely used in recent years, and shown great great peformance in a number of different domains, including image generation and anomaly detection, etc.. This paper aims to explore neural network model compression method based on VAE. The experiment uses different neural network models for MNIST recognition as compression targets, including Feedforward Neural Network (FNN), Convolutional Neural Network (CNN), Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM). These models are the most basic models in deep learning, and other more complex and advanced models are based on them or inherit their features and evolve. In the experiment, the first step is to train the models mentioned above, each trained model will have different accuracy and number of total parameters. And then the variants of parameters for each model are processed as training data in VAEs separately, and the trained VAEs are tested by the true model parameters. The experimental results show that using the latent space as a representation of the model compression can improve the compression rate compared to some traditional methods such as pruning and quantization, meanwhile the accuracy is not greatly affected using the model parameters reconstructed based on the latent space. In the future, a variety of different large-scale deep learning models will be used more widely, so exploring different ways to save time and space on saving or transferring models will become necessary, and the use of VAE in this paper can provide a basis for these further explorations.

cs.LG↗

Transfer learning-assisted inverse modeling in nanophotonics based on mixture density networks

The simulation of nanophotonic structures relies on electromagnetic solvers, which play a crucial role in understanding their behavior. However, these solvers often come with a significant computational cost, making their application in design tasks, such as optimization, impractical. To address this challenge, machine learning techniques have been explored for accurate and efficient modeling and design of photonic devices. Deep neural networks, in particular, have gained considerable attention in this field. They can be used to create both forward and inverse models. An inverse modeling approach avoids the need for coupling a forward model with an optimizer and directly performs the prediction of the optimal design parameters values. In this paper, we propose an inverse modeling method for nanophotonic structures, based on a mixture density network model enhanced by transfer learning. Mixture density networks can predict multiple possible solutions at a time including their respective importance as Gaussian distributions. However, multiple challenges exist for mixture density network models. An important challenge is that an upper bound on the number of possible simultaneous solutions needs to be specified in advance. Also, another challenge is that the model parameters must be jointly optimized, which can result computationally expensive. Moreover, optimizing all parameters simultaneously can be numerically unstable and can lead to degenerate predictions. The proposed approach allows overcoming these limitations using transfer learning-based techniques, while preserving a high accuracy in the prediction capability of the design solutions given an optical response as an input. A dimensionality reduction step is also explored. Numerical results validate the proposed method.

cs.LG↗

On local rigidity theorems with respect to the scalar curvature

By using the Ricci flow, we study local rigidity theorems regarding scalar curvature, isoperimetric constant and best constant of $L^2$ logarithmic Sobolev inequality. Precisely, we prove that if a metric $g$ on an open set $V$ in an $n$-dimensional Riemannian manifold satisfies $$ \int_V R(g) dvol_g \ge 0 \text{\ \ and\ \ } I(V)\ge I(\mathbb{R}^n), $$ or $$ \int_V R(g) dvol_g \ge 0 \text{\ \ and\ \ } S(V)\ge S(\mathbb{R}^n),$$ then $g=g_{\mathbb{R}^n}$ on $V$, where $R(g)$ is the scalar curvature of $g$, $\mathbb{R}^n$ is Euclidean space, $ I(V)$ is the isoperimetric constant of $V$ and $S(V)$ is best constant of $L^2$ logarithmic Sobolev inequality of $V$. Moreover,we also obtain the local $\mathbb{R}^n$-rigidity about local Perelman's $ν$-entropy, and local $\mathbb{S}^n$-rigidity (resp. $\mathbb{H}^n$-rigidity) theorems regarding the cases concerning $R(g)\ge n(n-1)$ (resp. $R(g)\ge -n(n-1) $), weighted isoperimetric constant and best constant of weighted $L^2$ logarithmic Sobolev inequality for the weighted metric $\left(\cos{\frac{d_g(p,x)}{2}}\right)^{-4}g$ (resp. $\left(\cosh{\frac{d_g(p,x)}{2}}\right)^{-4}g$).

math.DG↗

Uncovering the Handwritten Text in the Margins: End-to-end Handwritten Text Detection and Recognition

The pressing need for digitization of historical documents has led to a strong interest in designing computerised image processing methods for automatic handwritten text recognition. However, not much attention has been paid on studying the handwritten text written in the margins, i.e. marginalia, that also forms an important source of information. Nevertheless, training an accurate and robust recognition system for marginalia calls for data-efficient approaches due to the unavailability of sufficient amounts of annotated multi-writer texts. Therefore, this work presents an end-to-end framework for automatic detection and recognition of handwritten marginalia, and leverages data augmentation and transfer learning to overcome training data scarcity. The detection phase involves investigation of R-CNN and Faster R-CNN networks. The recognition phase includes an attention-based sequence-to-sequence model, with ResNet feature extraction, bidirectional LSTM-based sequence modeling, and attention-based prediction of marginalia. The effectiveness of the proposed framework has been empirically evaluated on the data from early book collections found in the Uppsala University Library in Sweden. Source code and pre-trained models are available at Github.

cs.CV↗

Sources of Hallucination by Large Language Models on Inference Tasks

Large Language Models (LLMs) are claimed to be capable of Natural Language Inference (NLI), necessary for applied tasks like question answering and summarization. We present a series of behavioral studies on several LLM families (LLaMA, GPT-3.5, and PaLM) which probe their behavior using controlled experiments. We establish two biases originating from pretraining which predict much of their behavior, and show that these are major sources of hallucination in generative LLMs. First, memorization at the level of sentences: we show that, regardless of the premise, models falsely label NLI test samples as entailing when the hypothesis is attested in training data, and that entities are used as ``indices'' to access the memorized data. Second, statistical patterns of usage learned at the level of corpora: we further show a similar effect when the premise predicate is less frequent than that of the hypothesis in the training data, a bias following from previous studies. We demonstrate that LLMs perform significantly worse on NLI test samples which do not conform to these biases than those which do, and we offer these as valuable controls for future LLM evaluation.

cs.CL↗

Pseudolocality theorems of Ricci flows on incomplete manifolds

In this paper we study the pseudolocality theorems of Ricci flows on incomplete manifolds. We prove that if a ball with its closure contained in an incomplete manifold has the small scalar curvature lower bound and almost Euclidean isoperimetric constant, or almost Euclidean local $\boldsymbolν$ constant, then we can construct a solution of Ricci flow in the ball which have the pseudolocality property. We also give two applications. First, we prove the short-time existence of Ricci flows on complete manifolds with scalar curvature bounded below uniformly and almost Euclidean isoperimetric inequality holds locally. Second, we show that any complete manifold with nonnegative scalar curvature and Euclidean isoperimetric inequality must be isometric to the Euclidean space.

math.DG↗

Magnetic-order-mediated carrier and phonon dynamics in MnBi2Te4

We investigate the quasiparticle dynamics in MnBi2Te4 single crystal using the ultrafast optical spectroscopy. Our results show that there exist anomalous dynamical optical responses below the antiferromagnetic (AFM) ordering temperature TN. In specific, we reveal that both the initial carrier decay and recombination processes can be modulated via introducing the AFM order in sub-picosecond and picosecond timescales, respectively. We also discover a long relaxation process emerging below TN with a timescale approaching to the nanosecond regime, and can be attributed to the T-dependent spin-lattice interaction. There also emerges an unusual phonon energy renormalization below TN , which is found to arise from its coupling the spin degree via the exchange interaction and magnetic anisotropy. Our findings provide key information for understanding the dynamical properties of non-equilibrium carrier, spin and lattice in MnBi2Te4.

cond-mat.mtrl-sci↗

Pseudolocality and uniqueness of Ricci flow on almost Euclidean noncompact manifolds

In this paper, we prove a pseudolocality-type theorem for $\mathcal L$-complete noncompact Ricci flow which may not have bounded sectional curvature; with the help of it we study the uniqueness of the Ricci flow on noncompact manifolds. In particular, we prove the strong uniqueness theorem for the $\mathcal L$-complete Ricci flow on the Euclidean space. This partially answers a question proposed by B-L.~Chen.

math.DG↗