SearcharxivSearch

arXiv subjects

Jeffrey Cheng

Publications and source records attributed to Jeffrey Cheng.

14 recordsLinked to original sources

The unique limit of the Glimm-Lax construction for Sobolev data and obstructions to 1-d convex integration

We consider a genuinely nonlinear $1$-d system of hyperbolic conservation laws with two unknowns. A famous construction of Glimm & Lax shows that global-in-time "Glimm-Lax" weak entropy solutions exist in this setting for any initial data with small $L^\infty$ norm [Mem. Amer. Math. Soc. (1970), no. 101]. Recent work in the $L^1$-stability theory by Bressan, Marconi & Vaidya has given the first partial uniqueness and stability results for these solutions [Arch. Ration. Mech. Anal. (2025), vol. 249]. In this paper, we build on these results by combining them with recent advances in the $L^2$-theory. We show that solutions with initial data in the Sobolev space $H^s$ for $s>0$ are unique in the full class of Glimm--Lax solutions that decay in total variation at a rate of $1/t$. As a secondary result, our techniques are also used to show the recent non-uniqueness result of Chen, Vasseur & Yu for continuous solutions (arxiv:2407.02927) cannot extend to $C^\alpha$ solutions for $\alpha > 1/2$, alongside some appropriate fractional Sobolev spaces $W^{s,p}$. An auxiliary result of independent interest is the development of a weighted relative entropy contraction for perturbations of rarefaction waves.

math.AP

Relative Entropy Contractions for Extremal Shocks of Nonlinear Hyperbolic Systems without Genuine Nonlinearity

We study extremal shocks of $1$-d hyperbolic systems of conservation laws which fail to be genuinely nonlinear. More specifically, we consider either $1$- or $n$-shocks in characteristic fields which are either concave-convex or convex-concave in the sense of LeFloch. We show that the theory of $a$-contraction can be applied to obtain $L^2$-stability up to shift for these shocks in a class of weak solutions to the conservation law whose shocks obey the Lax entropy condition. Our results apply in particular to the $2 \times 2$ system of nonlinear elastodynamics.

math.AP

CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers?

A core part of scientific peer review involves providing expert critiques that directly assess the scientific claims a paper makes. While it is now possible to automatically generate plausible (if generic) reviews, ensuring that these reviews are sound and grounded in the papers' claims remains challenging. To facilitate LLM benchmarking on these challenges, we introduce CLAIMCHECK, an annotated dataset of NeurIPS 2023 and 2024 submissions and reviews mined from OpenReview. CLAIMCHECK is richly annotated by ML experts for weakness statements in the reviews and the paper claims that they dispute, as well as fine-grained labels of the validity, objectivity, and type of the identified weaknesses. We benchmark several LLMs on three claim-centric tasks supported by CLAIMCHECK, requiring models to (1) associate weaknesses with the claims they dispute, (2) predict fine-grained labels for weaknesses and rewrite the weaknesses to enhance their specificity, and (3) verify a paper's claims with grounded reasoning. Our experiments reveal that cutting-edge LLMs, while capable of predicting weakness labels in (2), continue to underperform relative to human experts on all other tasks.

cs.CL

Is That Your Final Answer? Test-Time Scaling Improves Selective Question Answering

Scaling the test-time compute of large language models has demonstrated impressive performance on reasoning benchmarks. However, existing evaluations of test-time scaling make the strong assumption that a reasoning system should always give an answer to any question provided. This overlooks concerns about whether a model is confident in its answer, and whether it is appropriate to always provide a response. To address these concerns, we extract confidence scores during reasoning for thresholding model responses. We find that increasing compute budget at inference time not only helps models answer more questions correctly, but also increases confidence in correct responses. We then extend the current paradigm of zero-risk responses during evaluation by considering settings with non-zero levels of response risk, and suggest a recipe for reporting evaluations under these settings.

cs.CL

Uniqueness & Weak-BV Stability in the Large for Isothermal Gas Dynamics

For the $1$-d isothermal Euler system, we consider the family of entropic BV solutions with possibly large, but finite, total variation. We show that these solutions are stable with respect to large perturbations in a class of weak solutions to the system which may not even be BV. The method is based on the construction of a modified front tracking algorithm, in which the theory of $a$-contraction with shifts for shocks is used as a building block. The main contribution is to construct the weight in the modified front tracking algorithm in a large-BV setting.

math.AP

Viscous Destabilization for Large Shocks of Conservation Laws

The recent theory of $a-$contraction with shifts provides $L^2$-stability for shock waves of $1-$D hyperbolic systems of conservation laws. The theory has been established at the inviscid level uniformly in the shock amplitude, and at the viscous level for small shocks. In this work, we investigate whether the $a-$contraction property holds uniformly in the shock amplitude for some specific systems with viscosity. We show that in some cases, the $a-$contraction fails for sufficiently large shocks. This showcases a "viscous destabilization" effect in the sense that the $a$-contraction property is verified for the inviscid model, but can fail for the viscous one. This also shows that the $a$-contraction property, even among small perturbations, is stronger than the classical notion of nonlinear stability, which is known to hold regardless of shock amplitude for viscous scalar conservation laws.

math.AP

Compressed Chain of Thought: Efficient Reasoning Through Dense Representations

Chain-of-thought (CoT) decoding enables language models to improve reasoning performance at the cost of high generation latency in decoding. Recent proposals have explored variants of contemplation tokens, a term we introduce that refers to special tokens used during inference to allow for extra computation. Prior work has considered fixed-length sequences drawn from a discrete set of embeddings as contemplation tokens. Here we propose Compressed Chain-of-Thought (CCoT), a framework to generate contentful and continuous contemplation tokens of variable sequence length. The generated contemplation tokens are compressed representations of explicit reasoning chains, and our method can be applied to off-the-shelf decoder language models. Through experiments, we illustrate how CCoT enables additional reasoning over dense contentful representations to achieve corresponding improvements in accuracy. Moreover, the reasoning improvements can be adaptively modified on demand by controlling the number of contemplation tokens generated.

cs.CL

$L^2$-stability $\&$ Minimal Entropy Conditions for Scalar Conservation Laws with Concave-Convex Fluxes

In this paper, we study stability properties of solutions to scalar conservation laws with a class of non-convex fluxes. Using the theory of $a$-contraction with shifts, we show $L^2$-stability for shocks among a class of large perturbations, and give estimates on the weight coefficient $a$ in regimes where the shock amplitude is both large and small. Then, we use these estimates as a building block to show a uniqueness theorem under minimal entropy conditions for weak solutions to the conservation law via a modified front tracking algorithm. The proof is inspired by an analogous program carried out in the $2 \times 2$ system setting by Chen, Golding, Krupa, $\&$ Vasseur.

math.AP

Dated Data: Tracing Knowledge Cutoffs in Large Language Models

Released Large Language Models (LLMs) are often paired with a claimed knowledge cutoff date, or the dates at which training data was gathered. Such information is crucial for applications where the LLM must provide up to date information. However, this statement only scratches the surface: do all resources in the training data share the same knowledge cutoff date? Does the model's demonstrated knowledge for these subsets closely align to their cutoff dates? In this work, we define the notion of an effective cutoff. This is distinct from the LLM designer reported cutoff and applies separately to sub-resources and topics. We propose a simple approach to estimate effective cutoffs on the resource-level temporal alignment of an LLM by probing across versions of the data. Using this analysis, we find that effective cutoffs often differ from reported cutoffs. To understand the root cause of this observation, we conduct a direct large-scale analysis on open pre-training datasets. Our analysis reveals two reasons for these inconsistencies: (1) temporal biases of CommonCrawl data due to non-trivial amounts of old data in new dumps and (2) complications in LLM deduplication schemes involving semantic duplicates and lexical near-duplicates. Overall, our results show that knowledge cutoffs are not as simple as they have seemed and that care must be taken both by LLM dataset curators as well as practitioners who seek to use information from these models.

cs.CL

Isometric embedding and spectral constraints for weighted graph metrics

A weighted graph $\phi G$ encodes a finite metric space $D_{\phi G}$. When is $D$ totally decomposable? When does it embed in $\ell_1$ space? When does its representing matrix have $\leq 1$ positive eigenvalue? We give useful lemmata and prove that these questions can be answered without examining $\phi$ if and only if $G$ has no $K_{2,3}$ minor. We also prove results toward the following conjecture. $D_{\phi G}$ has $\leq n$ positive eigenvalues for all $\phi$, if and only if $G$ has no $K_{2,3,...,3}$ minor, with $n$ threes.

math.CO

Graphical distances & inertia

We study the inertia of distance matrices of weighted graphs. Our novel congruence-based proof of the inertia of weighted trees extends to a proof for the inertia of weighted unicyclic graphs whose cycle is a triangle. Partial results are given on the inertia of other rationally weighted unicylic graphs.

math.CO

Augmenting Supervised Learning by Meta-learning Unsupervised Local Rules

The brain performs unsupervised learning and (perhaps) simultaneous supervised learning. This raises the question as to whether a hybrid of supervised and unsupervised methods will produce better learning. Inspired by the rich space of Hebbian learning rules, we set out to directly learn the unsupervised learning rule on local information that best augments a supervised signal. We present the Hebbian-augmented training algorithm (HAT) for combining gradient-based learning with an unsupervised rule on pre-synpatic activity, post-synaptic activities, and current weights. We test HAT's effect on a simple problem (Fashion-MNIST) and find consistently higher performance than supervised learning alone. This finding provides empirical evidence that unsupervised learning on synaptic activities provides a strong signal that can be used to augment gradient-based methods. We further find that the meta-learned update rule is a time-varying function; thus, it is difficult to pinpoint an interpretable Hebbian update rule that aids in training. We do find that the meta-learner eventually degenerates into a non-Hebbian rule that preserves important weights so as not to disturb the learner's convergence.

cs.LG

Bilingual is At Least Monolingual (BALM): A Novel Translation Algorithm that Encodes Monolingual Priors

State-of-the-art machine translation (MT) models do not use knowledge of any single language's structure; this is the equivalent of asking someone to translate from English to German while knowing neither language. BALM is a framework incorporates monolingual priors into an MT pipeline; by casting input and output languages into embedded space using BERT, we can solve machine translation with much simpler models. We find that English-to-German translation on the Multi30k dataset can be solved with a simple feedforward network under the BALM framework with near-SOTA BLEU scores.

cs.CL

AI Reasoning Systems: PAC and Applied Methods

Learning and logic are distinct and remarkable approaches to prediction. Machine learning has experienced a surge in popularity because it is robust to noise and achieves high performance; however, ML experiences many issues with knowledge transfer and extrapolation. In contrast, logic is easily intepreted, and logical rules are easy to chain and transfer between systems; however, inductive logic is brittle to noise. We then explore the premise of combining learning with inductive logic into AI Reasoning Systems. Specifically, we summarize findings from PAC learning (conceptual graphs, robust logics, knowledge infusion) and deep learning (DSRL, $\partial$ILP, DeepLogic) by reproducing proofs of tractability, presenting algorithms in pseudocode, highlighting results, and synthesizing between fields. We conclude with suggestions for integrated models by combining the modules listed above and with a list of unsolved (likely intractable) problems.

cs.AI