SearcharxivSearch

arXiv subjects

Yue Sheng

Publications and source records attributed to Yue Sheng.

8 recordsLinked to original sources

Context and Pixel Aware Large Language Model for Video Quality Assessment

Video quality assessment (VQA) is a challenging research topic with broad applications. Traditional hand-crafted and discriminative learning-based VQA models mainly focus on pixel-level distortions and lack contextual understanding, while recent multimodal large language models (MLLMs) struggle with sensitivity to small distortions or handle quality scoring and description as separate tasks. To address these shortcomings, we introduce CP-LLM: a Context- and Pixel-aware Large Language Model. CP-LLM is a novel multimodal LLM architecture featuring dual vision encoders designed to independently analyze perceptual quality at both high-level (video context) and low-level (pixel distortion) granularity, along with a language decoder that subsequently reasons about the interplay between these aspects. This design enables CP-LLM to simultaneously produce robust quality scores and interpretable quality descriptions, with enhanced sensitivity to pixel distortions (e.g., compression artifacts). Experiment results demonstrate that CP-LLM achieves state-of-the-art cross-dataset performance on VQA benchmarks and superior robustness to pixel distortions.

cs.CV

Code-Based English Models Surprising Performance on Chinese QA Pair Extraction Task

In previous studies, code-based models have consistently outperformed text-based models in reasoning-intensive scenarios. When generating our knowledge base for Retrieval-Augmented Generation (RAG), we observed that code-based models also perform exceptionally well in Chinese QA Pair Extraction task. Further, our experiments and the metrics we designed discovered that code-based models containing a certain amount of Chinese data achieve even better performance. Additionally, the capabilities of code-based English models in specified Chinese tasks offer a distinct perspective for discussion on the philosophical "Chinese Room" thought experiment.

cs.CL

Distributed linear regression by averaging

Distributed statistical learning problems arise commonly when dealing with large datasets. In this setup, datasets are partitioned over machines, which compute locally, and communicate short messages. Communication is often the bottleneck. In this paper, we study one-step and iterative weighted parameter averaging in statistical linear models under data parallelism. We do linear regression on each machine, send the results to a central server, and take a weighted average of the parameters. Optionally, we iterate, sending back the weighted average and doing local ridge regressions centered at it. How does this work compared to doing linear regression on the full data? Here we study the performance loss in estimation, test error, and confidence interval length in high dimensions, where the number of parameters is comparable to the training data size. We find the performance loss in one-step weighted averaging, and also give results for iterative averaging. We also find that different problems are affected differently by the distributed framework. Estimation error and confidence interval length increase a lot, while prediction error increases much less. We rely on recent results from random matrix theory, where we develop a new calculus of deterministic equivalents as a tool of broader interest.

math.ST

Accelerated Gradient Flow: Risk, Stability, and Implicit Regularization

Acceleration and momentum are the de facto standard in modern applications of machine learning and optimization, yet the bulk of the work on implicit regularization focuses instead on unaccelerated methods. In this paper, we study the statistical risk of the iterates generated by Nesterov's accelerated gradient method and Polyak's heavy ball method, when applied to least squares regression, drawing several connections to explicit penalization. We carry out our analyses in continuous-time, allowing us to make sharper statements than in prior work, and revealing complex interactions between early stopping, stability, and the curvature of the loss function.

stat.ML

Selecting the number of components in PCA via random signflips

Principal component analysis (PCA) is a foundational tool in modern data analysis, and a crucial step in PCA is selecting the number of components to keep. However, classical selection methods (e.g., scree plots, parallel analysis, etc.) lack statistical guarantees in the increasingly common setting of large-dimensional data with heterogeneous noise, i.e., where each entry may have a different noise variance. Moreover, it turns out that these methods, which are highly effective for homogeneous noise, can fail dramatically for data with heterogeneous noise. This paper proposes a new method called signflip parallel analysis (FlipPA) for the setting of approximately symmetric noise: it compares the data singular values to those of "empirical null" matrices generated by flipping the sign of each entry randomly with probability one-half. We develop a rigorous theory for FlipPA, showing that it has nonasymptotic type I error control and that it consistently selects the correct rank for signals rising above the noise floor in the large-dimensional limit (even when the noise is heterogeneous). We also rigorously explain why classical permutation-based parallel analysis degrades under heterogeneous noise. Finally, we illustrate that FlipPA compares favorably to state-of-the-art methods via numerical simulations and an illustration on data coming from astronomy.

math.ST

WONDER: Weighted one-shot distributed ridge regression in high dimensions

In many areas, practitioners need to analyze large datasets that challenge conventional single-machine computing. To scale up data analysis, distributed and parallel computing approaches are increasingly needed. Here we study a fundamental and highly important problem in this area: How to do ridge regression in a distributed computing environment? Ridge regression is an extremely popular method for supervised learning, and has several optimality properties, thus it is important to study. We study one-shot methods that construct weighted combinations of ridge regression estimators computed on each machine. By analyzing the mean squared error in a high dimensional random-effects model where each predictor has a small effect, we discover several new phenomena. 1. Infinite-worker limit: The distributed estimator works well for very large numbers of machines, a phenomenon we call "infinite-worker limit". 2. Optimal weights: The optimal weights for combining local estimators sum to more than unity, due to the downward bias of ridge. Thus, all averaging methods are suboptimal. We also propose a new Weighted ONe-shot DistributEd Ridge regression (WONDER) algorithm. We test WONDER in simulation studies and using the Million Song Dataset as an example. There it can save at least 100x in computation time, while nearly preserving test accuracy.

math.ST

Rational Solutions of the Painlevé-III Equation

All of the six Painlevé equations except the first have families of rational solutions, which are frequently important in applications. The third Painlevé equation in generic form depends on two parameters $m$ and $n$, and it has rational solutions if and only if at least one of the parameters is an integer. We use known algebraic representations of the solutions to study numerically how the distributions of poles and zeros behave as $n\in\mathbb{Z}$ increases and how the patterns vary with $m\in\mathbb{C}$. This study suggests that it is reasonable to consider the rational solutions in the limit of large $n\in\mathbb{Z}$ with $m\in\mathbb{C}$ being an auxiliary parameter. To analyze the rational solutions in this limit, algebraic techniques need to be supplemented by analytical ones, and the main new contribution of this paper is to develop a Riemann-Hilbert representation of the rational solutions of Painlevé-III that is amenable to asymptotic analysis. Assuming further that $m$ is a half-integer, we derive from the Riemann-Hilbert representation a finite dimensional Hankel system for the rational solution in which $n\in\mathbb{Z}$ appears as an explicit parameter.

math.CA

Rational Solutions of the Painlevé-II Equation Revisited

The rational solutions of the Painlevé-II equation appear in several applications and are known to have many remarkable algebraic and analytic properties. They also have several different representations, useful in different ways for establishing these properties. In particular, Riemann-Hilbert representations have proven to be useful for extracting the asymptotic behavior of the rational solutions in the limit of large degree (equivalently the large-parameter limit). We review the elementary properties of the rational Painlevé-II functions, and then we describe three different Riemann-Hilbert representations of them that have appeared in the literature: a representation by means of the isomonodromy theory of the Flaschka-Newell Lax pair, a second representation by means of the isomonodromy theory of the Jimbo-Miwa Lax pair, and a third representation found by Bertola and Bothner related to pseudo-orthogonal polynomials. We prove that the Flaschka-Newell and Bertola-Bothner Riemann-Hilbert representations of the rational Painlevé-II functions are explicitly connected to each other. Finally, we review recent results describing the asymptotic behavior of the rational Painlevé-II functions obtained from these Riemann-Hilbert representations by means of the steepest descent method.

nlin.SI