SearcharxivSearch

arXiv subjects

Trung Vu

Publications and source records attributed to Trung Vu.

15 recordsLinked to original sources

Quantum Harish-Chandra bimodules at roots of unity and affine Hecke category

The category of Harish-Chandra bimodules for quantum groups was first appeared in the works about topological quantum field theory of surfaces. In this paper, we study this category when the quantum parameter q is an odd order root of unity. We relate the category to the category of affine Soergel bimodules and to non-commutative Springer resolution.

math.RT

On De Concini-Kac forms of quantum groups

Quantum groups of semisimple Lie algebras at roots of unity admit several different forms. Among them is the De Concini-Kac form, which is the easiest to define but, perhaps, hardest to study. In this paper, we propose a suitable modification to the De Concini-Kac form, namely the even part algebra, which has some appealing features. Notably, it behaves uniformly with respect to the order of the roots of unity and admits an adjoint action of the Lusztig form. We revisit several results due to De Concini-Kac-Procesi and Tanisaki for the even part algebra. Namely, we give conceptual definitions of the Frobenius and Harish-Chandra centers and describe the entire center in terms of these two subalgebras getting a complete quantum analog of the Veldkamp theorem on the center of the universal enveloping algebras in positive characteristic. We investigate the Azumaya locus of the even part algebra over its center. We also show that the locally finite part of the even part algebra under the adjoint action of the Lusztig form is isomorphic to the reflection equation algebra, which is the quantized coordinate algebra with the product twisted by $R$-matrix. Some results on Lusztig forms at roots of unity are revisited and proved in greater generality including Kempf vanishing theorem and good filtrations on the quantized coordinate algebra.

math.RT

On the functor relating Harish-Chandra bimodules and Soergel bimodules

In the 90's Soergel constructed a functor that relates Harish-Chandra bimodules to Soergel bimodules. We revisit this functor and relate it to the restriction functor constructed by Losev between Harish-Chandra bimodules and bimodules over $W$-algebra associated to the regular nilpotent element. We compute the images of certain Harish-Chandra bimodules under the restriction functor and provide alternative proofs for many properties of the functor constructed by Soergel. Blocks of Harish-Chandra bimodules with integral central characters were studied in the works of Soergel and Stroppel. We will generalize results of Soergel and Stroppel in the case of integral blocks to general blocks.

math.RT

OpenThoughts: Data Recipes for Reasoning Models

Reasoning models have made rapid progress on many benchmarks involving math, code, and science. Yet, there are still many open questions about the best training recipes for reasoning since state-of-the-art models often rely on proprietary datasets with little to no public information available. To address this, the goal of the OpenThoughts project is to create open-source datasets for training reasoning models. After initial explorations, our OpenThoughts2-1M dataset led to OpenThinker2-32B, the first model trained on public reasoning data to match DeepSeek-R1-Distill-32B on standard reasoning benchmarks such as AIME and LiveCodeBench. We then improve our dataset further by systematically investigating each step of our data generation pipeline with 1,000+ controlled experiments, which led to OpenThoughts3. Scaling the pipeline to 1.2M examples and using QwQ-32B as teacher yields our OpenThoughts3-7B model, which achieves state-of-the-art results: 53% on AIME 2025, 51% on LiveCodeBench 06/24-01/25, and 54% on GPQA Diamond - improvements of 15.3, 17.2, and 20.5 percentage points compared to the DeepSeek-R1-Distill-Qwen-7B. All of our datasets and models are available on https://openthoughts.ai.

cs.LG

Constrained Independent Vector Analysis with Reference for Multi-Subject fMRI Analysis

Independent component analysis (ICA) is now a widely used solution for the analysis of multi-subject functional magnetic resonance imaging (fMRI) data. Independent vector analysis (IVA) generalizes ICA to multiple datasets, i.e., to multi-subject data, and in addition to higher-order statistical information in ICA, it leverages the statistical dependence across the datasets as an additional type of statistical diversity. As such, it preserves variability in the estimation of single-subject maps but its performance might suffer when the number of datasets increases. Constrained IVA is an effective way to bypass computational issues and improve the quality of separation by incorporating available prior information. Existing constrained IVA approaches often rely on user-defined threshold values to define the constraints. However, an improperly selected threshold can have a negative impact on the final results. This paper proposes two novel methods for constrained IVA: one using an adaptive-reverse scheme to select variable thresholds for the constraints and a second one based on a threshold-free formulation by leveraging the unique structure of IVA. We demonstrate that our solutions provide an attractive solution to multi-subject fMRI analysis both by simulations and through analysis of resting state fMRI data collected from 98 subjects -- the highest number of subjects ever used by IVA algorithms. Our results show that both proposed approaches obtain significantly better separation quality and model match while providing computationally efficient and highly reproducible solutions.

eess.SP

Better Generalization with Semantic IDs: A Case Study in Ranking for Recommendations

Randomly-hashed item ids are used ubiquitously in recommendation models. However, the learned representations from random hashing prevents generalization across similar items, causing problems of learning unseen and long-tail items, especially when item corpus is large, power-law distributed, and evolving dynamically. In this paper, we propose using content-derived features as a replacement for random ids. We show that simply replacing ID features with content-based embeddings can cause a drop in quality due to reduced memorization capability. To strike a good balance of memorization and generalization, we propose to use Semantic IDs -- a compact discrete item representation learned from frozen content embeddings using RQ-VAE that captures the hierarchy of concepts in items -- as a replacement for random item ids. Similar to content embeddings, the compactness of Semantic IDs poses a problem of easy adaption in recommendation models. We propose novel methods for adapting Semantic IDs in industry-scale ranking models, through hashing sub-pieces of of the Semantic-ID sequences. In particular, we find that the SentencePiece model that is commonly used in LLM tokenization outperforms manually crafted pieces such as N-grams. To the end, we evaluate our approaches in a real-world ranking model for YouTube recommendations. Our experiments demonstrate that Semantic IDs can replace the direct use of video IDs by improving the generalization ability on new and long-tail item slices without sacrificing overall model quality.

cs.IR

Recommender Systems with Generative Retrieval

Modern recommender systems perform large-scale retrieval by first embedding queries and item candidates in the same unified space, followed by approximate nearest neighbor search to select top candidates given a query embedding. In this paper, we propose a novel generative retrieval approach, where the retrieval model autoregressively decodes the identifiers of the target candidates. To that end, we create semantically meaningful tuple of codewords to serve as a Semantic ID for each item. Given Semantic IDs for items in a user session, a Transformer-based sequence-to-sequence model is trained to predict the Semantic ID of the next item that the user will interact with. To the best of our knowledge, this is the first Semantic ID-based generative model for recommendation tasks. We show that recommender systems trained with the proposed paradigm significantly outperform the current SOTA models on various datasets. In addition, we show that incorporating Semantic IDs into the sequence-to-sequence model enhances its ability to generalize, as evidenced by the improved retrieval performance observed for items with no prior interaction history.

cs.IR

21st Century Global and Regional Surface Temperature Projections

Many regions across the globe broke their surface temperature records in recent years, further sparking concerns about the impending arrival of "tipping points" later in the 21st century. This study analyzes observed global surface temperature trends in three target latitudinal regions: the Arctic Circle, the Tropics, and the Antarctic Circle. We show that global warming is accelerating unevenly across the planet, with the Arctic warming at approximately three times the average rate of our world. We further analyzed the reliability of latitude-dependent surface temperature simulations from a suite of Coupled Model Intercomparison Project Phase 6 models and their multi-model mean. We found that GISS-E2-1-G and FGOALS-g3 were the best-performing models based on their statistical abilities to reproduce observational, latitude-dependent data. Surface temperatures were projected from ensemble simulations of the Shared Socioeconomic Pathway 2-4.5 (SSP2-4.5). We estimate when the climate will warm by 1.5, 2.0, and 2.5 degrees C relative to the preindustrial period, globally and regionally. GISS-E2-1-G projects that global surface temperature anomalies would reach 1.5, 2.0, and 2.5 degrees C in 2024 (+/-1.34), 2039 (+/-2.83), and 2057 (+/-5.03) respectively, while FGOALS-g3 predicts these "tipping points" would arrive in 2024 (+/-2.50), 2054 (+/-7.90), and 2087 (+/-10.55) respectively. Our results reaffirm a dramatic, upward trend in projected climate warming acceleration, with upward concavity in 21st century projections of the Arctic, which could lead to catastrophic consequences across the Earth. Further studies are necessary to determine the most efficient solutions to reduce global warming acceleration and maintain a low SSP, both globally and regionally.

physics.ao-ph

On Local Linear Convergence of Projected Gradient Descent for Unit-Modulus Least Squares

The unit-modulus least squares (UMLS) problem has a wide spectrum of applications in signal processing, e.g., phase-only beamforming, phase retrieval, radar code design, and sensor network localization. Scalable first-order methods such as projected gradient descent (PGD) have recently been studied as a simple yet efficient approach to solving the UMLS problem. Existing results on the convergence of PGD for UMLS often focus on global convergence to stationary points. As a non-convex problem, only a sublinear convergence rate has been established. However, these results do not explain the fast convergence of PGD frequently observed in practice. This manuscript presents a novel analysis of convergence of PGD for UMLS, justifying the linear convergence behavior of the algorithm near the solution. By exploiting the local structure of the objective function and the constraint set, we establish an exact expression for the convergence rate and characterize the conditions for linear convergence. Simulations show that our theoretical analysis corroborates numerical examples. Furthermore, variants of PGD with adaptive step sizes are proposed based on the new insight revealed in our convergence analysis. The variants show substantial acceleration in practice.

math.OC

A Closed-Form Bound on the Asymptotic Linear Convergence of Iterative Methods via Fixed Point Analysis

In many iterative optimization methods, fixed-point theory enables the analysis of the convergence rate via the contraction factor associated with the linear approximation of the fixed-point operator. While this factor characterizes the asymptotic linear rate of convergence, it does not explain the non-linear behavior of these algorithms in the non-asymptotic regime. In this letter, we take into account the effect of the first-order approximation error and present a closed-form bound on the convergence in terms of the number of iterations required for the distance between the iterate and the limit point to reach an arbitrarily small fraction of the initial distance. Our bound includes two terms: one corresponds to the number of iterations required for the linearized version of the fixed-point operator and the other corresponds to the overhead associated with the approximation error. With a focus on the convergence in the scalar case, the tightness of the proposed bound is proven for positively quadratic first-order difference equations.

eess.SY

On Asymptotic Linear Convergence of Projected Gradient Descent for Constrained Least Squares

Many recent problems in signal processing and machine learning such as compressed sensing, image restoration, matrix/tensor recovery, and non-negative matrix factorization can be cast as constrained optimization. Projected gradient descent is a simple yet efficient method for solving such constrained optimization problems. Local convergence analysis furthers our understanding of its asymptotic behavior near the solution, offering sharper bounds on the convergence rate compared to global convergence analysis. However, local guarantees often appear scattered in problem-specific areas of machine learning and signal processing. This manuscript presents a unified framework for the local convergence analysis of projected gradient descent in the context of constrained least squares. The proposed analysis offers insights into pivotal local convergence properties such as the conditions for linear convergence, the region of convergence, the exact asymptotic rate of convergence, and the bound on the number of iterations needed to reach a certain level of accuracy. To demonstrate the applicability of the proposed approach, we present a recipe for the convergence analysis of projected gradient descent and demonstrate it via a beginning-to-end application of the recipe on four fundamental problems, namely, linear equality-constrained least squares, sparse recovery, least squares with the unit norm constraint, and matrix completion.

math.OC

On Asymptotic Linear Convergence Rate of Iterative Hard Thresholding for Matrix Completion

Iterative hard thresholding (IHT) has gained in popularity over the past decades in large-scale optimization. However, convergence properties of this method have only been explored recently in non-convex settings. In matrix completion, existing works often focus on the guarantee of global convergence of IHT via standard assumptions such as incoherence property and uniform sampling. While such analysis provides a global upper bound on the linear convergence rate, it does not describe the actual performance of IHT in practice. In this paper, we provide a novel insight into the local convergence of a specific variant of IHT for matrix completion. We uncover the exact linear rate of IHT in a closed-form expression and identify the region of convergence in which the algorithm is guaranteed to converge. Furthermore, we utilize random matrix theory to study the linear rate of convergence of IHTSVD for large-scale matrix completion. We find that asymptotically, the rate can be expressed in closed form in terms of the relative rank and the sampling rate. Finally, we present various numerical results to verify the aforementioned theoretical analysis.

math.OC

Perturbation expansions and error bounds for the truncated singular value decomposition

Truncated singular value decomposition is a reduced version of the singular value decomposition in which only a few largest singular values are retained. This paper presents a novel perturbation analysis for the truncated singular value decomposition for real matrices. First, we describe perturbation expansions for the singular value truncation of order $r$. We extend perturbation results for the singular subspace decomposition to derive the first-order perturbation expansion of the truncated operator about a matrix with rank greater than or equal to $r$. Observing that the first-order expansion can be greatly simplified when the matrix has exact rank $r$, we further show that the singular value truncation admits a simple second-order perturbation expansion about a rank-$r$ matrix. Second, we introduce the first-known error bound on the linear approximation of the truncated singular value decomposition of a perturbed rank-$r$ matrix. Our bound only depends on the least singular value of the unperturbed matrix and the norm of the perturbation matrix. Intriguingly, while the singular subspaces are known to be extremely sensitive to additive noises, the newly established error bound holds universally for perturbations with arbitrary magnitude. Finally, we demonstrate an application of our results to the analysis of the mean squared error associated with the TSVD-based matrix denoising solution.

math.NA

Exact Linear Convergence Rate Analysis for Low-Rank Symmetric Matrix Completion via Gradient Descent

Factorization-based gradient descent is a scalable and efficient algorithm for solving low-rank matrix completion. Recent progress in structured non-convex optimization has offered global convergence guarantees for gradient descent under certain statistical assumptions on the low-rank matrix and the sampling set. However, while the theory suggests gradient descent enjoys fast linear convergence to a global solution of the problem, the universal nature of the bounding technique prevents it from obtaining an accurate estimate of the rate of convergence. In this paper, we perform a local analysis of the exact linear convergence rate of gradient descent for factorization-based matrix completion for symmetric matrices. Without any additional assumptions on the underlying model, we identify the deterministic condition for local convergence of gradient descent, which only depends on the solution matrix and the sampling set. More crucially, our analysis provides a closed-form expression of the asymptotic rate of convergence that matches exactly with the linear convergence observed in practice. To the best of our knowledge, our result is the first one that offers the exact rate of convergence of gradient descent for matrix factorization in Euclidean space for matrix completion.

math.OC

Parkinson's Disease Digital Biomarker Discovery with Optimized Transitions and Inferred Markov Emissions

We search for digital biomarkers from Parkinson's Disease by observing approximate repetitive patterns matching hypothesized step and stride periodic cycles. These observations were modeled as a cycle of hidden states with randomness allowing deviation from a canonical pattern of transitions and emissions, under the hypothesis that the averaged features of hidden states would serve to informatively characterize classes of patients/controls. We propose a Hidden Semi-Markov Model (HSMM), a latent-state model, emitting 3D-acceleration vectors. Transitions and emissions are inferred from data. We fit separate models per unique device and training label. Hidden Markov Models (HMM) force geometric distributions of the duration spent at each state before transition to a new state. Instead, our HSMM allows us to specify the distribution of state duration. This modified version is more effective because we are interested more in each state's duration than the sequence of distinct states, allowing inclusion of these durations the feature vector.

q-bio.QM