Searcharxiv⌕ Search

arXiv subjects

Cheng Li

Publications and source records attributed to Cheng Li.

At least 145 records · Page 8Linked to original sources

Gamma Analytical Modeling Evolution (GAME) I: The physical implications of deriving the stellar mass functions from z=0 to z=8

The $Γ$ growth model is an effective parameterization employed across various scientific disciplines and scales to depict growth. It has been demonstrated that the cosmic star formation rate density (CSFRD) can also be described broadly by this pattern, i.e. $\frac{dM(T)}{dT} = M_{z,0}\, \times \frac{β^α}{Γ(α)} \, T^{α-1} e^{-β\, T }$ M$_{\odot}$ Gyr$^{-1}$, where $M_{z,0}$ is the stellar mass at $z$ = 0, $α= 3.0$, $β= 0.5 $ Gyr$^{-1}$ and $T$ describes time. We use the identical $Γ$ growth pattern given by the CSFRD to extend the present day (z = 0) stellar mass bins $M_{\ast}(T)$ of the Galaxy Stellar Mass Function (GSMF) and investigate if we are able to reproduce observations for the high redshift GSMFs. Surprisingly, our scheme describes successfully the evolution of the GSMF over 13.5 Gyrs, especially for objects with intermediate and low masses. We observe some deviations that manifest {\it solely} at very high redshifts ($z > 1.5$, i.e. more than 9.5 Gyr ago) and {\it specifically} for very small and exceedingly massive objects. We discuss the possible solutions (e.g. impacts of mergers) for these offsets. Our formalism suggests that the evolution of the GSMF is set by simple (few parameters) and physically motivated arguments. The parameters $β$ and $α$ are theoretically consistent within a multi-scale context and are determined from the dynamical time scale ($β$) and the radial distribution of the accreting matter ($α$). We demonstrate that both our formalism and state-of-the-art simulations are consistent with recent GSMFs derived from JWST data at high redshifts.

astro-ph.GA↗

Pigeonhole Stochastic Gradient Langevin Dynamics for Large Crossed Mixed Effects Models

Large crossed mixed effects models with imbalanced structures and missing data pose major computational challenges for standard Bayesian posterior sampling algorithms, as the computational complexity is usually superlinear in the number of observations. We propose two efficient subset-based stochastic gradient MCMC algorithms for such crossed mixed effects models, which facilitate scalable inference on both the variance components and regression coefficients. The first algorithm is developed for balanced design without missing observations, where we leverage the closed-form expression of the precision matrix for the full data matrix. The second algorithm, which we call the pigeonhole stochastic gradient Langevin dynamics (PSGLD), is developed for both balanced and unbalanced designs with potentially a large proportion of missing observations. Our PSGLD algorithm imputes the latent crossed random effects by running short Markov chains and then samples the model parameters of variance components and regression coefficients at each MCMC iteration. We provide theoretical guarantees by showing the convergence of the output distribution from the proposed algorithms to the target non-log-concave posterior distribution. A variety of numerical experiments based on both synthetic and real data demonstrate that the proposed algorithms can significantly reduce the computational cost of the standard MCMC algorithms and better balance the approximation accuracy and computational efficiency.

stat.CO↗

Detection of two TeV gamma-ray outbursts from NGC 1275 by LHAASO

The Water Cherenkov Detector Array (WCDA) is one of the components of Large High Altitude Air Shower Observatory (LHAASO) and can monitor any sources over two-thirds of the sky for up to 7 hours per day with >98\% duty cycle. In this work, we report the detection of two outbursts of the Fanaroff-Riley I radio galaxy NGC 1275 that were detected by LHAASO-WCDA between November 2022 and January 2023 with statistical significance of 5.2~$σ$ and 8.3~$σ$. The observed spectral energy distribution in the range from 500 GeV to 3 TeV is fitted by a power-law with a best-fit spectral index of $α=-3.37\pm0.52$ and $-3.35\pm0.29$, respectively. The outburst flux above 0.5~TeV was ($4.55\pm 4.21)\times~10^{-11}~\rm cm^{-2}~s^{-1}$ and ($3.45\pm 1.78)\times~10^{-11}~\rm cm^{-2}~s^{-1}$, corresponding to 60\%, 45\% of Crab Nebula flux. Variation analysis reveals the variability time-scale of days at the TeV energy band. A simple test by one-zone synchrotron self-Compton model reproduces the data in the gamma-ray band well.

astro-ph.HE↗

MILLION: Mastering Long-Context LLM Inference Via Outlier-Immunized KV Product Quantization

Large language models (LLMs) are increasingly utilized for complex tasks requiring longer context lengths, with some models supporting up to 128K or 1M tokens. This trend, however, presents significant challenges in inference speed and memory management. Quantization emerges as a promising approach to address the widening gap between LLM size and memory capacity. However, traditional quantization schemes often yield suboptimal compression results for KV caches due to two key factors: i) On-the-fly quantization and de-quantization, causing significant performance overhead; ii) Prevalence of outliers in KV values, challenging low-bitwidth uniform quantization. To this end, we propose MILLION, a novel quantization framework achieving low-bitwidth KV cache through product quantization. First, we conduct a thorough analysis of KV cache distribution, revealing the limitations of existing quantization schemes. Second, we introduce a non-uniform quantization algorithm based on product quantization, which efficiently compresses data while preserving accuracy. Third, we develop a high-performance GPU inference framework with efficient attention kernel and pipeline design for MILLION that leverages sparse computation and asynchronous quantization, significantly enhancing inference speed. Comprehensive evaluation results demonstrate that MILLION can achieve 4 bits quantization with trivial perplexity and accuracy loss, and achieve 2.09x end-to-end performance gains at 32K context length. Code is released at https://github.com/ZongwuWang/MILLION.

cs.DC↗

KunlunBaize: LLM with Multi-Scale Convolution and Multi-Token Prediction Under TransformerX Framework

Large language models have demonstrated remarkable performance across various tasks, yet they face challenges such as low computational efficiency, gradient vanishing, and difficulties in capturing complex feature interactions. To address these limitations, a novel framework has been proposed. This framework incorporates a learnable dense residual skip connection mechanism, a TransformerX module a transformer based component integrating multiscale convolution and adaptive activation functions and a multitoken prediction interaction module. The learnable dense residual connections enhance information flow and feature capture across layers. Within the TransformerX module, large convolutional kernels aggregate semantic information from extensive text segments, while smaller convolutions focus on local word order and syntactic structures. The adaptive activation function dynamically adjusts its parameters based on the semantic features of the input text, improving the model's ability to handle diverse semantic expressions and complex relationships. The multitoken prediction module boosts data utilization and accelerates inference by predicting multiple future tokens. These components significantly enhance the performance and efficiency of large language models.

cs.CL↗

Video-VoT-R1: An efficient video inference model integrating image packing and AoE architecture

In the field of video-language pretraining, existing models face numerous challenges in terms of inference efficiency and multimodal data processing. This paper proposes a KunLunBaize-VoT-R1 video inference model based on a long-sequence image encoder, along with its training and application methods. By integrating image packing technology, the Autonomy-of-Experts (AoE) architecture, and combining the video of Thought (VoT), a large language model (LLM) trained with large-scale reinforcement learning, and multiple training techniques, the efficiency and accuracy of the model in video inference tasks are effectively improved. Experiments show that this model performs outstandingly in multiple tests, providing a new solution for video-language understanding.

cs.AI↗

Mapping Dust Attenuation at Kiloparsec Scales. II. Attenuation Curves from Near-Ultraviolet to Near-Infrared

This is the second paper in a series that utilize IFS from MaNGA, NUV imaging from Swift/UVOT and NIR imaging from 2MASS to study dust attenuation properties on kpc scales in nearby galaxies. We apply the method developed in Paper I (Zhou et al. 2023) to the updated SWiM_v4.2 catalog, and measure the optical attenuation curve and the attenuation in three NUV bands for 2487 spaxels selected from 91 galaxies with S/N>20 and $A_V$>0.25. We classify all spaxels into two subsets: star-forming (SF) regions and non-SF regions. We explore the correlations of optical opacity ($A_V$) and the optical and NUV slopes of attenuation curves ($A_B/A_V$ and $A_{w2}/A_{w1}$) with a broad range of stellar and emission-line properties, including specific surface brightness of H$α$ emission, stellar age, stellar and gas-phase metallicity, and diagnostics of recent star formation history. When comparing SF and non-SF regions, we find that $A_V$ and $A_B/A_V$ exhibit similar correlations with all the stellar population and emission-line properties considered, while the NUV slopes in SF regions tend to be flatter than those in non-SF regions. The NUV slope $A_{w2}/A_{w1}$ exhibits an anti-correlation with specific surface brightness of H$α$ emission, a trend that is primarily driven by the positive correlation between $A_{w2}/A_{w1}$ and $Σ_\ast$. The NUV slope flattens in SF regions that contain young stellar populations and have experienced recent star formation, but it shows no obvious dependence on stellar or gas-phase metallicity. The spatially resolved dust attenuation properties exhibit no clear correlations with the inclination of host galaxies or the galactocentric distance of the regions. This finding reinforces the conclusion from Paper I that dust attenuation is primarily regulated by local processes on kpc scales or smaller, rather than by global processes at galactic scales.

astro-ph.GA↗

BigMac: A Communication-Efficient Mixture-of-Experts Model Structure for Fast Training and Inference

The Mixture-of-Experts (MoE) structure scales the Transformer-based large language models (LLMs) and improves their performance with only the sub-linear increase in computation resources. Recently, a fine-grained DeepSeekMoE structure is proposed, which can further improve the computing efficiency of MoE without performance degradation. However, the All-to-All communication introduced by MoE has become a bottleneck, especially for the fine-grained structure, which typically involves and activates more experts, hence contributing to heavier communication overhead. In this paper, we propose a novel MoE structure named BigMac, which is also fine-grained but with high communication efficiency. The innovation of BigMac is mainly due to that we abandon the \textbf{c}ommunicate-\textbf{d}escend-\textbf{a}scend-\textbf{c}ommunicate (CDAC) manner used by fine-grained MoE, which leads to the All-to-All communication always taking place at the highest dimension. Instead, BigMac designs an efficient \textbf{d}escend-\textbf{c}ommunicate-\textbf{c}ommunicate-\textbf{a}scend (DCCA) manner. Specifically, we add a descending and ascending projection at the entrance and exit of the expert, respectively, which enables the communication to perform at a very low dimension. Furthermore, to adapt to DCCA, we re-design the structure of small experts, ensuring that the expert in BigMac has enough complexity to address tokens. Experimental results show that BigMac achieves comparable or even better model quality than fine-grained MoEs with the same number of experts and a similar number of total parameters. Equally importantly, BigMac reduces the end-to-end latency by up to 3.09$\times$ for training and increases the throughput by up to 3.11$\times$ for inference on state-of-the-art AI computing frameworks including Megatron, Tutel, and DeepSpeed-Inference.

cs.LG↗

AVD2: Accident Video Diffusion for Accident Video Description

Traffic accidents present complex challenges for autonomous driving, often featuring unpredictable scenarios that hinder accurate system interpretation and responses. Nonetheless, prevailing methodologies fall short in elucidating the causes of accidents and proposing preventive measures due to the paucity of training data specific to accident scenarios. In this work, we introduce AVD2 (Accident Video Diffusion for Accident Video Description), a novel framework that enhances accident scene understanding by generating accident videos that aligned with detailed natural language descriptions and reasoning, resulting in the contributed EMM-AU (Enhanced Multi-Modal Accident Video Understanding) dataset. Empirical results reveal that the integration of the EMM-AU dataset establishes state-of-the-art performance across both automated metrics and human evaluations, markedly advancing the domains of accident analysis and prevention. Project resources are available at https://an-answer-tree.github.io

cs.CV↗

Internal Heat and Energy Imbalance of Uranus

With its extreme axial tilt, radiant energy budget and internal heat of Uranus remain among the most intriguing mysteries of our Solar System. Here, we present the global radiant energy budget spanning a complete orbital period, revealing significant seasonal variations driven primarily by the highly variable solar flux. Despite these fluctuations, emitted thermal power consistently exceeds absorbed solar power, indicating a net energy loss and ongoing global cooling. Based on the seasonal variations of radiant energy budget, we determine a statistically significant internal heat flux. This finding resolves a long-standing debate over whether Uranus possesses internal heat. We also examine the energy budget of the weather layer by combining the internal heat with the radiant energies, revealing significant energy imbalances at both global and hemispheric scales. These global and hemispheric imbalances should be considered in theoretical and numerical models. The Uranus flagship mission, as recommended by the recent survey, will provide crucial observations to address more unresolved questions and advance our understanding of this enigmatic ice giant.

astro-ph.EP↗

CSST Large Scale Structure Analysis Pipeline: III. Emission-line Redshift Measurement for Slitless Spectra

The China Space Station Telescope (CSST) is a forthcoming space-based optical telescope designed to co-orbit with the Chinese Space Station. With a planned slitless spectroscopic survey spanning a broad wavelength range of $255-1000$nm and an average spectral resolution exceeding 200, the CSST holds significant potential for cosmic large-scale structure analysis. In this study, we focus on redshift determinations from slitless spectra through emission line analysis within the CSST framework. Our tailored redshift measurement process involves identifying emission lines in one-dimensional slitless spectra, aligning observed wavelengths with their rest-frame counterparts from prominent galaxy emissions, and calculating wavelength shifts to determine redshifts accurately. To validate our redshift measurement algorithm, we leverage simulated spectra generated by the CSST emulator for slitless spectroscopy. The outcomes demonstrate a remarkable redshift completeness exceeding 95 per cent for emission line galaxies (ELGs), alongside a purity surpassing 85 per cent. The redshift uncertainty remains impressively below than $\sim 0.001$. Notably, when concentrating on galaxies with more than three matched emission lines, the completeness of ELGs and the purity of measurable galaxies can reach 98 per cent and 97 per cent, respectively. Furthermore, we explore the influence of parameters like magnitude, spectral signal-to-noise ratio, and redshift on redshift completeness and purity. The discussion also delves into redshift degeneracies stemming from emission-line matching confusion. Our developed redshift measurement process will be applied to extensive simulated datasets and forthcoming CSST slitless spectroscopic observations for further cosmological and extragalactic analyses.

astro-ph.CO↗

Enhancing Scalability in Bayesian Nonparametric Factor Analysis of Spatiotemporal Data

This article introduces novel and practicable Bayesian factor analysis frameworks that are computationally feasible for moderate to large spatiotemporal data. Previous Bayesian analysis of spatiotemporal data has utilized a Bayesian factor model with separable temporal latent factors and spatial factor loadings, along with stick-breaking process priors on the loadings to enable clustering of spatial locations. Such a flexible Bayesian model, however, faces a prohibitively high computational cost in posterior sampling when the spatial and temporal dimensions increase to a couple hundred. We address this computational challenge with several speed-up proposals. We integrate a new slice sampling algorithm that permits varying numbers of spatial mixture components across all latent factors and guarantees them to be non-increasing through the posterior sampling iterations, thus effectively reducing the number of mixture parameters. Additionally, we introduce a spatial latent nearest-neighbor Gaussian process prior and new sequential updating algorithms for the spatially varying latent variables in the stick-breaking process prior. Our new models and sampling algorithms exhibit significantly enhanced computational scalability and storage efficiency and possess powerful inferential capabilities for both spatiotemporal prediction and clustering of spatial locations with similar temporal trajectories. The improvement in computational efficiency and inferential performance is substantiated by extensive simulation experiments.

stat.ME↗

MentalArena: Self-play Training of Language Models for Diagnosis and Treatment of Mental Health Disorders

Mental health disorders are one of the most serious diseases in the world. Most people with such a disease lack access to adequate care, which highlights the importance of training models for the diagnosis and treatment of mental health disorders. However, in the mental health domain, privacy concerns limit the accessibility of personalized treatment data, making it challenging to build powerful models. In this paper, we introduce MentalArena, a self-play framework to train language models by generating domain-specific personalized data, where we obtain a better model capable of making a personalized diagnosis and treatment (as a therapist) and providing information (as a patient). To accurately model human-like mental health patients, we devise Symptom Encoder, which simulates a real patient from both cognition and behavior perspectives. To address intent bias during patient-therapist interactions, we propose Symptom Decoder to compare diagnosed symptoms with encoded symptoms, and dynamically manage the dialogue between patient and therapist according to the identified deviations. We evaluated MentalArena against 6 benchmarks, including biomedicalQA and mental health tasks, compared to 6 advanced models. Our models, fine-tuned on both GPT-3.5 and Llama-3-8b, significantly outperform their counterparts, including GPT-4o. We hope that our work can inspire future research on personalized care. Code is available in https://github.com/Scarelette/MentalArena/tree/main

cs.CL↗

De-singularity Subgradient for the $q$-th-Powered $\ell_p$-Norm Weber Location Problem

The Weber location problem is widely used in several artificial intelligence scenarios. However, the gradient of the objective does not exist at a considerable set of singular points. Recently, a de-singularity subgradient method has been proposed to fix this problem, but it can only handle the $q$-th-powered $\ell_2$-norm case ($1\leqslant q<2$), which has only finite singular points. In this paper, we further establish the de-singularity subgradient for the $q$-th-powered $\ell_p$-norm case with $1\leqslant q\leqslant p$ and $1\leqslant p<2$, which includes all the rest unsolved situations in this problem. This is a challenging task because the singular set is a continuum. The geometry of the objective function is also complicated so that the characterizations of the subgradients, minimum and descent direction are very difficult. We develop a $q$-th-powered $\ell_p$-norm Weiszfeld Algorithm without Singularity ($q$P$p$NWAWS) for this problem, which ensures convergence and the descent property of the objective function. Extensive experiments on six real-world data sets demonstrate that $q$P$p$NWAWS successfully solves the singularity problem and achieves a linear computational convergence rate in practical scenarios.

math.OC↗

Spatial-Angular Representation Learning for High-Fidelity Continuous Super-Resolution in Diffusion MRI

Diffusion magnetic resonance imaging (dMRI) often suffers from low spatial and angular resolution due to inherent limitations in imaging hardware and system noise, adversely affecting the accurate estimation of microstructural parameters with fine anatomical details. Deep learning-based super-resolution techniques have shown promise in enhancing dMRI resolution without increasing acquisition time. However, most existing methods are confined to either spatial or angular super-resolution, limiting their effectiveness in capturing detailed microstructural features. Furthermore, traditional pixel-wise loss functions struggle to recover intricate image details essential for high-resolution reconstruction. To address these challenges, we propose SARL-dMRI, a novel Spatial-Angular Representation Learning framework for high-fidelity, continuous super-resolution in dMRI. SARL-dMRI explores implicit neural representations and spherical harmonics to model continuous spatial and angular representations, simultaneously enhancing both spatial and angular resolution while improving microstructural parameter estimation accuracy. To further preserve image fidelity, a data-fidelity module and wavelet-based frequency loss are introduced, ensuring the super-resolved images remain consistent with the original input and retain fine details. Extensive experiments demonstrate that, compared to five other state-of-the-art methods, our method significantly enhances dMRI data resolution, improves the accuracy of microstructural parameter estimation, and provides better generalization capabilities. It maintains stable performance even under a 45$\times$ downsampling factor.

eess.IV↗

Cosmological distance forecasts for the CSST Galaxy Survey using BAO peaks

The measurement of cosmological distances using baryon acoustic oscillations (BAO) is crucial for studying the universe's expansion. The Chinese Space Station Telescope (CSST) galaxy redshift survey, with its vast volume and sky coverage, provides an opportunity to address key challenges in cosmology. However, redshift uncertainties in galaxy surveys can degrade both angular and radial distance estimates. In this study, we forecast the precision of BAO distance measurements using mock CSST galaxy samples, applying a two-point correlation function (2PCF) wedge approach to mitigate redshift errors. We simulate redshift uncertainties of $σ_0 = 0.003$ and $σ_0 = 0.006$, representative of expected CSST errors, and examine their effects on the BAO peak and distance scaling factors, $α_\perp$ and $α_\parallel$, across redshift bins within $0.0 < z \leqslant 1.0$. The wedge 2PCF method proves more effective in detecting the BAO peak compared to the monopole 2PCF, particularly for $σ_0 = 0.006$. Constraints on the BAO peaks show that $α_\perp$ is well constrained around 1.0, regardless of $σ_0$, with precision between 1% and 3% across redshift bins. In contrast, $α_\parallel$ measurements are more sensitive to increases in $σ_0$. For $σ_0 = 0.003$, the results remain close to the fiducial value, with uncertainties ranging between 4% and 9%; for $σ_0 = 0.006$, significant deviations from the fiducial value are observed. We also study the ability to measure parameters $(Ω_m, H_0r_\mathrm{d})$ using distance measurements, proving robust constraints as a cosmological probe under CSST-like redshift uncertainties.

astro-ph.CO↗

Stochastically Constrained Best Arm Identification with Thompson Sampling

We consider the problem of the best arm identification in the presence of stochastic constraints, where there is a finite number of arms associated with multiple performance measures. The goal is to identify the arm that optimizes the objective measure subject to constraints on the remaining measures. We will explore the popular idea of Thompson sampling (TS) as a means to solve it. To the best of our knowledge, it is the first attempt to extend TS to this problem. We will design a TS-based sampling algorithm, establish its asymptotic optimality in the rate of posterior convergence, and demonstrate its superior performance using numerical examples.

cs.LG↗

Reasoning-Enhanced Self-Training for Long-Form Personalized Text Generation

Personalized text generation requires a unique ability of large language models (LLMs) to learn from context that they often do not encounter during their standard training. One way to encourage LLMs to better use personalized context for generating outputs that better align with the user's expectations is to instruct them to reason over the user's past preferences, background knowledge, or writing style. To achieve this, we propose Reasoning-Enhanced Self-Training for Personalized Text Generation (REST-PG), a framework that trains LLMs to reason over personal data during response generation. REST-PG first generates reasoning paths to train the LLM's reasoning abilities and then employs Expectation-Maximization Reinforced Self-Training to iteratively train the LLM based on its own high-reward outputs. We evaluate REST-PG on the LongLaMP benchmark, consisting of four diverse personalized long-form text generation tasks. Our experiments demonstrate that REST-PG achieves significant improvements over state-of-the-art baselines, with an average relative performance gain of 14.5% on the benchmark.

cs.CL↗