Searcharxiv⌕ Search

arXiv subjects

Ping Li

Publications and source records attributed to Ping Li.

At least 55 records · Page 3Linked to original sources

Gradient Descent Finds Over-Parameterized Neural Networks with Sharp Generalization for Nonparametric Regression

We study nonparametric regression by an over-parameterized two-layer neural network trained by gradient descent (GD) in this paper. We show that, if the neural network is trained by GD with early stopping, then the trained network renders a sharp rate of the nonparametric regression risk of $\mathcal{O}(ε_n^2)$, which is the same rate as that for the classical kernel regression trained by GD with early stopping, where $ε_n$ is the critical population rate of the Neural Tangent Kernel (NTK) associated with the network and $n$ is the size of the training data. It is remarked that our result does not require distributional assumptions about the covariate as long as the covariate is bounded, in a strong contrast with many existing results which rely on specific distributions of the covariates such as the spherical uniform data distribution or distributions satisfying certain restrictive conditions. The rate $\mathcal{O}(ε_n^2)$ is known to be minimax optimal for specific cases, such as the case that the NTK has a polynomial eigenvalue decay rate which happens under certain distributional assumptions on the covariates. Our result formally fills the gap between training a classical kernel regression model and training an over-parameterized but finite-width neural network by GD for nonparametric regression without distributional assumptions on the bounded covariate. We also provide confirmative answers to certain open questions or address particular concerns in the literature of training over-parameterized neural networks by GD with early stopping for nonparametric regression, including the characterization of the stopping time, the lower bound for the network width, and the constant learning rate used in GD.

stat.ML↗

Orbital and Pulsation Analysis of 42 Heartbeat Stars Discovered in TESS Data

Heartbeat stars (HBSs) are ideal laboratories for studying the formation and evolution of binary stars in eccentric orbits and their mutual tidal interactions. We present 42 new HBSs discovered based on TESS-SPOC and QLP data. Their physical parameters have been obtained through modeling with appropriate models. Subsequently, Tidally excited oscillations (TEOs) are detected in ten systems, and their pulsation phases and modes are identified. Most pulsation phases can be explained by the dominant being spherical harmonic degree $l=2$ and azimuthal order $m=0$ or $\pm2$. For TIC 156846634, the harmonic with large deviation ($>3σ$) from the expected adiabatic phase can be expected to be a traveling wave or significantly nonadiabatic. The harmonic numbers $n$ = 16 in TIC 184413651 may not be considered as a TEO candidate due to its large deviation ($>2σ$) from the adiabatic expectation. Moreover, TIC 92828790 shows no TEOs but exhibits a significant $γ$\,Dor-type pulsation. The eccentricity-period ($e-P$) relation also shows a positive correlation between eccentricity and period, as well as the existence of orbital circularization. The Hertzsprung-Russell diagram shows that TESS HBSs have higher temperatures and greater luminosities than Kepler HBSs, possibly due to selection effects. This significantly enhances the detectability of massive HBSs and those containing TEOs.

astro-ph.SR↗

Hamiltonian circle action, invariant hypersurface and the complex projective space

Let $M$ be a $2n$-dimensional closed symplectic manifold admitting a Hamiltonian circle action with isolated fixed points. We show that if $M$ contains an $S^1$-invariant symplectic hypersurface $D$ such that $M\setminus D$ is a homology cell, which is satisfied when $M\setminus D$ is contractible, then $M$ and $D$ are homotopy complex projective spaces with standard Chern classes and the $S^1$-representations on the fixed-point set of $(M,D)$ are the same as those arising from the standard linear actions on $(\mathbb{P}^n,\mathbb{P}^{n-1})$, provided that $n \not \equiv 3 \pmod 4$. This can be viewed as the transformation group analogue to a recent result obtained by Peternell and the author, where the latter was conjectured by Fujita more than four decades ago.

math.DG↗

Beyond Single Images: Retrieval Self-Augmented Unsupervised Camouflaged Object Detection

At the core of Camouflaged Object Detection (COD) lies segmenting objects from their highly similar surroundings. Previous efforts navigate this challenge primarily through image-level modeling or annotation-based optimization. Despite advancing considerably, this commonplace practice hardly taps valuable dataset-level contextual information or relies on laborious annotations. In this paper, we propose RISE, a RetrIeval SElf-augmented paradigm that exploits the entire training dataset to generate pseudo-labels for single images, which could be used to train COD models. RISE begins by constructing prototype libraries for environments and camouflaged objects using training images (without ground truth), followed by K-Nearest Neighbor (KNN) retrieval to generate pseudo-masks for each image based on these libraries. It is important to recognize that using only training images without annotations exerts a pronounced challenge in crafting high-quality prototype libraries. In this light, we introduce a Clustering-then-Retrieval (CR) strategy, where coarse masks are first generated through clustering, facilitating subsequent histogram-based image filtering and cross-category retrieval to produce high-confidence prototypes. In the KNN retrieval stage, to alleviate the effect of artifacts in feature maps, we propose Multi-View KNN Retrieval (MVKR), which integrates retrieval results from diverse views to produce more robust and precise pseudo-masks. Extensive experiments demonstrate that RISE outperforms state-of-the-art unsupervised and prompt-based methods. Code is available at https://github.com/xiaohainku/RISE.

cs.CV↗

Sharp Generalization for Nonparametric Regression in Interpolation Space by Over-Parameterized Neural Networks Trained with Preconditioned Gradient Descent and Early Stopping

We study nonparametric regression using an over-parameterized two-layer neural networks trained with algorithmic guarantees in this paper. We consider the setting where the training features are drawn uniformly from the unit sphere in $\RR^d$, and the target function lies in an interpolation space commonly studied in statistical learning theory. We demonstrate that training the neural network with a novel Preconditioned Gradient Descent (PGD) algorithm, equipped with early stopping, achieves a sharp regression rate of $\cO(n^{-\frac{2αs'}{2αs'+1}})$ when the target function is in the interpolation space $\bth{\cH_K}^{s'}$ with $s' \ge 3$. This rate is even sharper than the currently known nearly-optimal rate of $\cO(n^{-\frac{2αs'}{2αs'+1}})\log^2(1/δ)$~\citep{Li2024-edr-general-domain}, where $n$ is the size of the training data and $δ\in (0,1)$ is a small probability. This rate is also sharper than the standard kernel regression rate of $\cO(n^{-\frac{2α}{2α+1}})$ obtained under the regular Neural Tangent Kernel (NTK) regime when training the neural network with the vanilla gradient descent (GD), where $2α= d/(d-1)$. Our analysis is based on two key technical contributions. First, we present a principled decomposition of the network output at each PGD step into a function in the reproducing kernel Hilbert space (RKHS) of a newly induced integral kernel, and a residual function with small $L^{\infty}$-norm. Second, leveraging this decomposition, we apply local Rademacher complexity theory to tightly control the complexity of the function class comprising all the neural network functions obtained in the PGD iterates. Our results further suggest that PGD enables the neural network to escape the linear NTK regime and achieve improved generalization.

stat.ML↗

Vanishing theorems and rational connectedness on holomorphic tensor fields

A vanishing theorem for uniformly RC $k$-positive Hermitian holomorphic vector bundles is established. It turns out that the holomorphic tangent bundle of a compact complex manifold equipped with a positive $k$-Ricci curvature Kähler metric is uniformly RC $k$-positive. Two main applications are presented. The first one is to deduce that spaces of some holomorphic tensor fields on such Kähler or more generally Kähler-like Hermitian manifolds are trivial, generalizing some recent results. The second one is to show that a compact Kähler manifold whose holomorphic tangent bundle can be endowed with either a uniformly RC $k$-positive Hermitian metric or a positive $k$-Ricci curvature Kähler-like Hermitian metric is projective and rationally connected.

math.DG↗

Rotating black holes in de Rham-Gabadadze-Tolley massive gravity: Analytic calculation procedure

In this paper, we explore the solutions of rotating black holes within the framework of de Rham-Gabadadze-Tolley (dRGT) massive gravity. We provide a detailed, step-by-step analytical derivation of these solutions. Our solutions are characterized by several parameters: mass $M$ , electric charge $Q_{*}$, angular momentum $a$, and a graviton mass $m$. This graviton mass term incorporates both a cosmological constant $Λ$ and a Stückelberg charge $S_{*}$ into the black hole parameters. These solutions may serve as potential candidates for astrophysical black holes.

gr-qc↗

SEVEN: Pruning Transformer Model by Reserving Sentinels

Large-scale Transformer models (TM) have demonstrated outstanding performance across various tasks. However, their considerable parameter size restricts their applicability, particularly on mobile devices. Due to the dynamic and intricate nature of gradients on TM compared to Convolutional Neural Networks, commonly used pruning methods tend to retain weights with larger gradient noise. This results in pruned models that are sensitive to sparsity and datasets, exhibiting suboptimal performance. Symbolic Descent (SD) is a general approach for training and fine-tuning TM. In this paper, we attempt to describe the noisy batch gradient sequences on TM through the cumulative process of SD. We utilize this design to dynamically assess the importance scores of weights.SEVEN is introduced by us, which particularly favors weights with consistently high sensitivity, i.e., weights with small gradient noise. These weights are tended to be preserved by SEVEN. Extensive experiments on various TM in natural language, question-answering, and image classification domains are conducted to validate the effectiveness of SEVEN. The results demonstrate significant improvements of SEVEN in multiple pruning scenarios and across different sparsity levels. Additionally, SEVEN exhibits robust performance under various fine-tuning strategies. The code is publicly available at https://github.com/xiaojinying/SEVEN.

cs.LG↗

LNPT: Label-free Network Pruning and Training

Pruning before training enables the deployment of neural networks on smart devices. By retaining weights conducive to generalization, pruned networks can be accommodated on resource-constrained smart devices. It is commonly held that the distance on weight norms between the initialized and the fully-trained networks correlates with generalization performance. However, as we have uncovered, inconsistency between this metric and generalization during training processes, which poses an obstacle to determine the pruned structures on smart devices in advance. In this paper, we introduce the concept of the learning gap, emphasizing its accurate correlation with generalization. Experiments show that the learning gap, in the form of feature maps from the penultimate layer of networks, aligns with variations of generalization performance. We propose a novel learning framework, LNPT, which enables mature networks on the cloud to provide online guidance for network pruning and learning on smart devices with unlabeled data. Our results demonstrate the superiority of this approach over supervised training.

cs.LG↗

TED: Accelerate Model Training by Internal Generalization

Large language models have demonstrated strong performance in recent years, but the high cost of training drives the need for efficient methods to compress dataset sizes. We propose TED pruning, a method that addresses the challenge of overfitting under high pruning ratios by quantifying the model's ability to improve performance on pruned data while fitting retained data, known as Internal Generalization (IG). TED uses an optimization objective based on Internal Generalization Distance (IGD), measuring changes in IG before and after pruning to align with true generalization performance and achieve implicit regularization. The IGD optimization objective was verified to allow the model to achieve the smallest upper bound on generalization error. The impact of small mask fluctuations on IG is studied through masks and Taylor approximation, and fast estimation of IGD is enabled. In analyzing continuous training dynamics, the prior effect of IGD is validated, and a progressive pruning strategy is proposed. Experiments on image classification, natural language understanding, and large language model fine-tuning show TED achieves lossless performance with 60-70\% of the data. Upon acceptance, our code will be made publicly available.

cs.LG↗

Planar Turán numbers of three configurations

The planar Tuán number of $H$, denoted by $ex_{\mathcal{P}}(n,H)$, is defined as the maximum number of edges in an $n$-vertex $H$-free planar graph. The exact value of $ex_{\mathcal{P}}(n,H)$ remains a mystery when $H$ is large (for example, $H$ is a long path or a long cycle), while tight bounds have been established for many small planar graphs such as cycles, paths, $Θ$-graphs and other small graphs formed by a union of them. One representative graph among such union graphs is $K_1+L$ where $L$ is a linear forest without isolated vertices. Previous works solved the cases when $L$ is a path or a matching. In this work, we first investigate the planar Turán number of the graph $K_1+L$ when $L$ is the disjoint union of a $P_2$ and $P_3$. Equivalently, $K_1+L$ represents a specific configuration formed by combining a $C_3$ and a $Θ_4$. We further consider the planar Turán numbers of the all graphs obtained by combining $C_3$ and $Θ_4$. Among the six possible such configurations, three have been resolved in earlier works. For the remaining three configurations (including $K_1+(P_2\dot{\cup}P_3)$), we derive tight bounds. Furthermore, we completely characterize all extremal graphs for the remaining two of these three cases.

math.CO↗

Heartbeat Stars Recognition Based on Recurrent Neural Networks: Method and Validation

Since the variety of their light curve morphologies, the vast majority of the known heartbeat stars (HBSs) have been discovered by manual inspection. Machine learning, which has already been successfully applied to the classification of variable stars based on light curves, offers another possibility for the automatic detection of HBSs. We propose a novel feature extraction approach for HBSs. First, the orbital frequencies are calculated automatically according to the Fourier spectra of the light curves. Then, the amplitudes of the first 100 harmonics are extracted. Finally, these harmonics are normalized as feature vectors of the light curve. A training data set of synthetic light curves is constructed using ELLC, and their features are fed into recurrent neural networks (RNNs) for supervised learning, with the expected output being the eccentricity of these light curves. The performance of the RNNs is evaluated using a test data set of synthetic light curves, achieving 95$\%$ accuracy. When applied to known HBSs from the OGLE, Kepler, and TESS surveys, the networks achieve an average accuracy of 86$\%$. This method successfully identifies four new HBSs within the eclipsing binary catalog of Kirk et al. The use of orbital harmonics as features for HBSs proves to be a practical approach that significantly reduces the computational cost of neural networks. RNNs show excellent performance in recognizing this type of time series data. This method not only allows efficient identification of HBSs but can also be extended to recognize other types of periodic variable stars.

astro-ph.SR↗

Adacc: An Adaptive Framework Unifying Compression and Activation Recomputation for LLM Training

Training large language models (LLMs) is often constrained by GPU memory limitations. To alleviate memory pressure, activation recomputation and data compression have been proposed as two major strategies. However, both approaches have limitations: recomputation introduces significant training overhead, while compression can lead to accuracy degradation and computational inefficiency when applied naively. In this paper, we propose Adacc, the first adaptive memory optimization framework that unifies activation recomputation and data compression to improve training efficiency for LLMs while preserving model accuracy. Unlike existing methods that apply static, rule-based strategies or rely solely on one technique, Adacc makes fine-grained, tensor-level decisions, dynamically selecting between recomputation, retention, and compression based on tensor characteristics and runtime hardware constraints. Adacc tackles three key challenges: (1) it introduces layer-specific compression algorithms that mitigate accuracy loss by accounting for outliers in LLM activations; (2) it employs a MILP-based scheduling policy to globally optimize memory strategies across layers; and (3) it integrates an adaptive policy evolution mechanism to update strategies during training in response to changing data distributions. Experimental results show that Adacc improves training throughput by 1.01x to 1.37x compared to state-of-the-art frameworks, while maintaining accuracy comparable to the baseline.

cs.LG↗

CaliDrop: KV Cache Compression with Calibration

Large Language Models (LLMs) require substantial computational resources during generation. While the Key-Value (KV) cache significantly accelerates this process by storing attention intermediates, its memory footprint grows linearly with sequence length, batch size, and model size, creating a bottleneck in long-context scenarios. Various KV cache compression techniques, including token eviction, quantization, and low-rank projection, have been proposed to mitigate this bottleneck, often complementing each other. This paper focuses on enhancing token eviction strategies. Token eviction leverages the observation that the attention patterns are often sparse, allowing for the removal of less critical KV entries to save memory. However, this reduction usually comes at the cost of notable accuracy degradation, particularly under high compression ratios. To address this issue, we propose \textbf{CaliDrop}, a novel strategy that enhances token eviction through calibration. Our preliminary experiments show that queries at nearby positions exhibit high similarity. Building on this observation, CaliDrop performs speculative calibration on the discarded tokens to mitigate the accuracy loss caused by token eviction. Extensive experiments demonstrate that CaliDrop significantly improves the accuracy of existing token eviction methods.

cs.CL↗

SGCap: Decoding Semantic Group for Zero-shot Video Captioning

Zero-shot video captioning aims to generate sentences for describing videos without training the model on video-text pairs, which remains underexplored. Existing zero-shot image captioning methods typically adopt a text-only training paradigm, where a language decoder reconstructs single-sentence embeddings obtained from CLIP. However, directly extending them to the video domain is suboptimal, as applying average pooling over all frames neglects temporal dynamics. To address this challenge, we propose a Semantic Group Captioning (SGCap) method for zero-shot video captioning. In particular, it develops the Semantic Group Decoding (SGD) strategy to employ multi-frame information while explicitly modeling inter-frame temporal relationships. Furthermore, existing zero-shot captioning methods that rely on cosine similarity for sentence retrieval and reconstruct the description supervised by a single frame-level caption, fail to provide sufficient video-level supervision. To alleviate this, we introduce two key components, including the Key Sentences Selection (KSS) module and the Probability Sampling Supervision (PSS) module. The two modules construct semantically-diverse sentence groups that models temporal dynamics and guide the model to capture inter-sentence causal relationships, thereby enhancing its generalization ability to video captioning. Experimental results on several benchmarks demonstrate that SGCap significantly outperforms previous state-of-the-art zero-shot alternatives and even achieves performance competitive with fully supervised ones. Code is available at https://github.com/mlvccn/SGCap_Video.

cs.CV↗

Planar Turán number of disjoint union of $C_3$ and $C_5$

The planar Turán number of $H$, denoted by $ex_{\mathcal{P}}(n,H)$, is the maximum number of edges in an $n$-vertex $H$-free planar graph. The planar Turán number of $k\geq 3$ vertex-disjoint union of cycles is the trivial value $3n-6$. Let $C_{\ell}$ denote the cycle of length $\ell$ and $C_{\ell}\cup C_t$ denote the union of disjoint cycles $C_{\ell}$ and $C_t$. The planar Turán number $ex_{\mathcal{P}}(n,H)$ is known if $H=C_{\ell}\cup C_k$, where $\ell,k\in \{3,4\}$. In this paper, we determine the value $ex_{\mathcal{P}}(n,C_3\cup C_5)=\lfloor\frac{8n-13}{3}\rfloor$ and characterize the extremal graphs when $n$ is sufficiently large.

math.CO↗

Downregulation of aquaporin 3 promotes hyperosmolarity-induced apoptosis of nucleus pulposus cells through PI3K/Akt/mTOR pathway suppression

Hyperosmolarity is a key contributor to nucleus pulposus cell (NPC) apoptosis during intervertebral disc degeneration (IVDD). Aquaporin 3 (AQP3), a membrane channel protein, regulates cellular osmotic balance by transporting water and osmolytes. Although AQP3 downregulation is associated with disc degeneration, its role in apoptosis under hyperosmotic conditions remains unclear. Here, we demonstrate that hyperosmolarity induces AQP3 depletion, suppresses the PI3K/AKT/mTOR signaling pathway, and promotes mitochondrial dysfunction and ROS accumulation in NPCs. Lentiviral overexpression of AQP3 restores this pathway, attenuates oxidative damage, and reduces apoptosis, preserving disc structure in IVDD rat models. In contrast, pharmacological inhibition of AQP3 exacerbates ECM catabolism and NP tissue loss. Our findings reveal that AQP3 deficiency under hyperosmolarity contributes to NPC apoptosis via suppression of PI3K/AKT/mTOR signaling, potentially creating a pathological cycle of disc degeneration. These results highlight AQP3 as a promising therapeutic target for IVDD.

q-bio.BM↗

An eco-friendly universal strategy via ribavirin to achieve highly efficient and stable perovskite solar cells

The grain boundaries of perovskite films prepared by the solution method are highly disordered, with a large number of defects existing at the grain boundaries. These defect sites promote the decomposition of perovskite. Here, we use ribavirin obtained through bacillus subtilis fermentation to regulate the crystal growth of perovskite, inducing changes in the work function and energy level structure of perovskite, which significantly reduces the defect density. Based on density functional theory calculations, the defect formation energies of VI, VMA, VPb, and PbI in perovskite are improved. This increases the open-circuit voltage of perovskite solar cells (PSCs) (ITO/PEDOT:PSS/perovskite/PCBM/BCP/Ag) from 1.077 to 1.151 V, and the PCE increases significantly from 17.05% to 19.86%. Unencapsulated PSCs were stored in the environment (humidity approximately 35+-5%) for long-term stability testing. After approximately 900 hours of storage, the PCE of the ribavirin-based device retains 84.33% of its initial PCE, while the control-based device retains only 13.44% of its initial PCE. The PCE of PSCs (ITO/SnO2/perovskite/Spiro-OMETAD/Ag) is increased from 20.16% to 22.14%, demonstrating the universality of this doping method. This universal doping strategy provides a new approach for improving the efficiency and stability of PSCs using green molecular doping strategies.

cond-mat.mtrl-sci↗