Searcharxiv⌕ Search

arXiv subjects

Yuguang Wang

Publications and source records attributed to Yuguang Wang.

14 recordsLinked to original sources

Collaborative Parameter Learning: Mitigating Forgetting via Parameter-Level Gradient Analysis

Catastrophic forgetting during knowledge injection impairs the ability of large language models to acquire new knowledge without overwriting previously mastered knowledge. Recent studies analyze forgetting from a gradient similarity perspective and mitigate forgetting through vector projection. However, these methods primarily characterize gradient similarity at the aggregate direction level, leaving the parameter wise contributions to forgetting underexplored. In this paper, we decompose gradient similarity into parameter wise contributions and identify two types of parameters during forgetting: Conflicting Parameters, whose updates contribute to forgetting and typically account for 50 percent to 75 percent of parameters, and Collaborative Parameters, whose updates mitigate forgetting and account for 25 percent to 50 percent. Based on this analysis, we propose Collaborative Parameter Learning, CPL, a parameter wise training rule that freezes Conflicting Parameters and updates only Collaborative Parameters. Experiments comparing CPL with seven baseline methods show that CPL learns 20.2% to 48.2% more questions with negligible forgetting, while reducing peak VRAM by approximately 3 GB per billion model parameters and computation time by 16.5 percent. Extensive evaluations on parameter consumption, out of set generalization, cross prompt generalization, multimodal tasks, open ended question answering, and multilingual settings demonstrate that CPL effectively mitigates forgetting across diverse scenarios.

cs.LG↗

ProtoCycle: Reflective Tool-Augmented Planning for Text-Guided Protein Design

Designing proteins that satisfy natural language functional requirements is a central goal in protein engineering. A straightforward baseline is to fine-tune generic instruction-tuned LLMs as direct text-to-sequence generators, but this is data- and compute-hungry. With limited supervision, LLMs can produce coherent plans in text yet fail to reliably realize them as sequences. This plan-execute gap motivates ProtoCycle, an agentic framework for protein design that uses LLMs primarily to drive a multi-round, feedback-driven decision cycle. ProtoCycle couples an LLM planner with a lightweight tool environment designed to emulate the iterative workflow of human protein engineering and uses LLM-driven reflection on tool feedback to revise plans. Trained with supervised trajectories and online reinforcement learning, ProtoCycle achieves strong language alignment while maintaining competitive foldability, and ablations show that reflection substantially improves sequence quality.

q-bio.QM↗

Magnetic-field-tunable commensurate multi-q charge orders on UTe2 (011) surface

The heavy-fermion superconductor UTe2 has attracted intense interest as a candidate for spin-triplet pairing. Recent scanning tunneling microscopy (STM) studies have reported complex charge orders (COs) on its (011) surface, but their origin and relationship with superconductivity remain controversial. Here, by performing temperature-, magnetic field-, and sample-dependent STM measurements, we identify multiple new CO wave vectors beyond those previously reported. All these CO wave vectors are strictly locked to integer multiples of 1/14 and 1/4 of the reciprocal lattice vectors of the UTe2 (011) surface, and multiple of them coexist in real space, collectively revealing a family of field-tunable, commensurate multi-q COs. These COs exist within an energy range much larger than the superconducting energy scale, their emergence suppresses the density of states near EF, yet show negligible coupling to bulk superconductivity and magnetic vortices. Our findings strongly disfavor the Fermi surface nesting or primary pair-density-wave pictures, but are consistent with a surface parent spin order.

cond-mat.supr-con↗

How Out-of-Distribution Detection Learning Theory Enhances Transformer: Learnability and Reliability

Transformers excel in natural language processing and computer vision tasks. However, they still face challenges in generalizing to Out-of-Distribution (OOD) datasets, i.e. data whose distribution differs from that seen during training. OOD detection aims to distinguish outliers while preserving in-distribution (ID) data performance. This paper introduces the OOD detection Probably Approximately Correct (PAC) Theory for transformers, which establishes the conditions for data distribution and model configurations for the OOD detection learnability of transformers. It shows that outliers can be accurately represented and distinguished with sufficient data under conditions. The theoretical implications highlight the trade-off between theoretical principles and practical training paradigms. By examining this trade-off, we naturally derived the rationale for leveraging auxiliary outliers to enhance OOD detection. Our theory suggests that by penalizing the misclassification of outliers within the loss function and strategically generating soft synthetic outliers, one can robustly bolster the reliability of transformer networks. This approach yields a novel algorithm that ensures learnability and refines the decision boundaries between inliers and outliers. In practice, the algorithm consistently achieves state-of-the-art (SOTA) performance across various data formats.

cs.LG↗

How Particle-System Random Batch Methods Enhance Graph Transformer: Memory Efficiency and Parallel Computing Strategy

Attention mechanism is a significant part of Transformer models. It helps extract features from embedded vectors by adding global information and its expressivity has been proved to be powerful. Nevertheless, the quadratic complexity restricts its practicability. Although several researches have provided attention mechanism in sparse form, they are lack of theoretical analysis about the expressivity of their mechanism while reducing complexity. In this paper, we put forward Random Batch Attention (RBA), a linear self-attention mechanism, which has theoretical support of the ability to maintain its expressivity. Random Batch Attention has several significant strengths as follows: (1) Random Batch Attention has linear time complexity. Other than this, it can be implemented in parallel on a new dimension, which contributes to much memory saving. (2) Random Batch Attention mechanism can improve most of the existing models by replacing their attention mechanisms, even many previously improved attention mechanisms. (3) Random Batch Attention mechanism has theoretical explanation in convergence, as it comes from Random Batch Methods on computation mathematics. Experiments on large graphs have proved advantages mentioned above. Also, the theoretical modeling of self-attention mechanism is a new tool for future research on attention-mechanism analysis.

cs.LG↗

Yin-Yang vortex on UTe2 (011) surface

UTe2 is a promising candidate for spin-triplet superconductor, yet its exact superconducting order parameter remains highly debated. Here, via scanning tunneling microscopy/spectroscopy, we observe a novel type of magnetic vortex with distinct dark-bright contrast in local density of states on UTe2 (011) surface under a perpendicular magnetic field, resembling the conjugate structure of Yin-Yang diagram in Taoism. Each Yin-Yang vortex contains a quantized magnetic flux, and the boundary between the Yin and Yang parts aligns with the crystallographic a-axis of UTe2. The vortex states exhibit intriguing behaviors -- a sharp zero-energy conductance peak exists at the Yang part, while a superconducting gap with pronounced coherence peaks exists at the Yin part, which is even sharper than those measured far from the vortex core or in the absence of magnetic field. By theoretical modeling, we show that the Yin-Yang vortices on UTe2 (011) surface can be explained by the asymmetric vortex-derived local distortion of the zero-energy surface states associated with spin-triplet pairing with appropriate d-vectors. Therefore, the observation of Yin-Yang vortex confirms the spin-triplet pairing in UTe2 and imposes constraints on the candidate d-vector for the spin-triplet pairing.

cond-mat.supr-con↗

Progressive Residual Extraction based Pre-training for Speech Representation Learning

Self-supervised learning (SSL) has garnered significant attention in speech processing, excelling in linguistic tasks such as speech recognition. However, jointly improving the performance of pre-trained models on various downstream tasks, each requiring different speech information, poses significant challenges. To this purpose, we propose a progressive residual extraction based self-supervised learning method, named ProgRE. Specifically, we introduce two lightweight and specialized task modules into an encoder-style SSL backbone to enhance its ability to extract pitch variation and speaker information from speech. Furthermore, to prevent the interference of reinforced pitch variation and speaker information with irrelevant content information learning, we residually remove the information extracted by these two modules from the main branch. The main branch is then trained using HuBERT's speech masking prediction to ensure the performance of the Transformer's deep-layer features on content tasks. In this way, we can progressively extract pitch variation, speaker, and content representations from the input speech. Finally, we can combine multiple representations with diverse speech information using different layer weights to obtain task-specific representations for various downstream tasks. Experimental results indicate that our proposed method achieves joint performance improvements on various tasks, such as speaker identification, speech recognition, emotion recognition, speech enhancement, and voice conversion, compared to excellent SSL methods such as wav2vec2.0, HuBERT, and WavLM.

eess.AS↗

AB$\mathbb{C}$MB: Deep Delensing Assisted Likelihood-Free Inference from CMB Polarization Maps

The existence of a cosmic background of primordial gravitational waves (PGWB) is a robust prediction of inflationary cosmology, but it has so far evaded discovery. The most promising avenue of its detection is via measurements of Cosmic Microwave Background (CMB) $B$-polarization. However, this is not straightforward due to (a) the fact that CMB maps are distorted by gravitational lensing and (b) the high-dimensional nature of CMB data, which renders likelihood-based analysis methods computationally extremely expensive. In this paper, we introduce an efficient likelihood-free, end-to-end inference method to directly infer the posterior distribution of the tensor-to-scalar ratio $r$ from lensed maps of the Stokes $Q$ and $U$ polarization parameters. Our method employs a generative model to delense the maps and utilizes the Approximate Bayesian Computation (ABC) algorithm to sample $r$. We demonstrate that our method yields unbiased estimates of $r$ with well-calibrated uncertainty quantification.

astro-ph.CO↗

Can neural networks learn persistent homology features?

Topological data analysis uses tools from topology -- the mathematical area that studies shapes -- to create representations of data. In particular, in persistent homology, one studies one-parameter families of spaces associated with data, and persistence diagrams describe the lifetime of topological invariants, such as connected components or holes, across the one-parameter family. In many applications, one is interested in working with features associated with persistence diagrams rather than the diagrams themselves. In our work, we explore the possibility of learning several types of features extracted from persistence diagrams using neural networks.

cs.LG↗

PAN: Path Integral Based Convolution for Deep Graph Neural Networks

Convolution operations designed for graph-structured data usually utilize the graph Laplacian, which can be seen as message passing between the adjacent neighbors through a generic random walk. In this paper, we propose PAN, a new graph convolution framework that involves every path linking the message sender and receiver with learnable weights depending on the path length, which corresponds to the maximal entropy random walk. PAN generalizes the graph Laplacian to a new transition matrix we call \emph{maximal entropy transition} (MET) matrix derived from a path integral formalism. Most previous graph convolutional network architectures can be adapted to our framework, and many variations and derivatives based on the path integral idea can be developed. Experimental results show that the path integral based graph neural networks have great learnability and fast convergence rate, and achieve state-of-the-art performance on benchmark tasks.

cs.LG↗

Approximation by boolean sums of Jackson operators on the sphere

This paper concerns the approximation by the Boolean sums of Jackson operators $\oplus^rJ_{k,s}(f)$ on the unit sphere $\mathbb S^{n-1}$ of $\mathbb{R}^{n}$. We prove the following the direct and inverse theorem for $\oplus^rJ_{k,s}(f)$: there are constants $C_1$ and $C_2$ such that \begin{equation*} C_1\|\oplus^rJ_{k,s}f-f\|_p \leq ω^{2r}(f,k^{-1})_p \leq C_2 \max_{v\geq k}\|\oplus^rJ_{k,s}f-f\|_p \end{equation*} for any positive integer $k$ and any $p$th Lebesgue integrable functions $f$ defined on $\mathbb S^{n-1}$, where $ω^{2r}(f,t)_p$ is the modulus of smoothness of degree $2r$ of $f$. We also prove that the saturation order for $\oplus^rJ_{k,s}$ is $k^{-2r}$.

math.CA↗

A study on effectiveness of extreme learning machine

Extreme learning machine (ELM), proposed by Huang et al., has been shown a promising learning algorithm for single-hidden layer feedforward neural networks (SLFNs). Nevertheless, because of the random choice of input weights and biases, the ELM algorithm sometimes makes the hidden layer output matrix H of SLFN not full column rank, which lowers the effectiveness of ELM. This paper discusses the effectiveness of ELM and proposes an improved algorithm called EELM that makes a proper selection of the input weights and bias before calculating the output weights, which ensures the full column rank of H in theory. This improves to some extend the learning rate (testing accuracy, prediction accuracy, learning time) and the robustness property of the networks. The experimental results based on both the benchmark function approximation and real-world problems including classification and regression applications show the good performances of EELM.

cs.NE↗

The Direct and Converse Inequalities for Jackson-Type Operators on Spherical Cap

Approximation on the spherical cap is different from that on the sphere which requires us to construct new operators. This paper discusses the approximation on the spherical cap. That is, so called Jackson-type operator $\{J_{k,s}^m\}_{k=1}^{\infty}$ is constructed to approximate the function defined on the spherical cap $D(x_0,γ)$. We thus establish the direct and inverse inequalities and obtain saturation theorems for $\{J_{k,s}^m\}_{k=1}^{\infty}$ on the cap $D(x_0,γ)$. Using methods of $K$-functional and multiplier, we obtain the inequality \begin{eqnarray*} C_1\:\| J_{k,s}^m(f)-f\|_{D,p}\leq ω^2\left(f,\:k^{-1}\right)_{D,p} \leq C_2 \max_{v\geq k}\| J_{v,s}^m(f) - f\|_{D,p} \end{eqnarray*} and that the saturation order of these operators is $O(k^{-2})$, where $ω^2\left(f,\:t\right)_{D,p}$ is the modulus of smoothness of degree 2, the constants $C_1$ and $C_2$ are independent of $k$ and $f$.

math.CA↗

Approximation by Semigroups of Spherical Operators

This paper discusses the approximation by %semigroups of operators of class ($\mathscr{C}_0$) on the sphere and focuses on a class of so called exponential-type multiplier operators. It is proved that such operators form a strongly continuous semigroup of contraction operators of class ($\mathscr{C}_0$), from which the equivalence between approximation for these operators and $K$-functionals introduced by the operators is given. As examples, the constructed $r$-th Boolean of generalized spherical Abel-Poisson operator and $r$-th Boolean of generalized spherical Weierstrass operator denoted by $\oplus^r V_t^γ$ and $\oplus^r W_t^κ$ separately ($r$ is any positive integer, $0<γ,κ\leq1$ and $t>0$) satisfy that $\|\oplus^r V_t^γf - f\|_{\mathcal{X}}\approx ω^{rγ}(f,t^{1/γ})_{\mathcal{X}}$ and $\|\oplus^r W_t^κf - f\|_{\mathcal{X}}\approx ω^{2rγ}(f,t^{1/(2κ)})_{\mathcal{X}}$, for all $f\in \mathcal{X}$, where $\mathcal{X}$ is a Banach space of continuous functions or $\mathcal{L}^p$-integrable functions ($1\leq p<\infty$) and $\|\cdot\|_{\mathcal{X}}$ is the norm on $\mathcal{X}$ and $ω^s(f,t)_{\mathcal{X}}$ is the moduli of smoothness of degree $s>0$ for $f\in \mathcal{X}$. The saturation order and saturation class of the regular exponential-type multiplier operators with positive kernels are also obtained. Moreover, it is proved that $\oplus^r V_t^γ$ and $\oplus^r W_t^κ$ have the same saturation class if $γ=2κ$.

math.CA↗