SearcharxivSearch

arXiv subjects

Yuzhu Chen

Publications and source records attributed to Yuzhu Chen.

13 recordsLinked to original sources

Geodesic Trapping and Escape of Active Particles on Curved Surfaces

Active particles on curved surfaces can become trapped along closed geodesics even without physical barriers. We show that escape from these geometric traps exposes a fundamental distinction between continuous and discrete reorientation. At high Péclet numbers, active Brownian particles escape efficiently through rotational diffusion, whereas run-and-tumble particles remain trapped much longer; at low Péclet numbers, both reduce to passive diffusion. Gaussian curvature controls escape by focusing or defocusing neighboring geodesics, producing distinct asymptotic scalings of the mean exit time.

cond-mat.soft

Drawback of Enforcing Equivariance and its Compensation via the Lens of Expressive Power

Equivariant neural networks encode the intrinsic symmetry of data as an inductive bias, which has achieved impressive performance in wide domains. However, the understanding to their expressive power remains premature. Focusing on 2-layer ReLU networks, this paper investigates the impact of enforcing equivariance constraints on the expressive power. By examining the boundary hyperplanes and the channel vectors, we constructively demonstrate that enforcing equivariance constraints could undermine the expressive power. Naturally, this drawback can be compensated for by enlarging the model size -- we further prove upper bounds on the required enlargement for compensation. Surprisingly, we show that the enlarged neural architectures have reduced hypothesis space dimensionality, implying even better generalizability.

cs.LG

Collective dynamics of active suspensions on curved viscous interfaces

Self-propelled particles can navigate complex environments, including viscous fluid interfaces with curved geometries. In this work, we study the emergent dynamics of a suspension of self-propelled particles confined to a stationary curved viscous interface. The evolution of the particle configurations is modeled using the Fokker-Planck equation on the curved surface, formulated using Cartan's moving frame method, and coupled to the bulk and surface Stokes equations with flows driven by an interfacial nematic active stress. Specifically, for a spherical vesicle, the flow field and the distribution of the particles are analyzed theoretically and numerically within the framework of spin-weighted functions and spin-weighted spherical harmonics, which provide a natural geometric description of the probability distribution function on the sphere. A linear stability analysis about the uniform, isotropic state is performed and predicts a finite-wavelength instability, with mode selection arising from the competition between the vesicle radius and the Saffman-Delbrück length. This instability and the associated mode-selection mechanism are also confirmed in nonlinear numerical simulations using a pseudo-spectral method based on spin-weighted spherical harmonics.

physics.flu-dyn

Generalisation of RLHF under Reward Shift and Clipped KL Regularisation

Alignment and adaptation in large language models heavily rely on reinforcement learning from human feedback (RLHF); yet, theoretical understanding of its generalisability remains premature, especially when the learned reward could shift, and the KL control is estimated and clipped. To address this issue, we develop generalisation theory for RLHF that explicitly accounts for (1) \emph{reward shift}: reward models are trained on preference data from earlier or mixed behaviour policies while RLHF optimises the current policy on its own rollouts; and (2) \emph{clipped KL regularisation}: the KL regulariser is estimated from sampled log-probability ratios and then clipped for stabilisation, resulting in an error to RLHF. We present generalisation bounds for RLHF, suggesting that the generalisation error stems from a sampling error from prompts and rollouts, a reward shift error, and a KL clipping error. We also discuss special cases of (1) initialising RLHF parameters with a uniform prior over a finite space, and (2) training RLHF by stochastic gradient descent, as an Ornstein-Uhlenbeck process. The theory yields practical implications in (1) optimal KL clipping threshold, and (2) budget allocation in prompts, rollouts, and preference data.

cs.LG

CoVeR: Conformal Calibration for Versatile and Reliable Autoregressive Next-Token Prediction

Autoregressive pre-trained models combined with decoding methods have achieved impressive performance on complex reasoning tasks. While mainstream decoding strategies such as beam search can generate plausible candidate sets, they often lack provable coverage guarantees, and struggle to effectively balance search efficiency with the need for versatile trajectories, particularly those involving long-tail sequences that are essential in certain real-world applications. To address these limitations, we propose \textsc{CoVeR}, a novel model-free decoding strategy wihtin the conformal prediction framework that simultaneously maintains a compact search space and ensures high coverage probability over desirable trajectories. Theoretically, we establish a PAC-style generalization bound, guaranteeing that \textsc{CoVeR} asymptotically achieves a coverage rate of at least $1 - α$ for any target level $α\in (0,1)$.

cs.LG

HRP: High-Rank Preheating for Superior LoRA Initialization

This paper studies the crucial impact of initialization in Low-Rank Adaptation (LoRA). Through theoretical analysis, we demonstrate that the fine-tuned result of LoRA is highly sensitive to initialization, which is likely to lead suboptimal low-rank results. While this issue can be mitigated by adjusting the initial direction towards the main singular vectors of the target $ΔW$, which is, however, typically unknown in real-world scenarios. To approximate this initial direction, we propose High-Rank Preheating (HRP), which first trains LoRA with a higher preheating rank for a few steps, then uses the main singular vectors of the derived $BA^\top$ as initialization for the main fine-tuning process. With only a modification in the initial direction, we prove that HRP makes LoRA achieve better fine-tuned results than random initialization in expectation, and the enhancement grows with the preheating rank. We validate our theoretical findings through extensive experiments in various models and tasks, where HRP significantly enhances LoRA's effectiveness and outperforms other initialization strategies and other LoRA variants.

cs.LG

A Theoretical Perspective: How to Prevent Model Collapse in Self-consuming Training Loops

High-quality data is essential for training large generative models, yet the vast reservoir of real data available online has become nearly depleted. Consequently, models increasingly generate their own data for further training, forming Self-consuming Training Loops (STLs). However, the empirical results have been strikingly inconsistent: some models degrade or even collapse, while others successfully avoid these failures, leaving a significant gap in theoretical understanding to explain this discrepancy. This paper introduces the intriguing notion of recursive stability and presents the first theoretical generalization analysis, revealing how both model architecture and the proportion between real and synthetic data influence the success of STLs. We further extend this analysis to transformers in in-context learning, showing that even a constant-sized proportion of real data ensures convergence, while also providing insights into optimal synthetic data sizing.

cs.LG

A Theoretical Survey on Foundation Models

Understanding the inner mechanisms of black-box foundation models (FMs) is essential yet challenging in artificial intelligence and its applications. Over the last decade, the long-running focus has been on their explainability, leading to the development of post-hoc explainable methods to rationalize the specific decisions already made by black-box FMs. However, these explainable methods have certain limitations in terms of faithfulness and resource requirement. Consequently, a new class of interpretable methods should be considered to unveil the underlying mechanisms of FMs in an accurate, comprehensive, heuristic, and resource-light way. This survey aims to review those interpretable methods that comply with the aforementioned principles and have been successfully applied to FMs. These methods are deeply rooted in machine learning theory, covering the analysis of generalization performance, expressive capability, and dynamic behavior. They provide a thorough interpretation of the entire workflow of FMs, ranging from the inference capability and training dynamics to their ethical implications. Ultimately, drawing upon these interpretations, this review identifies the next frontier research directions for FMs.

cs.LG

Segmented Model-Based Hydrogen Delivery Control for PEM Fuel Cells: a Port-Hamiltonian Approach

This paper proposes an extended interconnection and damping assignment passivity-based control technique (IDA-PBC) to control the pressure dynamics in the fuel delivery subsystem (FDS) of proton exchange membrane fuel cells. The fuel cell stack is a distributed parameter model which can be modeled by partial differential equations PDEs). In this paper, the segmentation concept is used to approximate the PDEs model by ordinary differential equations (ODEs) model. Therefore, each segments are having multiple ODEs to obtain the lump-sum model of the segments. Subsequently, a generalized multi-input multi-output lumped parameters model is developed in port-Hamiltonian framework based on mass balance to minimize the modeling error. The modeling errors arises due to the difference between spatially distributed pressures in FDS segments, and also due to the difference between the actual stack pressure and the measured output pressure of the anode. The segments interconnection feasibilities are ensured by maintaining passivity of each segment. With consideration of re-circulation and bleeding of the anode in the modeling, an extended energy-shaping and output tracking IDA-PBC based state-feedback controller is proposed to control the spatially distributed pressure dynamics in the anode. Furthermore, a sliding mode observer of high order is designed to estimate the unmeasurable pressures in FDS with known disturbances. Performance recovery of output feedback control is accomplished with explicit stability analysis. The effectiveness of the proposed IDA-PBC approach is validated by the simulation results.

eess.SY

Cell motility modes are selected by the interplay of mechanosensitive adhesion and membrane tension

The initiation of directional cell motion requires symmetry breaking that can happen both with or without external stimuli. During cell crawling, forces generated by the cytoskeleton and their transmission through mechanosensitive adhesions to the extracellular substrate play a crucial role. In a recently proposed 1D model (Sens, PNAS 2020), a mechanical feedback loop between force-sensitive adhesions and cell tension was shown to be sufficient to explain spontaneous symmetry breaking and multiple motility patterns through stick-slip dynamics, without the need to account for signaling networks or active polar gels. We extended this model to 2D to study the interplay between cell shape and mechanics during crawling. Through a local force balance along a deformable boundary, we show that the membrane tension coupled with shape change can regulate the spatiotemporal evolution of the stochastic binding of mechanosensitive adhesions. Linear stability analysis identified the unstable parameter regimes where spontaneous symmetry breaking can take place. sing simulations to solve the fully coupled nonlinear system of equations, we show that starting from a randomly perturbed circular shape, this instability can lead to keratocyte-like shapes. Simulations predict that different adhesion kinetics and membrane tension can result in different cell motility modes including gliding, zigzag, rotating, and sometimes chaotic movements. Thus, using a minimal model of cell motility, we identify that the interplay between adhesions and tension can select emergent motility modes.

q-bio.CB

Identities of the Kauffman Monoid $\mathcal{K}_3$

We give a transparent combinatorial characterization of the identities satisfied by the Kauffman monoid $\mathcal{K}_3$. Our characterization leads to a polynomial time algorithm to check whether a given identity holds in $\mathcal{K}_3$.

math.GR

The finite basis problem for the monoid of 2 by 2 upper triangular tropical matrices

For each positive $n$, let $u_n = v_n$ denote the identity obtained from the Adjan identity $(xy) (yx) (xy) (xy) (yx) = (xy) (yx) (yx) (xy) (yx)$ by substituting $(xy) \rightarrow (x_1 x_2 \dots x_n)$ and $(yx) \rightarrow (x_n \dots x_2 x_1)$. We show that every monoid which satisfies $u_n = v_n$ for each positive $n$ and generates the variety containing the bicyclic monoid is nonfinitely based. This implies that the monoid of 2 by 2 upper triangular tropical matrices over the tropical semiring is nonfinitely based.

math.GR

The Finite Basis Problem for Kauffman Monoids

We prove a sufficient condition under which a semigroup admits no finite identity basis. As an application, it is shown that the identities of the Kauffman monoid $\mathcal{K}_n$ are nonfinitely based for each $n\ge 3$. This result holds also for the case when $\mathcal{K}_n$ is considered as an involution semigroup under either of its natural involutions.

math.GR