SearcharxivSearch

arXiv subjects

Lang Yu

Publications and source records attributed to Lang Yu.

At least 19 recordsLinked to original sources

Sparse Recovery via $\ell_1^2-\eta\ell_2^2$ Minimization

The weighted difference of squared norms (WDSN) penalty $\ell_1^2-\eta\ell_2^2$ with $0\leq \eta\leq 1$ has attracted considerable attention due to its strong sparsity-promoting ability and favorable reconstruction performance in compressed sensing and inverse problems. However, exact recovery guarantees and restricted isometry property (RIP) analysis for WDSN minimization have not yet been established. In this paper, we address this gap. First, we establish sufficient conditions for the exact recovery of $k$-sparse signals based on the null space property (NSP). Then, under the $\delta_{2k}$-RIP condition, we derive stable recovery guarantees for both $k$-sparse signals and general signals, and characterize upper bounds on the reconstruction error. Furthermore, we propose a WDSN-based regularized model to handle both noiseless and noisy observations in a unified framework. To design an efficient algorithm, we derive an explicit formula for the proximal operator of the WDSN functional. Based on this proximal solver, we develop a suitable variable-splitting scheme within the alternating direction method of multipliers (ADMM) and establish its global convergence under some mild conditions. Finally, numerical experiments show that the proposed method outperforms the iterative half variation method in both noiseless and noisy sparse recovery tasks.

math.OC

Sparse Recovery via $\ell_p^p/\ell_q^p$ Ratio Minimization: Theory and Algorithm

The constrained $\ell_p^p/\ell_q^p$ ratio model is scale invariant and is therefore attractive for sparse signal recovery. However, its nonconvex, nonsmooth, and fractional structure makes a unified theoretical and algorithmic analysis challenging for $0 1$. This paper develops a unified framework for this general model, covering deterministic exact recovery, stable recovery for sparse and compressible signals, and convergence analysis of a fractional algorithm. We first establish two deterministic sufficient conditions for exact recovery: a local optimality criterion and a null-space condition ensuring uniform recovery. For the $\ell_1/\ell_q$ subfamily, this null-space condition is further converted into high-probability sample-complexity bounds for isotropic sub-Gaussian matrix. We then study noisy recovery. Under the $k$-sparsity assumption, we improve the RIP-based stable recovery theory by relaxing the required sufficient condition and deriving sharper reconstruction-error bounds. For compressible signals, we establish RIP--ROP based error estimates whose constants are independent of the ambient dimension, improving prior bounds with explicit dimension-dependent factors [1]. An RIP-only variant is also derived. On the algorithmic side, we propose a prox-linear Dinkelbach framework that directly handles the fractional structure of the constrained problem and prove its convergence. Numerical experiments demonstrate that suitable choices of $(p,q)$ are effective for high-dynamic-range sparse signals and coherent sensing matrices.

math.OC

AblateCell: A Reproduce-then-Ablate Agent for Virtual Cell Repositories

Systematic ablations are essential to attribute performance gains in AI Virtual Cells, yet they are rarely performed because biological repositories are under-standardized and tightly coupled to domain-specific data and formats. While recent coding agents can translate ideas into implementations, they typically stop at producing code and lack a verifier that can reproduce strong baselines and rigorously test which components truly matter. We introduce AblateCell, a reproduce-then-ablate agent for virtual cell repositories that closes this verification gap. AblateCell first reproduces reported baselines end-to-end by auto-configuring environments, resolving dependency and data issues, and rerunning official evaluations while emitting verifiable artifacts. It then conducts closed-loop ablation by generating a graph of isolated repository mutations and adaptively selecting experiments under a reward that trades off performance impact and execution cost. Evaluated on three single-cell perturbation prediction repositories (CPA, GEARS, BioLORD), AblateCell achieves 88.9% (+29.9% to human expert) end-to-end workflow success and 93.3% (+53.3% to heuristic) accuracy in recovering ground-truth critical components. These results enable scalable, repository-grounded verification and attribution directly on biological codebases.

cs.AI

SCALE:Scalable Conditional Atlas-Level Endpoint transport for virtual cell perturbation prediction

Virtual-cell models aim to predict how cell populations respond to perturbations, but control and treated cells are measured as unpaired populations, complicating the learning of perturbation-specific effects. We present SCALE, a conditional transport model that represents cells as unordered sets and predicts treated populations without cell-level matching. A shared set-aware encoder and conditional DiT backbone learn latent transport, making endpoint supervision directly delta-aligned without an auxiliary delta objective. Across genetic, chemical, developmental and immune perturbations, SCALE recovered gene-expression changes, response directions and population structure. In CRISPR data with dominant cell-line effects, SCALE outperformed competing methods across seven metrics and maintained separation among gene-target representations rather than collapsing them into a shared region. SCALE further prioritized cytokines predicted to produce distinct immune activation and inflammatory responses. Experiments using matched PBMC samples from three donors confirmed these predicted differences. Together, SCALE enables perturbation-specific prediction from unpaired populations and supports experimental prioritization.

cs.LG

HarmonyCell: Automating Single-Cell Perturbation Modeling under Semantic and Distribution Shifts

Single-cell perturbation studies face dual heterogeneity bottlenecks: (i) semantic heterogeneity--identical biological concepts encoded under incompatible metadata schemas across datasets; and (ii) statistical heterogeneity--distribution shifts from biological variation demanding dataset-specific inductive biases. We propose HarmonyCell, an end-to-end agent framework resolving each challenge through a dedicated mechanism: an LLM-driven Semantic Unifier autonomously maps disparate metadata into a canonical interface without manual intervention; and an adaptive Monte Carlo Tree Search engine operates over a hierarchical action space to synthesize architectures with optimal statistical inductive biases for distribution shifts. Evaluated across diverse perturbation tasks under both semantic and distribution shifts, HarmonyCell achieves a 95% valid execution rate on heterogeneous input datasets (versus 0% for general agents) while matching or even exceeding expert-designed baselines in rigorous out-of-distribution evaluations. This dual-track orchestration enables scalable automatic virtual cell modeling without dataset-specific engineering.

cs.AI

Difference-of-Convex Elastic Net for Compressed Sensing

This work proposes a novel and unified sparse recovery framework, termed the difference of convex Elastic Net (DCEN). This framework effectively balances strong sparsity promotion with solution stability, and is particularly suitable for high-dimensional variable selection involving highly correlated features. Built upon a difference-of-convex (DC) structure, DCEN employs two continuously tunable parameters to unify classical and state-of-the-art models--including LASSO, Elastic Net, Ridge, and $\ell_1-\alpha\ell_2$--as special cases. Theoretically, sufficient conditions for exact and stable recovery are established under the restricted isometry property (RIP), an oracle inequality and recovery bound are derived for the global solution, and a closed-form expression of the DCEN regularization proximal operator is obtained. Moreover, two efficient optimization algorithms are developed based on the DC algorithm (DCA) and the alternating direction method of multipliers (ADMM). Within the Kurdyka-\L{}ojasiewicz (K\L{}) framework, the global convergence of DCA and its linear convergence rate are rigorously established. Furthermore, DCEN is extended to image reconstruction by incorporating total variation (TV) regularization, yielding the DCEN-TV model, which is efficiently solved via the Split Bregman method. Numerical experiments demonstrate that DCEN consistently outperforms state-of-the-art methods in sparse signal recovery, high-dimensional variable selection under strong collinearity, and Magnetic Resonance Imaging (MRI) image reconstruction, achieving superior recovery accuracy and robustness.

math.OC

An Efficient ADMM Method for Ratio-Type Nonconvex and Nonsmooth Minimization in Sparse Recovery

Sparse signal recovery based on nonconvex and nonsmooth optimization problems has significant applications and demonstrates superior performance in signal processing and machine learning. This work deals with a scale-invariant $\ell_{1/2}/\ell_{2}$ sparse minimization with nonconvex, nonseparable, ratio-type regularization to enhance the accuracy and stability of sparse recovery. Within the framework of the null space property, we analyze the conditions for exact and stable recovery in constrained minimization problem. For the unconstrained regularized minimization problem, we develop an alternating direction method of multipliers (ADMM) based on a splitting strategy and rigorously analyze its global convergence and linear convergence rate under reasonable assumptions. Numerical experiments demonstrate that the proposed method consistently outperforms existing approaches across diverse noise levels and measurement settings. Furthermore, experiments on neural network sparsity and generalization performance demonstrate that the method effectively improves prediction accuracy.

math.OC

Scaling functions in the soft-wall AdS/QCD models

We investigate the static scaling behavior of the chiral condensate near the two-flavor critical point within the framework of the soft-wall AdS/QCD. The scaling functions are extracted from the chiral order parameters and are found to precisely match those obtained through mean-field calculations. Additionally, it is also checked that the scaling functions are independent of the specific construction of the holographic model. Furthermore, we develop the formalism for calculating the chiral susceptibility and demonstrate that the pseudo-critical temperatures obey the scaling law for moderate quark masses. It is shown that the temperature scaling could be comparable with those obtained from Dyson-Schwinger equations and lattice simulations. These findings could help improve the effectiveness of the soft-wall AdS/QCD.

hep-ph

Protein-SE(3): Benchmarking SE(3)-based Generative Models for Protein Structure Design

SE(3)-based generative models have shown great promise in protein geometry modeling and effective structure design. However, the field currently lacks a modularized benchmark to enable comprehensive investigation and fair comparison of different methods. In this paper, we propose Protein-SE(3), a new benchmark based on a unified training framework, which comprises protein scaffolding tasks, integrated generative models, high-level mathematical abstraction, and diverse evaluation metrics. Recent advanced generative models designed for protein scaffolding, from multiple perspectives like DDPM (Genie1 and Genie2), Score Matching (FrameDiff and RfDiffusion) and Flow Matching (FoldFlow and FrameFlow) are integrated into our framework. All integrated methods are fairly investigated with the same training dataset and evaluation metrics. Furthermore, we provide a high-level abstraction of the mathematical foundations behind the generative models, enabling fast prototyping of future algorithms without reliance on explicit protein structures. Accordingly, we release the first comprehensive benchmark built upon unified training framework for SE(3)-based protein structure design, which is publicly accessible at https://github.com/BruthYU/protein-se3.

cs.LG

Optimizing Question Semantic Space for Dynamic Retrieval-Augmented Multi-hop Question Answering

Retrieval-augmented generation (RAG) is usually integrated into large language models (LLMs) to mitigate hallucinations and knowledge obsolescence. Whereas,conventional one-step retrieve-and-read methods are insufficient for multi-hop question answering, facing challenges of retrieval semantic mismatching and the high cost in handling interdependent subquestions. In this paper, we propose Optimizing Question Semantic Space for Dynamic Retrieval-Augmented Multi-hop Question Answering (Q-DREAM). Q-DREAM consists of three key modules: (1) the Question Decomposition Module (QDM), which decomposes multi-hop questions into fine-grained subquestions; (2) the Subquestion Dependency Optimizer Module (SDOM), which models the interdependent relations of subquestions for better understanding; and (3) the Dynamic Passage Retrieval Module (DPRM), which aligns subquestions with relevant passages by optimizing the semantic embeddings. Experimental results across various benchmarks demonstrate that Q-DREAM significantly outperforms existing RAG methods, achieving state-of-the-art performance in both in-domain and out-of-domain settings. Notably, Q-DREAM also improves retrieval efficiency while maintaining high accuracy compared with recent baselines.

cs.IR

MELO: Enhancing Model Editing with Neuron-Indexed Dynamic LoRA

Large language models (LLMs) have shown great success in various Natural Language Processing (NLP) tasks, whist they still need updates after deployment to fix errors or keep pace with the changing knowledge in the world. Researchers formulate such problem as Model Editing and have developed various editors focusing on different axes of editing properties. However, current editors can hardly support all properties and rely on heavy computational resources. In this paper, we propose a plug-in Model Editing method based on neuron-indexed dynamic LoRA (MELO), which alters the behavior of language models by dynamically activating certain LoRA blocks according to the index built in an inner vector database. Our method satisfies various editing properties with high efficiency and can be easily integrated into multiple LLM backbones. Experimental results show that our proposed MELO achieves state-of-the-art editing performance on three sequential editing tasks (document classification, question answering and hallucination correction), while requires the least trainable parameters and computational cost.

cs.CL

Exploring the chiral and deconfinement phase transitions in a self-consistent PNJL model

In this work, we study the chiral and deconfinement phase transitions in a two-flavor Polyakov loop extended Nambu--Jona-Lasinio (PNJL) model. And note that the self-consistent mean field approximation is employed by introducing an arbitrary parameter $\alpha$ to measure the weights of the Fierz-transformed interaction channels. By making use of this model, we systematically investigate the chiral and deconfinement phase transition lines (as well as the chiral ones in the NJL model for comparison) under different values of $\alpha$. It is found that, the increasing of $\alpha$ helps to enhance the chiral (pseudo)critical temperature at fixed chemical potential, and also to enhance the chiral (pseudo)critical chemical potential at fixed temperature. And the critical end point (CEP) vanishes when $\alpha$ becomes large enough. Besides, we find that the incorporation of Polyakov loop increases $T_{CEP}$ but does not change $\mu_{CEP}$ for small values of $\alpha$.

nucl-th

Counterfactual reasoning: Testing language models' understanding of hypothetical scenarios

Current pre-trained language models have enabled remarkable improvements in downstream tasks, but it remains difficult to distinguish effects of statistical correlation from more systematic logical reasoning grounded on the understanding of real world. We tease these factors apart by leveraging counterfactual conditionals, which force language models to predict unusual consequences based on hypothetical propositions. We introduce a set of tests from psycholinguistic experiments, as well as larger-scale controlled datasets, to probe counterfactual predictions from five pre-trained language models. We find that models are consistently able to override real-world knowledge in counterfactual scenarios, and that this effect is more robust in case of stronger baseline world knowledge -- however, we also find that for most models this effect appears largely to be driven by simple lexical cues. When we mitigate effects of both world knowledge and lexical cues to test knowledge of linguistic nuances of counterfactuals, we find that only GPT-3 shows sensitivity to these nuances, though this sensitivity is also non-trivially impacted by lexical associative factors.

cs.CL

Counterfactual reasoning: Do language models need world knowledge for causal understanding?

Current pre-trained language models have enabled remarkable improvements in downstream tasks, but it remains difficult to distinguish effects of statistical correlation from more systematic logical reasoning grounded on understanding of the real world. In this paper we tease these factors apart by leveraging counterfactual conditionals, which force language models to predict unusual consequences based on hypothetical propositions. We introduce a set of tests drawn from psycholinguistic experiments, as well as larger-scale controlled datasets, to probe counterfactual predictions from a variety of popular pre-trained language models. We find that models are consistently able to override real-world knowledge in counterfactual scenarios, and that this effect is more robust in case of stronger baseline world knowledge -- however, we also find that for most models this effect appears largely to be driven by simple lexical cues. When we mitigate effects of both world knowledge and lexical cues to test knowledge of linguistic nuances of counterfactuals, we find that only GPT-3 shows sensitivity to these nuances, though this sensitivity is also non-trivially impacted by lexical associative factors.

cs.CL

"No, they did not": Dialogue response dynamics in pre-trained language models

A critical component of competence in language is being able to identify relevant components of an utterance and reply appropriately. In this paper we examine the extent of such dialogue response sensitivity in pre-trained language models, conducting a series of experiments with a particular focus on sensitivity to dynamics involving phenomena of at-issueness and ellipsis. We find that models show clear sensitivity to a distinctive role of embedded clauses, and a general preference for responses that target main clause content of prior utterances. However, the results indicate mixed and generally weak trends with respect to capturing the full range of dynamics involved in targeting at-issue versus not-at-issue content. Additionally, models show fundamental limitations in grasp of the dynamics governing ellipsis, and response selections show clear interference from superficial factors that outweigh the influence of principled discourse constraints.

cs.CL

Impacts of (inverse) magnetic catalysis on screening masses of neutral pions and sigma mesons in hot and magnetized quark matter

We investigate the screening masses of neutral pions and sigma mesons in hot and magnetized quark matter in the framework of a two-flavor lattice-improved Nambu-Jona-Lasinio (NJL) model with a magnetic field dependent coupling constant, which is determined by utilizing the results from lattice QCD simulations. Since such model can well reproduce inverse magnetic catalysis (IMC), by comparing with the standard NJL model, we systemically analyze the impacts of IMC on the temperature and magnetic field dependences of the longitudinal and transverse screening masses of the chiral partners, i.e. {\pi}^0 and {\sigma} mesons, as well as the screening mass differences between them. Particularly, it is found that the eB dependences of two alternative (pseudo)critical temperatures for the chiral transition defined by {\sigma}-{\pi}^0 meson screening mass differences are consistent with that defined by the quark condensate.

hep-ph

On the Interplay Between Fine-tuning and Composition in Transformers

Pre-trained transformer language models have shown remarkable performance on a variety of NLP tasks. However, recent research has suggested that phrase-level representations in these models reflect heavy influences of lexical content, but lack evidence of sophisticated, compositional phrase information. Here we investigate the impact of fine-tuning on the capacity of contextualized embeddings to capture phrase meaning information beyond lexical content. Specifically, we fine-tune models on an adversarial paraphrase classification task with high lexical overlap, and on a sentiment classification task. After fine-tuning, we analyze phrasal representations in controlled settings following prior work. We find that fine-tuning largely fails to benefit compositionality in these representations, though training on sentiment yields a small, localized benefit for certain models. In follow-up analyses, we identify confounding cues in the paraphrase dataset that may explain the lack of composition benefits from that task, and we discuss potential factors underlying the localized benefits from sentiment training.

cs.CL

Qubitization of Bosons

A binary mapping from Fock space of bosonic state to qubits is given. Based on the binary mapping, we construte an algorithm of qubitization of bosons with complexity O(log(N)). As an example, the algorithm of qubitization of bosons in matrix product state to simulate real time dynamics of Yukawa coupling is realized. The calculation error bar is estimated by random sampling method. This proposal may be achieved in superconductivity noisy intermediate--scale quantum computer not far future.

quant-ph