SearcharxivSearch

arXiv subjects

Fei Ye

Publications and source records attributed to Fei Ye.

At least 19 recordsLinked to original sources

Robust Dynamic Expansion for Continual Learning under Backdoor Attacks via Purification and Selective Recovery

Continual learning (CL) enables models to acquire new knowledge from sequentially arriving tasks while retaining previously learned knowledge. However, in practical scenarios, task streams collected from untrusted sources may contain backdoor-poisoned samples, posing a critical challenge to the stability, plasticity, and security of continual learners. In this work, we investigate a challenging setting termed Continual Learning Under Backdoor Attack (CLUBA), where each incremental task may involve a small proportion of maliciously manipulated training samples. Unlike conventional continual learning or backdoor defense scenarios, CLUBA requires models to simultaneously mitigate catastrophic forgetting, preserve adaptation capability, and prevent the absorption of malicious supervision during sequential updates. To address this challenge, we propose a robust dynamic-expansion framework that integrates sample purification, selective recovery, and robust expert routing into a unified continual learning paradigm. Specifically, we introduce Bi-Prototype Purification (BPP) to identify suspicious samples by exploiting semantic discrepancies in feature space. Based on purified data, Gradient Discrepancy-based Robustness Optimization (GDBRO) selectively recovers informative poisoned samples through pseudo-label correction and gradient consistency evaluation, improving robustness while maintaining model plasticity. Furthermore, Robust Feature Consistency-based Expert Selection (RFCBES) constructs perturbation-aware class prototypes to enable reliable expert routing under corrupted or shifted inputs.

cs.LG

Singlet-doublet transitions and Josephson currents in a superconducting ring with a quantum dot

We investigate the ground state properties of a superconducting ring embedded with a quantum dot (QD) by using a variational wave-function approach. A theoretical formulation for the treatment of the finite-U Anderson impurity coupled with a superconducting ring are presented. We demonstrate singlet-doublet transitions of the ground state for this system with the QD in the mixed valence regime. It is shown that the supercurrent in the superconductor ring shows oscillations with the external enclosed magnetic flux and exhibits abrupt jumps at the singlet-doublet phase transition points.

cond-mat.mes-hall

On the Chern-Ricci form of a twisted almost K\"{a}hler structure

Let $(M,g,J,\omega)$ be an almost K\"{a}hler manifold. For any smooth function $f$ on $M$, one can associate an automorphism $\psi\in \mbox{Aut}(TM)$ for which the K\"{a}hler form is invariant. Using $\psi$, one can ``twist" the metric $g$ and almost complex structure $J$ to obtain a new almost K\"{a}hler structure $(g^\psi,J^\psi,\omega)$ on $M$. Let $\widetilde{D}$ denote the Chern connection of $(g^\psi,J^\psi,\omega)$ and let $K^{-1}$ denote the anti-canonical bundle of $(TM,J^\psi)$. In the current paper, we give an explicit formula for the local connection 1-form $\alpha$ associated to the pair $(K^{-1},\widetilde{D})$. The Chern-Ricci form of $(g^\psi,J^\psi,\omega)$ is then $\rho_{\widetilde{D}}=-d\alpha$. We note that under certain conditions the aforementioned formula assumes a simpler form when applied to the calculation of $\alpha$. We illustrate this with some examples.

math.DG

Berezinskii-Kosterlitz-Thouless Quantum Supercriticality in XXZ Heisenberg Spin Chain

Quantum fluctuations can give rise to a singular quantum critical point (QCP) in the ground state, whose influence extends to finite temperatures, forming a quantum critical regime (QCR). Recently, it has been shown that in the quantum Ising model, the symmetry-breaking, longitudinal field can induce a quantum supercritical regime (QSR) emanating from the QCP, which hosts a universally enhanced quantum supercritical magnetocaloric effect (MCE). In this paper, we show that the QSR also emerges in the spin-1/2 XXZ model, in both the form of Ising and Berezinskii-Kosterlitz-Thouless (BKT) supercriticality. Using ground-state and finite-temperature tensor-network methods, we investigate quantum supercritical phenomena near a BKT QCP. We reveal a quantum supercritical crossover scaling $T \propto h^{2/3}$ and a Gr\"uneisen ratio scaling $\Gamma_h \propto T^{-3/2}$ for the BKT QCP, which differ from the corresponding Ising supercritical scalings. Nevertheless, we find that the scaling function $\phi_{\Gamma}(x)$ of the singular Gr\"uneisen ratio for both BKT and Ising cases can be approximately described by the same expression $\phi_{\Gamma}(x) \approx x/(1+x^2)$. Our work extends the study of quantum supercritical phenomena from the Ising to the XXZ Heisenberg model, thereby revealing the presence of BKT quantum supercriticality and broadening the scope of quantum supercritical physics.

cond-mat.str-el

Non-volatile Multistate Magnetic Switching via Spin-orbit Torque and Intrinsic Anisotropy

While current-induced bistate spin-orbit torque (SOT) switching has been well established, deterministic electrical control of multiple magnetic states remains a central challenge in spintronics. Here, we realize a conceptually new multistate SOT device in a SrIrO_3/SrRuO_3 bilayer, hosting four intrinsically stable yet electrically distinguishable magnetic states, including two in-plane canted (IP_c^$\pm$) and two out-of-plane canted (OP_c^$\pm$) states. Pulsed current excitations fully map all twelve deterministic transitions among the four states, establishing a robust switching protocol defined by two characteristic current densities. In-situ scanning nitrogen-vacancy (NV) center magnetometry provides direct real-space evidence for the previously unobserved IP_c^$\pm$ states, and spin dynamics simulations uncover a two-step switching pathway, driven by the concerted action of spin torques and the effective anisotropy field within the fourfold anisotropy landscape. Our demonstration of the intrinsic multistate SOT device directly addresses the density bottleneck of conventional bistate SOT technology, establishing a powerful paradigm for compact, high-speed, and energy-efficient multistate spintronics.

cond-mat.mes-hall

On Vanishing Theorems and Bogomolov's Inequality on Surfaces in Positive Characteristic

In this paper, we study the equivalence between Bogomolov's instability theorem and the Miyaoka-Sakai theorem on surfaces in positive characteristic. We show that Bogomolov's instability theorem can be derived from Miyaoka-Sakai theorem. Conversely, it implies a partial version of the Miyaoka-Sakai theorem that lacks the vanishing conclusion. This partial version is still sufficient to deduce the Mumford-Ramanujam vanishing theorem. Additionally, we identify a class of surfaces in positive characteristic for which the Miyaoka-Sakai theorem (or a weaker variant), or the Kawamata-Viehweg vanishing theorem holds. In particular, we present a new proof of the Kawamata-Viehweg vanishing theorem on smooth del Pezzo surfaces. As an application of the Miyaoka-Sakai theorem, we obtain Reider-type results concerning Fujita's conjecture.

math.AG

SeedProteo: Accurate De Novo All-Atom Design of Protein Binders

We present SeedProteo, a diffusion-based model for de novo all-atom protein design. We demonstrate how to repurpose a cutting-edge folding architecture into a powerful generative design framework by effectively integrating self-conditioning features. Extensive benchmarks highlight the model's capabilities across two distinct tasks: in unconditional generation, SeedProteo exhibits superior length generalization and structural diversity, maintaining robustness for long sequences and complex topologies; in binder design, it achieves state-of-the-art performance among open-source methods, attaining the highest in-silico design success rates, structural diversity and novelty. Finally, we validate SeedProteo through wet-lab assays on two therapeutic targets, achieving hit rates of 70%-80% and picomolar-level binding affinities, establishing leading results. To facilitate community adoption, we provide public access to SeedProteo via a webserver (https://seedfold.io/proteinDesign).

q-bio.BM

SeedFold: Scaling Biomolecular Structure Prediction

Highly accurate biomolecular structure prediction is a key component of developing biomolecular foundation models, and one of the most critical aspects of building foundation models is identifying the recipes for scaling the model. In this work, we present SeedFold, a folding model that successfully scales up the model capacity. Our contributions are threefold: first, we identify an effective width-scaling strategy for the Pairformer to increase representation capacity; second, we introduce a novel linear triangular attention that reduces computational complexity to enable efficient scaling; finally, we construct a large-scale distillation dataset to substantially enlarge the training set. Experiments on FoldBench show that SeedFold outperforms AlphaFold3 on most protein-related tasks.

q-bio.BM

AInsteinBench: Benchmarking Coding Agents on Scientific Repositories

We introduce AInsteinBench, a large-scale benchmark for evaluating whether large language model (LLM) agents can operate as scientific computing development agents within real research software ecosystems. Unlike existing scientific reasoning benchmarks which focus on conceptual knowledge, or software engineering benchmarks that emphasize generic feature implementation and issue resolving, AInsteinBench evaluates models in end-to-end scientific development settings grounded in production-grade scientific repositories. The benchmark consists of tasks derived from maintainer-authored pull requests across six widely used scientific codebases, spanning quantum chemistry, quantum computing, molecular dynamics, numerical relativity, fluid dynamics, and cheminformatics. All benchmark tasks are carefully curated through multi-stage filtering and expert review to ensure scientific challenge, adequate test coverage, and well-calibrated difficulty. By leveraging evaluation in executable environments, scientifically meaningful failure modes, and test-driven verification, AInsteinBench measures a model's ability to move beyond surface-level code generation toward the core competencies required for computational scientific research.

cs.SE

Interplay of spin-orbit coupling and trigonal crystal field enhances superconductivity in $LaAlO_3/KTaO_3$ (111)

In conventional superconductors, bulk physical properties typically degrade as the film thickness approaches the two-dimensional (2D) limit. Here in the (111) oriented LaAlO3/KTaO3 (LAO/KTO) heterostructure, we demonstrate experimental evidence that reducing the conducting layer thickness at the interface significantly enhances superconducting transition temperature Tc, in direct contrast to conventional wisdom. From the sum frequency generation (SFG) spectroscopy and superconducting upper-critical field measurements, both the trigonal symmetry and spin orbit scattering are enhanced with the increased Tc. We attribute the enhanced superconductivity (SC) to the synergic interplay between spin-orbit coupling (SOC) and trigonal crystal field, resulting in an enhanced electron-phonon coupling. Furthermore, we show the existence of unconventional SC: the approaching linear temperature dependence of normal state resistance with increasing Tc and the existence of a quantum critical point (QCP) near the superconducting phase. Our findings provide important insight into the underlying mechanism of the strong orientation-dependent KTO interface SC.

cond-mat.supr-con

VeriEquivBench: An Equivalence Score for Ground-Truth-Free Evaluation of Formally Verifiable Code

Formal verification is the next frontier for ensuring the correctness of code generated by Large Language Models (LLMs). While methods that co-generate code and formal specifications in formal languages, like Dafny, can, in principle, prove alignment with user intent, progress is bottlenecked by specification quality evaluation. Current benchmarks rely on matching against ground-truth specifications, a manual and expertise-intensive process that has limited existing datasets to a few hundred simple problems and also suffers from a reliability issue. To address this, we introduce VeriEquivBench, a new benchmark with $2,389$ complex algorithmic problems that probe the limitations of current models in both code generation and formal reasoning. Our evaluation framework replaces ground-truth matching with a formally grounded metric, the equivalence score, and rigorously verifies the quality of generated specifications and code. Our results show that generating formally verifiable code remains a profound challenge for state-of-the-art LLMs. This underscores both the difficulty of the task and the need for benchmarks like VeriEquivBench to drive progress toward scalable and reliable coding agents.

cs.PL

MLego: Interactive and Scalable Topic Exploration Through Model Reuse

With massive texts on social media, users and analysts often rely on topic modeling techniques to quickly extract key themes and gain insights. Traditional topic modeling techniques, such as Latent Dirichlet Allocation (LDA), provide valuable insights but are computationally expensive, making them impractical for real-time data analysis. Although recent advances in distributed training and fast sampling methods have improved efficiency, real-time topic exploration remains a significant challenge. In this paper, we present MLego, an interactive query framework designed to support real-time topic modeling analysis by leveraging model materialization and reuse. Instead of retraining models from scratch, MLego efficiently merges materialized topic models to construct approximate results at interactive speeds. To further enhance efficiency, we introduce a hierarchical plan search strategy for single queries and an optimized query reordering technique for batch queries. We integrate MLego into a visual analytics prototype system, enabling users to explore large-scale textual datasets through interactive queries. Extensive experiments demonstrate that MLego significantly reduces computation costs while maintaining high-quality topic modeling results. MLego enhances existing visual analytics approaches, which primarily focus on user-driven topic modeling, by enabling real-time, query-driven exploration. This complements traditional methods and bridges the gap between scalable topic modeling and interactive data analysis.

cs.DB

Classification of autoimmune diseases from Peripheral blood TCR repertoires by multimodal multi-instance learning

T cell receptor (TCR) repertoires encode critical immunological signatures for autoimmune diseases, yet their clinical application remains limited by sequence sparsity and low witness rates. We developed EAMil, a multi-instance deep learning framework that leverages TCR sequencing data to diagnose systemic lupus erythematosus (SLE) and rheumatoid arthritis (RA) with exceptional accuracy. By integrating PrimeSeq feature extraction with ESMonehot encoding and enhanced gate attention mechanisms, our model achieved state-of-the-art performance with AUCs of 98.95% for SLE and 97.76% for RA. EAMil successfully identified disease-associated genes with over 90% concordance with established differential analyses and effectively distinguished disease-specific TCR genes. The model demonstrated robustness in classifying multiple disease categories, utilizing the SLEDAI score to stratify SLE patients by disease severity as well as to diagnose the site of damage in SLE patients, and effectively controlling for confounding factors such as age and gender. This interpretable framework for immune receptor analysis provides new insights for autoimmune disease detection and classification with broad potential clinical applications across immune-mediated conditions.

cs.LG

Dynamic Dual Buffer with Divide-and-Conquer Strategy for Online Continual Learning

Online Continual Learning (OCL) involves sequentially arriving data and is particularly challenged by catastrophic forgetting, which significantly impairs model performance. To address this issue, we introduce a novel framework, Online Dynamic Expandable Dual Memory (ODEDM), that integrates a short-term memory for fast memory and a long-term memory structured into sub-buffers anchored by cluster prototypes, enabling the storage of diverse and category-specific samples to mitigate forgetting. We propose a novel K-means-based strategy for prototype identification and an optimal transport-based mechanism to retain critical samples, prioritising those exhibiting high similarity to their corresponding prototypes. This design preserves semantically rich information. Additionally, we propose a Divide-and-Conquer (DAC) optimisation strategy that decomposes memory updates into subproblems, thereby reducing computational overhead. ODEDM functions as a plug-and-play module that can be seamlessly integrated with existing rehearsal-based approaches. Experimental results under both standard and imbalanced OCL settings show that ODEDM consistently achieves state-of-the-art performance across multiple datasets, delivering substantial improvements over the DER family as well as recent methods such as VR-MCL and POCL.

cs.LG

Elucidating the Design Space of Multimodal Protein Language Models

Multimodal protein language models (PLMs) integrate sequence and token-based structural information, serving as a powerful foundation for protein modeling, generation, and design. However, the reliance on tokenizing 3D structures into discrete tokens causes substantial loss of fidelity about fine-grained structural details and correlations. In this paper, we systematically elucidate the design space of multimodal PLMs to overcome their limitations. We identify tokenization loss and inaccurate structure token predictions by the PLMs as major bottlenecks. To address these, our proposed design space covers improved generative modeling, structure-aware architectures and representation learning, and data exploration. Our advancements approach finer-grained supervision, demonstrating that token-based multimodal PLMs can achieve robust structural modeling. The effective design methods dramatically improve the structure generation diversity, and notably, folding abilities of our 650M model by reducing the RMSD from 5.52 to 2.36 on PDB testset, even outperforming 3B baselines and on par with the specialized folding models. Project page and code: https://bytedance.github.io/dplm/dplm-2.1/.

cs.LG

Self-Controlled Dynamic Expansion Model for Continual Learning

Continual Learning (CL) epitomizes an advanced training paradigm wherein prior data samples remain inaccessible during the acquisition of new tasks. Numerous investigations have delved into leveraging a pre-trained Vision Transformer (ViT) to enhance model efficacy in continual learning. Nonetheless, these approaches typically utilize a singular, static backbone, which inadequately adapts to novel tasks, particularly when engaging with diverse data domains, due to a substantial number of inactive parameters. This paper addresses this limitation by introducing an innovative Self-Controlled Dynamic Expansion Model (SCDEM), which orchestrates multiple distinct trainable pre-trained ViT backbones to furnish diverse and semantically enriched representations. Specifically, by employing the multi-backbone architecture as a shared module, the proposed SCDEM dynamically generates a new expert with minimal parameters to accommodate a new task. A novel Collaborative Optimization Mechanism (COM) is introduced to synergistically optimize multiple backbones by harnessing prediction signals from historical experts, thereby facilitating new task learning without erasing previously acquired knowledge. Additionally, a novel Feature Distribution Consistency (FDC) approach is proposed to align semantic similarity between previously and currently learned representations through an optimal transport distance-based mechanism, effectively mitigating negative knowledge transfer effects. Furthermore, to alleviate over-regularization challenges, this paper presents a novel Dynamic Layer-Wise Feature Attention Mechanism (DLWFAM) to autonomously determine the penalization intensity on each trainable representation layer. An extensive series of experiments have been conducted to evaluate the proposed methodology's efficacy, with empirical results corroborating that the approach attains state-of-the-art performance.

cs.LG

Simulating Bulk Gap in Chiral Projected Entangled-Pair States

Projected entangled-pair states (PEPS) have proven effective in capturing chiral spin liquid ground states, yet the presence of long-range ``gossamer'' correlation tails raises concerns about their ability to accurately describe bulk gaps. Here, we address this challenge and demonstrate that PEPS can reliably characterize gapped bulk excitations in chiral topological phases. Using a variational principle for excited states within a local mode approximation, we establish that correlation functions decaying faster than $r^{-2}$ are not necessarily related to gapless modes and thus long-range ``gossamer'' correlation tails in chiral PEPS do not contradict the presence of a bulk gap. This framework is validated in the spin-$\frac{1}{2}$ Kitaev model with a chiral term, where PEPS yields excitation gaps that agree well with exact solutions. Extending our approach to the $\mathbb{Z}_3$ Kitaev model, we present compelling evidence for its chiral ground state and accurately resolve its gapped excitations. These findings thus solidify PEPS as a powerful tool for studying both ground and excited states in chiral topological systems, thereby bridging a key gap in the understanding of their bulk properties.

cond-mat.str-el

Learning Dynamic Representations via An Optimally-Weighted Maximum Mean Discrepancy Optimization Framework for Continual Learning

Continual learning has emerged as a pivotal area of research, primarily due to its advantageous characteristic that allows models to persistently acquire and retain information. However, catastrophic forgetting can severely impair model performance. In this study, we address network forgetting by introducing a novel framework termed Optimally-Weighted Maximum Mean Discrepancy (OWMMD), which imposes penalties on representation alterations via a Multi-Level Feature Matching Mechanism (MLFMM). Furthermore, we propose an Adaptive Regularization Optimization (ARO) strategy to refine the adaptive weight vectors, which autonomously assess the significance of each feature layer throughout the optimization process, The proposed ARO approach can relieve the over-regularization problem and promote the future task learning. We conduct a comprehensive series of experiments, benchmarking our proposed method against several established baselines. The empirical findings indicate that our approach achieves state-of-the-art performance.

cs.LG