SearcharxivSearch

arXiv subjects

Weiwei Wang

Publications and source records attributed to Weiwei Wang.

At least 19 recordsLinked to original sources

Divisive Normalization Shapes Low-Rank Slow Manifolds for Continuous Working Memory

The ability to robustly maintain and update continuous variables is a hallmark of working memory. While classical continuous attractor networks suffer from severe fine-tuning fragility, standard artificial recurrent neural networks (RNNs) like GRUs and LSTMs typically fail to stably learn continuous manifolds, instead shattering the state space into discretized point attractors. To bridge this gap, we draw inspiration from divisive normalization, a canonical neural computation widely observed across cortical circuits, and propose the Recurrent Divisive Normalization Network (RDNN), a minimal and algebraically isolated model of dynamic division. Through dynamical systems analysis on canonical working memory tasks, we demonstrate that this biophysical constraint allows the network to converge to robust, high-fidelity slow manifolds. Furthermore, we analyze the gradient dynamics of divisive normalization during Backpropagation Through Time (BPTT), showing that it introduces an activity-dependent local gradient scaling. This scaling dampens parameter updates in highly active regimes, which empirically aligns with a significant self-compression of the network's effective rank, confining the recurrent dynamics to a tight, low-dimensional subspace while avoiding the optimization pathologies associated with explicit low-rank factorization. Finally, ablations demonstrate that while subtractive inhibition can maintain static memories, divisive normalization is mathematically essential to prevent manifold shattering under time-varying inputs. Our findings identify divisive normalization not merely as a biological artifact, but as a critical computational mechanism for learning high-fidelity continuous representations.

q-bio.NC

Bochner-Yano Type Theorems for Conformal Killing Vector Fields under Curvature Pinching Conditions

In this article, we investigate conformal Killing vector fields on closed Riemannian manifolds under a curvature pinching condition. By establishing a new Bochner-type identity for the $1$-form dual to a conformal Killing vector field, we derive a sharp gradient estimate via a Moser iteration procedure. Based on this estimate, we prove that, under a suitable upper bound on the Ricci curvature, every nontrivial conformal Killing vector field must be nowhere vanishing. Consequently, on even-dimensional manifolds with non-zero Euler characteristic, every conformal Killing vector field vanishes identically, which in turn implies that the conformal transformation group of such a manifold is finite. Our results extend the classical rigidity theorems of Yano and Bochner from the setting of non-positive Ricci curvature to that of small positive Ricci curvature, and generalize the recent results of Chen and Han from Killing vector fields to conformal Killing vector fields.

math.DG

Ferrimagnetic Skyrmions in a Tetragonal Mn1.9Co0.1Sb Single Crystal at Room Temperature

The development of room temperature small-sized ferrimagnetic skyrmion materials is significant for topological spintronic device applications. As a room temperature ferrimagnetic material, the tetragonal Mn1.9Co0.1Sb crystal exhibits multiple phase transitions, including spin reorientation transitions. However, the magnetic spin textures and their evolution mechanisms during magnetic phase transitions in Mn1.9Co0.1Sb crystals remain unexplored. Using Lorentz transmission electron microscopy, we discovered and verified dipolar skyrmion behavior and its magnetic evolution at room temperature. We established a stable phase diagram of magnetic textures as functions of temperature and magnetic field, while also investigating the evolution mechanisms of spin textures across multiple temperature-induced magnetic phase transitions. Through micromagnetic simulations, a ferrimagnetic configuration with in-plane ferromagnetic coupling and interlayer antiferromagnetic arrangement was established, which stands in contrast to synthetic ferrimagnetic/antiferromagnetic systems that exhibit interlayer antiferromagnetic coupling via the Ruderman-Kittel-Kasuya-Yosida (RKKY) interaction. We determined that the intrinsic frequency of ferrimagnetic skyrmions can reach the THz regime due to strong interlayer antiparallel exchange interactions. These findings highlight the diversity of room temperature ferrimagnetic skyrmion regulation behaviors in Mn1.9Co0.1Sb and their dynamic evolution characteristics, opening new avenues for developing novel spintronic devices with enhanced functionalities capable of operating under ambient conditions.

cond-mat.mtrl-sci

BRIGHT: A Collaborative Generalist-Specialist Foundation Model for Breast Pathology

Generalist pathology foundation models (PFMs), pretrained on large-scale multi-organ datasets, have demonstrated remarkable predictive capabilities across diverse clinical applications. However, their proficiency on the full spectrum of clinically essential tasks within a specific organ system remains an open question due to the lack of large-scale validation cohorts for a single organ as well as the absence of a tailored training paradigm that can effectively translate broad histomorphological knowledge into the organ-specific expertise required for specialist-level interpretation. In this study, we propose BRIGHT, the first PFM specifically designed for breast pathology, trained on over 51,000 breast whole-slide images derived from a cohort of over 40,000 patients across 19 hospitals. BRIGHT employs a collaborative generalist-specialist framework to capture both universal and organ-specific features. To comprehensively evaluate the performance of PFMs on breast oncology, we curate the largest multi-institutional cohorts to date for downstream task development and evaluation, comprising over 25,000 WSIs across 10 hospitals. The validation cohorts cover the full spectrum of breast pathology across 25 distinct clinical tasks spanning diagnosis, biomarker prediction, treatment response and survival prediction. Extensive experiments demonstrate that BRIGHT outperforms five leading generalist PFMs, achieving state-of-the-art (SOTA) performance in 25 of 25 internal validation tasks and in 4 of 11 external validation tasks with excellent heatmap interpretability. By evaluating on large-scale validation cohorts, this study not only demonstrates BRIGHT's clinical utility in breast oncology but also validates a collaborative generalist-specialist paradigm, providing a scalable template for developing PFMs on a specific organ system, accelerating the translation of foundation models into ...

cs.CV

Fair Regression under Demographic Parity: A Unified Framework

We propose a unified framework for fair regression tasks formulated as risk minimization problems subject to a demographic parity constraint. Unlike many existing approaches that are limited to specific loss functions or rely on challenging non-convex optimization, our framework is applicable to a broad spectrum of regression tasks. Examples include linear regression with squared loss, binary classification with cross-entropy loss, quantile regression with pinball loss, and robust regression with Huber loss. We derive a novel characterization of the fair risk minimizer, which yields a computationally efficient estimation procedure for general loss functions. Theoretically, we establish the asymptotic consistency of the proposed estimator and derive its convergence rates under mild assumptions. We illustrate the method's versatility through detailed discussions of several common loss functions. Numerical results demonstrate that our approach effectively minimizes risk while satisfying fairness constraints across various regression settings.

stat.ME

Searth Transformer: A Transformer Architecture Incorporating Earth's Geospheric Physical Priors for Global Mid-Range Weather Forecasting

Accurate global medium-range weather forecasting is fundamental to Earth system science. Most existing Transformer-based forecasting models adopt vision-centric architectures that neglect the Earth's spherical geometry and zonal periodicity. In addition, conventional autoregressive training is computationally expensive and limits forecast horizons due to error accumulation. To address these challenges, we propose the Shifted Earth Transformer (Searth Transformer), a physics-informed architecture that incorporates zonal periodicity and meridional boundaries into window-based self-attention for physically consistent global information exchange. We further introduce a Relay Autoregressive (RAR) fine-tuning strategy that enables learning long-range atmospheric evolution under constrained memory and computational budgets. Based on these methods, we develop YanTian, a global medium-range weather forecasting model. YanTian achieves higher accuracy than the high-resolution forecast of the European Centre for Medium-Range Weather Forecasts and performs competitively with state-of-the-art AI models at one-degree resolution, while requiring roughly 200 times lower computational cost than standard autoregressive fine-tuning. Furthermore, YanTian attains a longer skillful forecast lead time for Z500 (10.3 days) than HRES (9 days). Beyond weather forecasting, this work establishes a robust algorithmic foundation for predictive modeling of complex global-scale geophysical circulation systems, offering new pathways for Earth system science.

cs.LG

Intelligence Degradation in Long-Context LLMs: Critical Threshold Determination via Natural Length Distribution Analysis

Large Language Models (LLMs) exhibit catastrophic performance degradation when processing contexts approaching certain critical thresholds, even when information remains relevant. This intelligence degradation-defined as over 30% drop in task performance-severely limits long-context applications. This degradation shows a common pattern: models maintain strong performance up to a critical threshold, then collapse catastrophically. We term this shallow long-context adaptation-models adapt for short to medium contexts but fail beyond critical thresholds. This paper presents three contributions: (1) Natural Length Distribution Analysis: We use each sample's natural token length without truncation or padding, providing stronger causal evidence that degradation results from context length itself. (2) Critical Threshold Determination: Through experiments on a mixed dataset (1,000 samples covering 5%-95% of context length), we identify the critical threshold for Qwen2.5-7B at 40-50% of maximum context length, where F1 scores drop from 0.55-0.56 to 0.3 (45.5% degradation), using five-method cross-validation. (3) Unified Framework: We consolidate shallow adaptation, explaining degradation patterns and providing a foundation for mitigation strategies. This work provides the first systematic characterization of intelligence degradation in open-source Qwen models, offering practical guidance for deploying LLMs in long-context scenarios.

cs.CL

Skyrmion Sliding Switch in a 90-nm-Wide Nanostructured Chiral Magnet

Magnetic skyrmions, renowned for their fascinating electromagnetic properties, hold potential for next-generation topological spintronic devices. Recent advancements have unveiled a rich tapestry of 3D topological magnetism. Nevertheless, the practical application of 3D topological magnetism in the development of topological spintronic devices remains a challenge. Here, we showcase the experimental utilization of 3D topological magnetism through the exploitation of skyrmion-edge attractive interactions in 90-nm-wide confined chiral FeGe and CoZnMn magnetic nanostructures. These attractive interactions result in two degenerate equilibrium positions, which can be naturally interpreted as binary bits for a skyrmion sliding switch. Our theory and simulation reveal current-driven spiral motions of skyrmions, governed by the anisotropic gradient of the potential landscape. Our experiments validate the theory that predicts a tunable threshold current density via magnetic field and temperature modulation of the energy barrier. Our results offer an approach for implementing universal on-off switch functions in 3D topological spintronic devices.

cond-mat.mes-hall

Solving LLM Repetition Problem in Production: A Comprehensive Study of Multiple Solutions

The repetition problem, where Large Language Models (LLMs) continuously generate repetitive content without proper termination, poses a critical challenge in production deployments, causing severe performance degradation and system stalling. This paper presents a comprehensive investigation and multiple practical solutions for the repetition problem encountered in real-world batch code interpretation tasks. We identify three distinct repetition patterns: (1) business rule generation repetition, (2) method call relationship analysis repetition, and (3) PlantUML diagram syntax generation repetition. Through rigorous theoretical analysis based on Markov models, we establish that the root cause lies in greedy decoding's inability to escape repetitive loops, exacerbated by self-reinforcement effects. Our comprehensive experimental evaluation demonstrates three viable solutions: (1) Beam Search decoding with early_stopping=True serves as a universal post-hoc mechanism that effectively resolves all three repetition patterns; (2) presence_penalty hyperparameter provides an effective solution specifically for BadCase 1; and (3) Direct Preference Optimization (DPO) fine-tuning offers a universal model-level solution for all three BadCases. The primary value of this work lies in combining first-hand production experience with extensive experimental validation. Our main contributions include systematic theoretical analysis of repetition mechanisms, comprehensive evaluation of multiple solutions with task-specific applicability mapping, identification of early_stopping as the critical parameter for Beam Search effectiveness, and practical production-ready solutions validated in real deployment environments.

cs.AI

Eigenvalue Estimate for the Rough Laplacian on $1$-Forms and its Applications

In this article, we establish a geometric lower bound for the first positive eigenvalue $\lambda^{(1)}_{1}$ of the rough Laplacian acting on $1$-forms for closed $2n$-dimensional Riemannian manifolds with nonvanishing Euler characteristic. In contrast to the case of functions, such a Li-Yau-type estimate does not hold in general, as evidenced by existing counterexamples. Under assumptions including a lower bound on Ricci curvature, an upper bound on diameter, and an $L^{2p}$-norm bound on the Riemann curvature tensor, we prove that $\lambda^{(1)}_{1}$ is bounded below by a positive constant depending on these parameters. As applications, we derive vanishing results for the Euler characteristic under certain Ricci curvature bounds and the presence of a nonzero Killing vector field, extending classical Bochner-type theorems.

math.DG

Real Time Detection and Quantitative Analysis of Spurious Forgetting in Continual Learning

Catastrophic forgetting remains a fundamental challenge in continual learning for large language models. Recent work revealed that performance degradation may stem from spurious forgetting caused by task alignment disruption rather than true knowledge loss. However, this work only qualitatively describes alignment, relies on post-hoc analysis, and lacks automatic distinction mechanisms. We introduce the shallow versus deep alignment framework, providing the first quantitative characterization of alignment depth. We identify that current task alignment approaches suffer from shallow alignment - maintained only over the first few output tokens (approximately 3-5) - making models vulnerable to forgetting. This explains why spurious forgetting occurs, why it is reversible, and why fine-tuning attacks are effective. We propose a comprehensive framework addressing all gaps: (1) quantitative metrics (0-1 scale) to measure alignment depth across token positions; (2) real-time detection methods for identifying shallow alignment during training; (3) specialized analysis tools for visualization and recovery prediction; and (4) adaptive mitigation strategies that automatically distinguish forgetting types and promote deep alignment. Extensive experiments on multiple datasets and model architectures (Qwen2.5-3B to Qwen2.5-32B) demonstrate 86.2-90.6% identification accuracy and show that promoting deep alignment improves robustness against forgetting by 3.3-7.1% over baselines.

cs.LG

Improvement of P\'{o}lya's conjecture for balls and cylinders

P\'{o}lya's conjecture on the eigenvalues of the Laplacian has been one of the core problems in spectral geometry. Building upon the recent breakthrough works on P\'{o}lya's conjecture for balls and annuli by Filonov, Levitin, Polterovich and Sher, we study several aspects of P\'{o}lya's conjecture for balls and cylinders: by refining the purely analytical portion of the proof in [2] for the Neumann P\'{o}lya's conjecture for the disk, we extend the regime of the spectral parameter that can be established without computer assistance; we obtain improvement of P\'{o}lya's conjecture for disks and balls; we obtain improvement of P\'{o}lya's conjecture for cylinders and confirm the Neumann P\'{o}lya's conjecture for cylinders in $\mathbb{R}^3$. As a supplementary effort, we study Weyl's law for cylinders.

math.CA

$\alpha$-decay half-lives and $\alpha$-cluster preformation factors of nuclei around $N=Z$ line

In this work, a microscopic effective nucleon-nucleon interaction based on the Dirac-Brueckner-Hartree-Fock $G$ matrix starting from a bare nucleon-nucleon interaction is used to explore the $\alpha$-decay half-lives of the nuclei near $N=Z$ line. Specifically, the $\alpha$-nucleus potential is constructed by doubly folding the effective nucleon-nucleon interaction with respect to the density distributions of both the $\alpha$-cluster and daughter nucleus. Moreover, the $\alpha$-cluster preformation factor is extracted by a cluster formation model. It is shown that the calculated half-lives can reproduce the experimental data well. Then, the $\alpha$-decay half-lives that are experimentally unavailable for the nuclei around $N=Z$ are predicted, which are helpful for searching for the new candidates of $\alpha$-decay in future experiments. In addition, by analyzing the proton-neutron correlation energy and two protons-two neutrons correlation energy of $Z=52$ and $Z=54$ isotopes, the $\alpha$-cluster preformation factor evolution with $N$ is explained. Furthermore, it is found that the two protons-two neutrons interaction plays more important role in $\alpha$-cluster preformation than the proton-neutron interaction. Meanwhile, proton-neutron interaction results in the odd-even effect of the $\alpha$-cluster preformation factor.

nucl-th

IllumFlow: Illumination-Adaptive Low-Light Enhancement via Conditional Rectified Flow and Retinex Decomposition

We present IllumFlow, a novel framework that synergizes conditional Rectified Flow (CRF) with Retinex theory for low-light image enhancement (LLIE). Our model addresses low-light enhancement through separate optimization of illumination and reflectance components, effectively handling both lighting variations and noise. Specifically, we first decompose an input image into reflectance and illumination components following Retinex theory. To model the wide dynamic range of illumination variations in low-light images, we propose a conditional rectified flow framework that represents illumination changes as a continuous flow field. While complex noise primarily resides in the reflectance component, we introduce a denoising network, enhanced by flow-derived data augmentation, to remove reflectance noise and chromatic aberration while preserving color fidelity. IllumFlow enables precise illumination adaptation across lighting conditions while naturally supporting customizable brightness enhancement. Extensive experiments on low-light enhancement and exposure correction demonstrate superior quantitative and qualitative performance over existing methods.

cs.CV

VQualA 2025 Challenge on Engagement Prediction for Short Videos: Methods and Results

This paper presents an overview of the VQualA 2025 Challenge on Engagement Prediction for Short Videos, held in conjunction with ICCV 2025. The challenge focuses on understanding and modeling the popularity of user-generated content (UGC) short videos on social media platforms. To support this goal, the challenge uses a new short-form UGC dataset featuring engagement metrics derived from real-world user interactions. This objective of the Challenge is to promote robust modeling strategies that capture the complex factors influencing user engagement. Participants explored a variety of multi-modal features, including visual content, audio, and metadata provided by creators. The challenge attracted 97 participants and received 15 valid test submissions, contributing significantly to progress in short-form UGC video engagement prediction.

cs.CV

A Unified Cortical Circuit Model with Divisive Normalization and Self-Excitation for Robust Representation and Memory Maintenance

Robust information representation and its persistent maintenance are fundamental for higher cognitive functions. Existing models employ distinct neural mechanisms to separately address noise-resistant processing or information maintenance, yet a unified framework integrating both operations remains elusive -- a critical gap in understanding cortical computation. Here, we introduce a recurrent neural circuit that combines divisive normalization with self-excitation to achieve both robust encoding and stable retention of normalized inputs. Mathematical analysis shows that, for suitable parameter regimes, the system forms a continuous attractor with two key properties: (1) input-proportional stabilization during stimulus presentation; and (2) self-sustained memory states persisting after stimulus offset. We demonstrate the model's versatility in two canonical tasks: (a) noise-robust encoding in a random-dot kinematogram (RDK) paradigm; and (b) approximate Bayesian belief updating in a probabilistic Wisconsin Card Sorting Test (pWCST). This work establishes a unified mathematical framework that bridges noise suppression, working memory, and approximate Bayesian inference within a single cortical microcircuit, offering fresh insights into the brain's canonical computation and guiding the design of biologically plausible artificial neural architectures.

q-bio.NC

Decomposition of the Curvature Operator and Applications to the Hopf Conjecture

In this article, we investigate the interplay between the curvature operator, Weyl curvature, and the Hopf conjecture on compact Riemannian manifolds of even dimension. By decomposing the curvature operator into Hermitian components, we develop eigenvalue criteria for sectional curvature and prove vanishing theorems for Betti numbers under integral bounds on the Weyl tensor. Our results confirm the Hopf conjecture for manifolds with sufficiently small Weyl curvature, including locally conformally flat cases, and provide new rigidity theorems under harmonic Weyl curvature conditions.

math.DG