SearcharxivSearch

arXiv subjects

Wenjie Zhou

Publications and source records attributed to Wenjie Zhou.

At least 19 recordsLinked to original sources

HyperTransfer: Understanding the Equivalence between Base Optimizer and Hyperball

Hyperball optimizers constrain parameter norms and update only their directions, establishing a distinct paradigm for neural network optimization. Although this geometry appears fundamentally different from that of conventional Base Optimizers, which update both parameter norms and directions, we show that the two paradigms are dynamically equivalent for scale-invariant networks. Building on this equivalence, we propose HyperTransfer, which constructs a Hyperball optimizer that reproduces the dynamics of a target Base Optimizer using only its initialization and learning-rate schedule, without running the target optimizer itself. We further derive the inverse mapping and extend the framework to non-scale-invariant networks. Experiments show that both HyperTransfer and the inverse mapping produce loss trajectories nearly identical to those of their targets, suggesting that Hyperball dynamics are governed primarily by the induced effective learning-rate schedule and optimizer state.

cs.LG

Anderson Lattice in Incommensurate $\bf{Nb_3Cl_8}$/Graphene van der Waals Heterostructures

The periodic Anderson model}, traditionally realized in rare-earth compounds with limited tunability, have hindered systematic exploration of correlated quantum phenomena. Here, we introduce a strategy for {realizing and }engineering {this model} in incommensurate van der Waals heterostructures by coupling a Mott insulator (Nb$_3$Cl$_8$) with itinerant electrons (from monolayer graphene), circumventing strict lattice-matching requirements. Through magnetotransport and slave spin mean-field calculations, we demonstrate the hybridization gap ($Δ\approx30$ meV), gate-tunable metal-insulator transition, and band-selective electron effective mass enhancement, hallmarks of Kondo coherence. The heterostructure exhibits a nearly order-of-magnitude enhancement in the effective electron mass between hybridized and conventional graphene-like regimes, alongside in-plane magnetic field-induced metal-insulator transitions. Top gate-temperature phase mapping reveals competing correlated states, including insulating and hidden-order phases. This work establishes an electrically tunable van der Waals platform for studying correlated states generated by coupling a Mott-insulating layer to an itinerant-electron system, providing a materials route for exploring low-dimensional correlated quantum phases.

cond-mat.str-el

TOI-2147 b and TOI-6019 b: Two eccentric warm Jupiters detected and characterized with TESS and MaHPS

The population of Jupiter-sized exoplanets with orbital periods between 10 and 200 days (WJs) exhibits a broad range of orbital eccentricities and system architectures, suggesting a diversity of formation and migration pathways. In this work, we report the detection and characterization of two new eccentric WJs, TOI-2147 b and TOI-6019 b, initially identified as planet candidates by the Transiting Exoplanet Survey Satellite (TESS). We combined TESS photometry with ground-based follow-up observations, including multiband photometry from LCOGT and MuSCAT2, high-angular-resolution speckle imaging, and high-precision radial velocity measurements from the high-resolution Manfred Hirt Planet Finder Spectrograph (MaHPS). Using these data, we were able to confirm the planetary nature of both candidates. TOI-2147 b has a radius of $10.5 \pm 0.3\,\mathrm{R}_\oplus$ and a mass of $116 \pm 22\,\mathrm{M}_\oplus$. It orbits its slightly metal-poor ($\mathrm{[Fe/H]} = -0.29^{+0.07}_{-0.08}$) G-type host star on an eccentric orbit ($e = 0.29 \pm 0.07$) with a period of 26.2 days. TOI-6019 b has a radius of $12.3 \pm 0.3\,\mathrm{R}_\oplus$ and a mass of $149 \pm 15\,\mathrm{M}_\oplus$. It orbits a slightly evolved, solar-metallicity G-type sub-giant with a period of 14.5 days on a significantly eccentric orbit ($e = 0.48^{+0.05}_{-0.04}$). Both planets have bulk densities below that of Jupiter, indicating mildly inflated radii, with interior structure modeling using GASTLI. This suggests that tidal heating from the nonzero eccentricities likely contributes to this inflation and disfavors large atmospheric metal enrichment. No significant signals from additional companions were detected in the radial velocity time series or transit timing variations. Together with the elevated eccentricities, this is consistent with a high-eccentricity migration origin for both systems.

astro-ph.EP

Visualizing the Impact of Quenched Disorder on 2D Electron Wigner Solids

Electron Wigner solids (WSs)1-12 provide an ideal system for understanding the competing effects of electron-electron and electron-disorder interactions, a central unsolved problem in condensed matter physics. Progress in this topic has been limited by a lack of single-defect-resolved experimental measurements as well as accurate theoretical tools to enable realistic experiment/theory comparison. Here we overcome these limitations by combining atomically-resolved scanning tunneling microscopy (STM) with neural-quantum-state quantum Monte Carlo (NQS-QMC) simulation of disordered 2D electron WSs to discover new disorder-induced physical regimes of correlated electron behavior. STM was used to image the electron density ($n_e$) dependent evolution of electron WSs in gate-tunable bilayer MoSe2 devices with varying long-range ($n_\mathrm{LR}$) and short-range ($n_\mathrm{SR}$) disorder densities. These images were compared to NQS-QMC simulations using realistic disorder maps extracted from experiment, thus allowing the roles of different disorder types to be disentangled. We identify two distinct physical regimes for disordered electron WSs that depend on $n_\mathrm{SR}$. For $n_\mathrm{SR} \lesssim n_e$ the WS behavior is dominated by long-range disorder and features extensive mixed solid-liquid phases, a new type of local re-entrant melting/crystallization, and prominent Friedel oscillations. In contrast, when $n_\mathrm{SR} \gg n_e$ these features are suppressed and a more robust amorphous WS phase emerges that persists to higher ne, highlighting the importance of short-range disorder in this regime. Our work establishes a powerful framework for studying disordered quantum solids via a combined experimental-theoretical approach.

cond-mat.str-el

Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach

Psychometric calibration for educational tests typically requires costly human response data. Large language models (LLMs) simulated examinees offer a promising route to early calibration, but their responses are too accurate and too uniform. We propose Cognitive Diagnostic Profiling (CDP), a zero-shot framework that prompts LLMs to simulate plausible examinees with diverse cognitive profiles: binary attribute-mastery patterns are rendered as natural-language profiles and sampled under an uninformative or an informative distribution. Using the Tatsuoka fraction-subtraction dataset (536 examinees, 15 items, five attributes), we evaluated eight LLM configurations under no-profile, uninformative-CDP, and informative-CDP conditions, assessing alignment with human examinees at the ability-distribution, mastery-profile, and item-difficulty levels. CDP improved all three levels: distributional overlap rose across configurations; weighted correlations between profile-level scores and human profile expectations reached 0.92 to 0.98; and item-difficulty recovery improved in rank order and absolute alignment, most for reasoning-enabled models; in the strongest case, Gemini 3.0 Flash (Thinking), one-parameter logistic (1PL) difficulty Spearman correlations rose from 0.24 to 0.86 and 0.90 and the root-mean-square error (RMSE) fell from 6.31 to 1.30 and 0.90; the informative condition helped most where profile-level alignment was strong. CDP brings LLM-simulated examinees into closer psychometric alignment with human examinees, making them practical for operational test development.

cs.CY

Temporal Fourier Optics Reveals Hidden Hybridized Light-Matter States

Spectral measurements provide fundamental insights into wave systems by revealing resonances, mode hybridization, and light-matter interactions. However, intrinsic dissipation and measurement-related spectral broadening often obscure the spectral signatures of the underlying hybridized light-matter states. Here, we establish a temporal Fourier optics framework based on a space-time Fourier correspondence, which interprets spectral broadening as the Fourier counterpart of temporal attenuation. This perspective introduces a temporal point-spread function (TPSF) that enables direct, synthesis-free reconstruction of the underlying spectral response from experimentally measured spectra by compensating for effective temporal decay before transformation back to the frequency domain. We experimentally validate the framework using deterministic single-molecule Au nanosphere dimers and open Au@Ag nanorod- and nanotriangle-based plasmonic nanocavities coupled to J-aggregate excitons. Across these distinct platforms, TPSF consistently resolves hidden upper and lower polaritonic branches, revealing hybridized light-matter states and strong coupling that remain inaccessible in conventional scattering spectra. The reconstructed spectra agree closely with the recently developed complex-frequency formalism while providing a substantially simpler and experimentally accessible implementation. More broadly, temporal Fourier optics establishes a general framework for recovering dissipation-obscured spectral information, opening new opportunities for spectroscopy, imaging, sensing, and inverse wave measurements across photonics and wave physics.

physics.optics

Cryogenic shock exfoliation for ultrahigh mobility rhombohedral graphite nanoelectronics

Rhombohedral multilayer graphene (RMG) offers a highly tunable platform for correlated electron physics, featuring field-effect control of magnetic, superconducting, and topological phases[1-24]. The promise of these materials has been held back by the limited abundance of rhombohedral stacking in natural graphite, which constrains both sample yield and useful area. Here we introduce 'cryogenic shock exfoliation' to produce large area rhombohedral graphene flakes which, combined with a low-pressure van der Waals assembly technique that preserves stacking order, enable highly uniform devices exceeding 1300 $μm^2$ with fabrication yields of 90%. Using scanning nanoSQUID-on-tip imaging, we demonstrate uniform spin magnetism over the full central 10 times 10 $μm^2$ area of our devices. Transverse magnetic focusing reveals a disorder mean free path exceeding 200 $μm$ at low temperatures. Within the flat surface bands of RMG[20], we observe a size-driven crossover from Poiseuille to porous electron flow in the intermediate-temperature regime of strong electron-electron hydrodynamics[16, 25], providing a further signature of ultrahigh device quality. Our approach overcomes a key materials bottleneck in the fabrication of mesoscopic rhombohedral graphene devices, paving the way for incorporating strongly correlated phases into two-dimensional nanoelectronics.

cond-mat.mes-hall

Rethinking Bregman Divergences in Kronecker-Factored Optimizers

Shampoo-style optimizers approximate gradient covariance matrices using Kronecker-factored structures. Recent work~\cite{lin2026understanding} showed that such approximations can be viewed as projections under Bregman matrix divergences, leading to different Kronecker-factored preconditioners. However, it remains unclear what role the choice of divergence plays when the covariance is not exactly Kronecker-factored. We study this question through the spectrum of the covariance matrix. We show that Frobenius, von Neumann, and LogDet divergences distribute the unavoidable Kronecker approximation error differently across the covariance spectrum. We further show that their Kronecker factors are governed by divergence-weighted residuals rather than the raw approximation error, explaining how these spectral preferences are realized in the resulting preconditioners. Empirically, we observe that the top covariance eigenspace is substantially better aligned with the Hessian matrix, while the tail spectrum is much noisier and unreliable. Motivated by these findings, we propose a subspace-aware Kronecker optimizer that applies eigenvalue-based preconditioning in the top subspace and uses an adaptive isotropic acceleration constant in the bottom subspace.

cs.LG

Extra-Merge: Tracing the Rank-1 Subspace of Model Merging in Language Model Pre-Training

Model merging has emerged as a lightweight paradigm for enhancing Large Language Models (LLMs), yet its underlying mechanisms remain poorly understood. In this work, we analyze late-stage pre-training trajectories and uncover a \textbf{Rank-1 Subspace} phenomenon: while raw optimization steps oscillate violently, consecutive \emph{merged} checkpoints collapse onto a stable, approximately one-dimensional linear manifold. We theoretically ground this observation in a \emph{river-valley} landscape analysis: averaging acts as a geometric low-pass filter that dampens high-curvature noise to reveal the optimal descent direction. Capitalizing on this insight, we propose \textbf{Extra-Merge}, a training-free strategy that extrapolates along this subspace to minimize loss without additional gradient updates. Extensive experiments across GPT-2 and LLaMA families (124M to 2B) demonstrate that Extra-Merge consistently outperforms standard merging baselines. Notably, it yields consistent zero-shot accuracy gains on Pythia-12B downstream tasks and generalizes effectively to the Muon optimizer \citep{jordan2024muon}.

cs.LG

The Stability of Singular Distribution: A Spectral Perspective on the Two-Phase Dynamics of Language Model Pre-training

Large language model pre-training typically exhibits a two-phase trajectory: a fast initial loss drop followed by a prolonged slow improvement. We identify an underlying spectral phenomenon, Stability of Singular Distribution (SoSD), where the trace-normalized singular value spectrum stabilizes early, even as parameter matrices continue to evolve. We demonstrate that synchronization between SoSD and the slow-descent regime is widely observed across diverse architectures (GPT-2, LLaMA) and settings, including various schedules (Step-wise, WSD, Cosine Decay), weight decays, and optimizers (AdamW, Muon). By analyzing a simplified Transformer, we prove that growing weight norms inevitably precipitate an early SoSD threshold, after which the rate of loss decrease becomes theoretically bounded by the variation in the singular distribution. We further interpret strategies like WSD and Muon through their ability to modulate the SoSD scale, offering a spectral lens for understanding efficient pre-training dynamics.

cs.LG

Benchmarking Real-Time Question Answering via Executable Code Workflows

Retrieving real-time information is a fundamental capability for search-integrated agents in real-world applications. However, existing benchmarks are predominantly static and therefore fail to capture the temporal dynamics of information and the continuously evolving nature of real-world knowledge. To address this limitation, we propose RT-QA, a dynamic evaluation framework that leverages executable code workflows to retrieve up-to-date answers at evaluation time. Specifically, we construct an agent-driven pipeline that autonomously generates code for web crawling and DOM-based answer extraction to produce real-time ground truth. To ensure robust evaluation over time, the pipeline further incorporates a self-repair mechanism to adapt to changes in web page structures. RT-QA spans 12 domains (e.g., Finance, Sports) with 320 Chinese questions categorized into three difficulty levels. Extensive evaluations of state-of-the-art models (e.g., GPT-5.2, GLM-4.7) reveal significant limitations in real-time adaptability: even the best models achieve only 46% accuracy. Our analysis highlights two primary failure modes: (1) Lazy Retrieval, where agents rely on search snippets instead of deeply scanning specific websites for information (20% of failures); and (2) Temporal Confusion, a cognitive error where agents retrieve a historical date (e.g., an event in 2024) and fail to re-anchor to the current time (2026) for subsequent reasoning. These findings suggest that future agents require not just better retrieval strategies, but robust temporal state management.

cs.IR

When and Why Grouping Attention Heads Accelerates Muon Optimization

Muon orthogonalizes matrix updates, but multi-head attention naturally operates at the level of heads. This granularity mismatch raises the question of whether Muon should be applied to the full attention projection, to individual heads, or to intermediate head groups. We study this question through a one-step descent comparison between full-matrix Muon and group-wise Muon. Our analysis reveals a trade-off between the \textbf{group-wise whitening gain} from group-wise updates and the \textbf{grouping-induced norm cost}, an additional update-norm cost caused by replacing full-matrix whitening with group-wise whitening. Motivated by this trade-off, we propose \textbf{Group Muon}, which treats head group size and grouping rule as optimizer hyperparameters. On GPT-2 Small trained on FineWeb, appropriate grouping improves validation loss over both full-QKV Muon and fully head-wise MuonSplit.

cs.LG

LMFPPO-UBP: Local Mean Field Proximal Policy Optimization with Unbalanced Punishment for Spatial Public Goods Games

Spatial public goods games are characterized by high-dimensional state spaces and localized externalities, which pose significant challenges for achieving stable and widespread cooperation. Traditional approaches often struggle to effectively capture neighborhood-level strategic interactions and dynamically align individual incentives with collective welfare. To resolve this issue, this paper introduces a novel intelligent decision-making framework called Local Mean-Field Proximal Policy Optimization with Unbalanced Punishment (LMFPPO-UBP). The conventional mean field concept is reformulated as a socio-statistical sensor embedded directly into the policy gradient space of deep reinforcement learning, allowing agents to adapt their strategies based on mesoscale neighborhood dynamics. Additionally, an unbalanced punishment mechanism is integrated to penalize defectors proportionally to the local density of cooperators, thereby reshaping the payoff structures without imposing direct costs on cooperative agents. Experimental results demonstrate that the LMFPPO-UBP promotes rapid and stable global cooperation even under low enhancement factors, consistently outperforming baseline methods such as Q-learning and Fermi update rules. Statistical analyses further validate the framework's effectiveness in lowering the cooperation threshold and achieving better coordinated outcomes.

cs.GT

BSFA: Leveraging the Subspace Dichotomy to Accelerate Neural Network Training

Recent studies \citep{gur2018gradient,song2024does, wen2024understanding} highlight a fundamental dichotomy in deep learning optimization: Although parameter updates along the top eigendirections of the loss Hessian (Dom-space) capture most of the update magnitude, they often contribute minimally to loss reduction. In contrast, updates in the orthogonal component (Bulk-space) have smaller magnitudes but drive most learning progress. In this work, we further advance the understanding of this phenomenon and introduce the \textbf{Bulk-Space-Filtration-Accelerator (BSFA)}, a novel plug-and-play framework. BSFA accelerates training by differentially scaling update components projected onto these distinct subspaces, simultaneously enhancing stability by moderating updates in the dominant subspace and boosting convergence speed by amplifying those in the bulk-space. To ensure BSFA is both practical and scalable for contemporary large models, we introduce two key innovations: an efficient estimator using Principal Component Analysis (PCA) on historical updates for fast subspace estimation, and a block-wise strategy that applies this estimation on a per-parameter-block basis. These designs make BSFA computationally tractable and highly effective. We demonstrate BSFA's acceleration across various tasks, notably achieving approximately 2$\times$ speedup when pre-training LLaMA-72M on WikiText-103 and LLaMA-134M on OpenWebText compared to vanilla AdamW.

cs.LG

Spectral localization of single-nanoparticle plasmons through photonic substrate engineering

Surface plasmon resonances (SPRs) are crucial for confining light beyond the diffraction limit, yet heavy metal losses often limit their spectral localization. Here, we propose a practical strategy for enabling the spectral localization of single-nanoparticle SPRs through photonic substrate engineering, which creates distinct optical pathways (OPs) to tailor the electromagnetic environments around plasmonic nanoparticles. By analyzing the multiplication factor spectrum of the projected local density of states, we can trace and control these OPs, enabling strong spatial and spectral confinement of single-nanoparticle SPRs. Simulations reveal that a photonic crystal substrate can reduce the mode volume by fivefold and boost the quality factor by over 80 times compared to a metal nanoparticle on a dielectric substrate. Proof-of-concept experiments using two types of leaking Fabry-Perot photonic substrates demonstrate active manipulation of SPRs in both "open" and "closed" OP states. This multidimensional photonic substrate engineering establishes a customizable platform for single-nanoparticle plasmonics, potentially transforming applications that were previously limited by spectral localization.

physics.optics

Mimicking the Physicist's Eye:A VLM-centric Approach for Physics Formula Discovery

Automated discovery of physical laws from observational data in the real world is a grand challenge in AI. Current methods, relying on symbolic regression or LLMs, are limited to uni-modal data and overlook the rich, visual phenomenological representations of motion that are indispensable to physicists. This "sensory deprivation" severely weakens their ability to interpret the inherent spatio-temporal patterns within dynamic phenomena. To address this gap, we propose VIPER-R1, a multimodal model that performs Visual Induction for Physics-based Equation Reasoning to discover fundamental symbolic formulas. It integrates visual perception, trajectory data, and symbolic reasoning to emulate the scientific discovery process. The model is trained via a curriculum of Motion Structure Induction (MSI), using supervised fine-tuning to interpret kinematic phase portraits and to construct hypotheses guided by a Causal Chain of Thought (C-CoT), followed by Reward-Guided Symbolic Calibration (RGSC) to refine the formula structure with reinforcement learning. During inference, the trained VIPER-R1 acts as an agent: it first posits a high-confidence symbolic ansatz, then proactively invokes an external symbolic regression tool to perform Symbolic Residual Realignment (SR^2). This final step, analogous to a physicist's perturbation analysis, reconciles the theoretical model with empirical data. To support this research, we introduce PhysSymbol, a new 5,000-instance multimodal corpus. Experiments show that VIPER-R1 consistently outperforms state-of-the-art VLM baselines in accuracy and interpretability, enabling more precise discovery of physical laws. Project page: https://jiaaqiliu.github.io/VIPER-R1/

cs.AI

On ordering of surjective cardinals

Let $\mathrm{Card}$ denote the class of cardinals. For all cardinals $\mathfrak{a}$ and $\mathfrak{b}$, $\mathfrak{a}\leqslant\mathfrak{b}$ means that there is an injection from a set of cardinality $\mathfrak{a}$ into a set of cardinality $\mathfrak{b}$, and $\mathfrak{a}\leqslant^\ast\mathfrak{b}$ means that there is a partial surjection from a set of cardinality $\mathfrak{b}$ onto a set of cardinality $\mathfrak{a}$. A doubly ordered set is a triple $\langle P,\preccurlyeq,\preccurlyeq^\ast\rangle$ such that $\preccurlyeq$ is a partial ordering on $P$, $\preccurlyeq^\ast$ is a preordering on $P$, and ${\preccurlyeq}\subseteq{\preccurlyeq^\ast}$. In 1966, Jech proved that for every partially ordered set $\langle P,\preccurlyeq\rangle$, there exists a model of $\mathsf{ZF}$ in which $\langle P,\preccurlyeq\rangle$ can be embedded into $\langle\mathrm{Card},\leqslant\rangle$. We generalize this result by showing that for every doubly ordered set $\langle P,\preccurlyeq,\preccurlyeq^\ast\rangle$, there exists a model of $\mathsf{ZF}$ in which $\langle P,\preccurlyeq,\preccurlyeq^\ast\rangle$ can be embedded into $\langle\mathrm{Card},\leqslant,\leqslant^\ast\rangle$.

math.LO

The Superconformal Index and Black Hole Instabilities

The superconformal index of ${\cal N}=4$ supersymmetric Yang-Mills theory with gauge group $\mathrm{U}(N)$ has provided powerful insights into the entropy of supersymmetric black holes in AdS$_5\times S^5$, including some sub-leading logarithmic and non-perturbative corrections. Recently, the phase space of supersymmetric solutions has been argued to contain configurations other than the asymptotically AdS$_5$ black hole. Such configurations include the so-called grey galaxies where the black hole at the center is surrounded by a gas of gravitons. By numerically evaluating the superconformal index of ${\cal N}=4$ supersymmetric Yang-Mills at small values of $N$, we detect systematic deviations from the entropy of black holes with two distinct angular momenta. We find that the giant graviton expansion of the index is a numerically efficient way of evaluating the index that complements the direct character evaluation and allows for explicit access to $N\le 15$ with up to two giant gravitons in the expansion. We find it remarkable that a supersymmetric quantity in field theory, usually thought of as a rigid counting observable, indeed contains information about different phases in the space of supersymmetric solutions on the gravity side.

hep-th