SearcharxivSearch

arXiv subjects

Lei Dong

Publications and source records attributed to Lei Dong.

At least 19 recordsLinked to original sources

The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models

A capability appears in a language model when the last parts of its circuit align in one stochastic attempt, and getting all but one right is worth nothing. We show this no-partial-credit joint alignment is the rate-limiting step of capability formation. Two fingerprints: in a shortcut-free apparatus a five-part circuit missing three waits as long as a three-part circuit missing three (1.19-1.37), so the wait counts missing parts, not size; and on Pythia across seven capabilities and three scales, ablating one part leaves a median 17% of the capability in 32 of 32 discriminating cells, where partial credit predicts 50-83% (p = 2e-10), while a random non-part head leaves 100%. One rare event whose barrier grows with missing parts yields a rate equation -- sites x attempts x drive x exp(-beta*K), minus destruction -- read three ways, each preregistered with frozen constants. Forward: a capability flat at baseline ignites at a step of our choosing once the mix passes a concentration floor (10/10 above, 0/12 below), and while still flat its arrival is datable from its precursor to 5% median error on six held-out models. Backward: the delay to learn a withheld capability grows with waiting until, past a critical step, it never ignites -- yet validation loss falls smoothly throughout, so standard monitors are blind to it. We locate the damage (heads commit to the base data) and isolate the cure: re-initializing only the query-key slices restores learnability (6/6) while the value slices do nothing (0/6). We prove the mechanism in a controlled gated-attention model: occupation forces a deadline whose consequences need no mixing assumption. Completed: SGD's noise fails the fluctuation-dissipation test, so we install one and anneal, melt and pin circuits on schedule. Scope: conjunction circuits in transformers to 1.4B.

cs.LG

Generating 2-Gray codes for grand Motzkin paths and grand Dyck paths with air pockets in constant amortized time

A grand Motzkin path with air pockets is a non-empty lattice path in the first and fourth quadrant of $\mathbb{Z}^2$, starting at the origin $(0,0)$, ending on the $x$-axis, and consisting of up-steps $(1, 1)$, horizontal steps $(1, 0)$, down-steps $(1, -k)$ where $k \geq 1$, and with no consecutive down-steps. A {grand Dyck path with air pockets} is a grand Motzkin path with air pockets that uses no horizontal steps. We present the first known 2-Gray codes for grand Motzkin paths with air pockets. Setting the number of horizontal steps to zero in our algorithm yields the first known 2-Gray codes for grand Dyck paths with air pockets. Our three-stage algorithm generates each path in constant amortized time per string, using $O(n^2)$ memory. We also provide enumeration formulae for grand Motzkin paths and grand Dyck paths with air pockets.

math.CO

A Geometric Analysis of Sign-Magnitude Asymmetry in a ReLU + RMSNorm Block under Ternary Quantization

Pre-norm Transformers with RMSNorm tolerate ternary {-1,0,+1} weight quantization with surprisingly small loss (Ma et al., 2024). We give a geometric explanation via sign-magnitude decomposition of weight perturbations. In a two-layer ReLU + RMSNorm model with i.i.d. Gaussian weights, sign-flips produce $\pi/(\pi-2) \approx 2.75$ times more transverse output energy than sign-preserving magnitude perturbations of equal Frobenius norm, as the flip rate $p \to 0$ (Theorem 3). The mechanism: ReLU creates a hidden-space directional asymmetry between the two perturbation types, which RMSNorm's transverse-projection Fr\'echet derivative selectively exposes. Sign-quantization error is itself a sign-preserving perturbation with angular alignment $\cos^2 \to 2/\pi$ (Theorem 4); its post-ReLU radial fraction ($0.365$) matches the pre-ReLU value $1-2/\pi$ within $0.4\%$, so ReLU is approximately transparent to ternary error. Multi-layer compounding of the $2.75\times$ factor is not experimentally supported; the gap to real-model sign sensitivity arises from outlier features violating delocalization. For an input dimension with amplitude $\alpha$, a single sign-flip produces post-ReLU energy amplified by $R \approx n\alpha^2$ relative to a delocalized entry. On TinyLlama-1.1B, at linear response ($p \leq 0.5\%$), count-matched NLL leverage stabilizes at $\sim 10\times \approx n\mathbb{E}[\alpha^2]$, matching the per-entry theory; the all-column NLL ratio of $5.0\times$ falls within $R_{\mathrm{col}} \leq 19$ ($67\times$ PPL gap reflects metric nonlinearity). Measured outlier $\alpha$ at layer 12 (median $0.024$, max $0.26$) confirms heavy-tailed concentration. The Bussgang constant $2/\pi$, RMSNorm geometry, and ReLU half-space structure together explain sign-magnitude asymmetry in pre-norm models, with $R \propto n\alpha^2$ accounting for real-model deviations.

cs.LG

Emergent Detailed Balance in Human Mobility under Temporal Coarse-Graining

A fundamental question in nonequilibrium statistical physics is whether effective equilibrium behavior can emerge at coarse-grained scales in strongly driven systems. Here, we investigate this question in the context of human mobility by analyzing five years of intercity flow data covering millions of travelers. While short-term flows are highly asymmetric, temporal coarse-graining reveals that over half of all city pairs converge toward effective flow balance, with normalized directional imbalance decaying as a power law. The remaining pairs either exhibit persistent drift-dominated currents or a crossover between these two extremes. A stochastic model decomposing mobility into directional drift and correlated fluctuations quantitatively captures the coexistence of all three regimes. Directly measured variance scaling of the fluctuation process confirms near-diffusive behavior with regime-dependent deviations. These results demonstrate that large-scale mobility networks exhibit a scale-dependent transition from broken to restored flow symmetry, with direct implications for modeling transport and spreading dynamics.

physics.soc-ph

Detection and manipulation of surface electric field noise of hexagonal boron nitride

Hexagonal boron nitride (hBN) spin defects off er transformative potential for quantum sensing through atomic-scale proximity to target samples, yet their performance is fundamentally limited by rapid coherence loss. While magnetic noise mechanisms have been extensively studied, another critical infl uence from surface electric fi eld noise remains unexplored in hBN systems. Here,we address this challenge and systematically investigate surface electric fi eld noise in hBN using shallow boron vacancy defects. The double-quantum spin relaxation behavior in response to magnetic fi elds and defect depths is examined, revealing that the relaxation rate follows a distinctive depth-related power-law dependence of ODMR splitting frequency. The relaxation is also demonstrated to be independent of the defect concentrations. Furthermore, the temperature dependence of the relaxation rate is investigated, showing a noticeable rise as the temperature increases from 296 K to 453 K, thus highlighting the infl uence of thermal eff ects on spin relaxation. To further suppress surface electric fi eld noise, we explore the eff ectiveness of passivation materials, including glycerol and PMMA. Notably, PMMA is more effi cient in mitigating surface electric fi eld noise. These experiments enhance the understanding of surface electric fi eld noise in hBN and provide a foundation for developing noise mitigation strategies in future research.

quant-ph

Switching exploration modes in human mobility

Recent advances in human mobility research have revealed consistent pairwise characteristics in movement behavior, yet existing mobility models often overlook the spatial and topological structure of mobility networks. By analyzing millions of devices' anonymized cell phone trajectories, we uncover a distinct modular organization within these networks, demonstrating that movements within spatial modules differ significantly from those between modules. This finding challenges the conventional assumption of uniform mobility dynamics and underscores the influence of heterogeneous environments on human movement. Inspired by switching behaviors in animal movement patterns, we introduce a novel "switch mechanism" to differentiate movement modes, allowing our model to accurately reproduce both the modular structures of trajectory networks and spatial mobility patterns. Our results provide new insights into the dynamics of human mobility and its impact on network formation, with broad applications in traffic prediction, disease transmission modeling, and urban planning. Beyond advancing the theoretical and practical understanding of mobility networks, this work opens new avenues for understanding societal dynamics at large.

physics.soc-ph

POINT: a web-based platform for pharmacological investigation enhanced by multi-omics networks and knowledge graphs

Network pharmacology (NP) explores pharmacological mechanisms through biological networks. Multi-omics data enable multi-layer network construction under diverse conditions, requiring integration into NP analyses. We developed POINT, a novel NP platform enhanced by multi-omics biological networks, advanced algorithms, and knowledge graphs (KGs) featuring network-based and KG-based analytical functions. In the network-based analysis, users can perform NP studies flexibly using 1,158 multi-omics biological networks encompassing proteins, transcription factors, and non-coding RNAs across diverse cell line-, tissue- and disease-specific conditions. Network-based analysis-including random walk with restart (RWR), GSEA, and diffusion profile (DP) similarity algorithms-supports tasks such as target prediction, functional enrichment, and drug screening. We merged networks from experimental sources to generate a pre-integrated multi-layer human network for evaluation. RWR demonstrated superior performance with a 33.1% average ranking improvement over the second-best algorithm, PageRank, in identifying known targets across 2,002 drugs. Additionally, multi-layer networks significantly improve the ability to identify FDA-approved drug-disease pairs compared to the single-layer network. For KG-based analysis, we compiled three high-quality KGs to construct POINT KG, which cross-references over 90% of network-based predictions. We illustrated the platform's capabilities through two case studies. POINT bridges the gap between multi-omics networks and drug discovery; it is freely accessible at http://point.gene.ac/.

q-bio.MN

Global evidence for a consistent spatial footprint of intra-urban centers

Urban space is highly heterogeneous, with economic and population activities concentrating in localized centers. However, the global organization of such intra-urban centers remains poorly understood due to the lack of consistent, comparable data. Here we develop a scalable geospatial framework using nighttime light observations to identify over 15,000 intra-urban centers worldwide. We uncover a robust regularity: despite differences in city size, geography, and development context, total urban area scales linearly with the number of centers, implying a roughly constant spatial footprint per center. This macroscopic regularity is underpinned by two independent sublinear scaling laws -- center number and urban area both scale with population at closely matched rates -- whose ratio cancels to produce the observed linear relationship. At the within-city level, this constancy manifests as a characteristic Voronoi coverage area per center that is consistent across regions, and centers are more regularly spaced than spatial null models predict. As a consequence, polycentric cities maintain stable accessibility as they expand. These findings provide a new empirical foundation for understanding the spatial organization of urban growth.

physics.soc-ph

Distilling human mobility models with symbolic regression

Human mobility is a fundamental aspect of social behavior, with broad applications in transportation, urban planning, and epidemic modeling. Represented by the gravity model and the radiation model, established analytical models for mobility phenomena are often discovered by analogy to physical processes. Such discoveries can be challenging and rely on intuition, while the potential of emerging social observation data in model discovery is largely unexploited. Here, we propose a systematic approach that leverages symbolic regression to automatically discover interpretable models from human mobility data. Our approach finds several well-known formulas, such as the distance decay effect and classical gravity models, as well as previously unknown ones, such as an exponential-power-law decay that can be explained by the maximum entropy principle. By relaxing the constraints on the complexity of model expressions, we further show how key variables of human mobility are progressively incorporated into the model, making this framework a powerful tool for revealing the underlying mathematical structures of complex social phenomena directly from observational data.

physics.soc-ph

Universal expansion of human mobility across urban scales

Human mobility is a fundamental process underpinning socioeconomic life and urban structure. Classic theories, such as egocentric activity spaces and central place theory, provide crucial insights into specific facets of movement, like home-centricity and hierarchical spatial organization. However, identifying universal characteristics or an underlying principle that quantitatively links these disparate perspectives has remained a challenge. Here, we reveal such a connection by analyzing the spatial structure of individual daily mobility trajectories using network-based modules. We discover a universal scaling law: the spatial extent (radius) of these mobility modules expands sublinearly with increasing distance from home, a pattern consistent across three orders of magnitude. Furthermore, we demonstrate that these modules precisely map onto the nested hierarchy of urban systems, corresponding to local, city-level, and regional scales as distance from home increases. These findings deepen our understanding of human mobility dynamics and demonstrate the profound connection between classical urban theory, human geography, and mobility studies.

physics.soc-ph

Noninvasive magnetic detection of 2D van der Waals room-temperature ferromagnet Fe3GaTe2 using divacancy spins in SiC

Room-temperature (RT) two-dimensional (2D) van der Waals (vdW) ferromagnets hold immense promise for next-generation spintronic devices for information storage and processing. To achieve high-density energy-efficient spintronic devices, it is essential to understand local magnetic properties of RT 2D vdW magnets. In this work, we realize noninvasive in situ magnetic detection in vdW-layered ferromagnet Fe3GaTe2 using divacancy spins quantum sensor in silicon carbide (SiC) at RT. The structural features and magnetic properties of the Fe3GaTe2 are characterized utilizing Raman spectrum, magnetization and magneto-transport measurements. Further detailed analysis of temperature- and magnetic field-dependent optically detected magnetic resonances of the PL6 divacancy near the Fe3GaTe2 reveal that, the Curie temperature (Tc) of Fe3GaTe2 is ~360K, and the magnetization increases with external magnetic fields. Additionally, spin relaxometry technology is employed to probe the magnetic fluctuations of Fe3GaTe2, revealing a peak in the spin relaxation rate around Tc. These experiments give insights into the intriguing local magnetic properties of 2D vdW RT ferromagnet Fe3GaTe2 and pave the way for the application of SiC quantum sensors in noninvasive in situ magnetic detection of related 2D vdW magnets.

quant-ph

Hyperuniform organization in human settlements

Quantifying the spatial organization of human settlements is fundamental to understanding the complexity of urban systems. However, the quantitative patterns of the distribution of villages, towns, and cities that lie between random and regular, are still largely unknown. Here, by analyzing the geographic location of settlements in diverse regions, we show that the apparently complex urban systems can be characterized by disordered hyperuniformity (with small density fluctuations), an intriguing pattern that has been identified in many physical and biological systems, but has rarely been documented in socio-economic systems. By introducing the mechanisms of spatial matching and competition, we develop a growth model that shows how settlements evolve towards hyperuniformity. Our model also predicts the heavy-tail population distribution across settlements, in agreement with empirical observations. These results provide insights into the self-organization of cities, and reveal the universality of spatial organization shared by social, physical, and biological systems.

physics.soc-ph

Superimposed Pilot-based Channel Estimation for RIS-Assisted IoT Systems Using Lightweight Networks

Conventional channel estimation (CE) for Internet of Things (IoT) systems encounters challenges such as low spectral efficiency, high energy consumption, and blocked propagation paths. Although superimposed pilot-based CE schemes and the reconfigurable intelligent surface (RIS) could partially tackle these challenges, limited researches have been done for a systematic solution. In this paper, a superimposed pilot-based CE with the reconfigurable intelligent surface (RIS)-assisted mode is proposed and further enhanced the performance by networks. Specifically, at the user equipment (UE), the pilot for CE is superimposed on the uplink user data to improve the spectral efficiency and energy consumption for IoT systems, and two lightweight networks at the base station (BS) alleviate the computational complexity and processing delay for the CE and symbol detection (SD). These dedicated networks are developed in a cooperation manner. That is, the conventional methods are employed to perform initial feature extraction, and the developed neural networks (NNs) are oriented to learn along with the extracted features. With the assistance of the extracted initial feature, the number of training data for network training is reduced. Simulation results show that, the computational complexity and processing delay are decreased without sacrificing the accuracy of CE and SD, and the normalized mean square error (NMSE) and bit error rate (BER) performance at the BS are improved against the parameter variance.

eess.SP

Constructing multi-level urban clusters based on population distributions and interactions

A city (or an urban cluster) is not an isolated spatial unit, but a combination of areas with closely linked socio-economic activities. However, so far, we lack a consistent and quantitative approach to define multi-level urban clusters through these socio-economic connections. Here, using granular population distribution and flow data from China, we propose a bottom-up aggregation approach to quantify urban clusters at multiple spatial scales. We reveal six 'phases' (i.e., levels) in the population density-flow diagram, each of which corresponds to a spatial configuration of urban clusters from large to small. Besides, our results show that Zipf's law appears only after the fifth level, confirming the spatially dependent nature of urban laws. Our approach does not need pre-defined administrative boundaries and can be applied effectively on a global scale.

physics.soc-ph

An efficient dosimetry method with a Faraday cup for small animal, small-field proton irradiation under conventional and ultra-high dose rates

Introduction: We developed and evaluated a method for dose calibration and monitoring under conventional and ultra-high dose rates for small animal experiments with small-field proton beams using a Faraday cup. Methods: We determined a relationship between dose and optical density (OD) of EBT-XD Gafchromic film using scanned 10x10 cm2 proton pencil beams delivered at clinical dose rates; the dose was measured with an Advanced Markus chamber. On a small animal proton irradiation platform, double-scattered pencil beams with 5 or 8 mm diameter brass collimation at conventional and ultra-high dose rates were delivered to the EBT-XD films. The proton fluence charges were collected by a Faraday cup placed downstream from the film. The average of the irradiated film ODs was related to the Faraday cup charges. A conversion from the Faraday cup charge to the average dose of the small-field proton beam was then obtained. Results: The relationship between the small-field average profile dose and Faraday cup charge was established for 10 and 15 Gy mice FLASH experiments. The film OD was found to be independent of dose rate. At small-animal treatments, the Faraday cup readings were conveniently used to QA and monitor the delivered dose and dose rates to the mice under conventional and ultra-high dose rates. Conclusion: The dose calibration and monitoring method with Faraday cup for small animal proton FLASH experiments is time-efficient and cost-effective and can be used for irradiations of various small field sizes. The same approach can also be adopted for clinical proton dosimetry for small-field irradiations.

physics.med-ph

Quantifying navigation complexity in transportation networks

The complexity of navigation in cities has increased with the expansion of urban areas, creating challenging transportation problems that drive many studies on the navigability of networks. However, due to the lack of individual mobility data, large-scale empirical analysis of the wayfinder's real-world navigation is rare. Here, using 225 million subway trips from three major cities in China, we quantify navigation difficulty from an information perspective. Our results reveal that 1) people conserve a small number of repeatedly used routes, and 2) the navigation information in the subnetworks formed by those routes is much smaller than the theoretical value in the global network, suggesting that the decision cost for actual trips is significantly smaller than the theoretical upper limit found in previous studies. By modeling routing behaviors in growing networks, we show that while the global network becomes difficult to navigate, navigability can be improved in subnetworks. We further present a universal linear relationship between the empirical and theoretical search information, which allows the two metrics to predict each other. Our findings demonstrate how large-scale observations can quantify real-world navigation behaviors and aid in evaluating transportation planning.

physics.soc-ph

MetroGAN: Simulating Urban Morphology with Generative Adversarial Network

Simulating urban morphology with location attributes is a challenging task in urban science. Recent studies have shown that Generative Adversarial Networks (GANs) have the potential to shed light on this task. However, existing GAN-based models are limited by the sparsity of urban data and instability in model training, hampering their applications. Here, we propose a GAN framework with geographical knowledge, namely Metropolitan GAN (MetroGAN), for urban morphology simulation. We incorporate a progressive growing structure to learn hierarchical features and design a geographical loss to impose the constraints of water areas. Besides, we propose a comprehensive evaluation framework for the complex structure of urban systems. Results show that MetroGAN outperforms the state-of-the-art urban simulation methods by over 20% in all metrics. Inspiringly, using physical geography features singly, MetroGAN can still generate shapes of the cities. These results demonstrate that MetroGAN solves the instability problem of previous urban simulation GANs and is generalizable to deal with various urban attributes.

cs.CY

Transfer Learning-based Channel Estimation in Orthogonal Frequency Division Multiplexing Systems Using Data-nulling Superimposed Pilots

Data-nulling superimposed pilot (DNSP) effectively alleviates the superimposed interference of superimposed training (ST)-based channel estimation (CE) in orthogonal frequency division multiplexing (OFDM) systems, while facing the challenges of the estimation accuracy and computational complexity. By developing the promising solutions of deep learning (DL) in the physical layer of wireless communication, we fuse the DNSP and DL to tackle these challenges in this paper. Nevertheless, due to the changes of wireless scenarios, the model mismatch of DL leads to the performance degradation of CE, and thus faces the issue of network retraining. To address this issue, a lightweight transfer learning (TL) network is further proposed for the DL-based DNSP scheme, and thus structures a TL-based CE in OFDM systems. Specifically, based on the linear receiver, the least squares estimation is first employed to extract the initial features of CE. With the extracted features, we develop a convolutional neural network (CNN) to fuse the solutions of DLbased CE and the CE of DNSP. Finally, a lightweight TL network is constructed to address the model mismatch. To this end, a novel CE network for the DNSP scheme in OFDM systems is structured, which improves its estimation accuracy and alleviates the model mismatch. The experimental results show that in all signal-to-noise-ratio (SNR) regions, the proposed method achieves lower normalized mean squared error (NMSE) than the existing DNSP schemes with minimum mean square error (MMSE)-based CE. For example, when the SNR is 0 decibel (dB), the proposed scheme achieves similar NMSE as that of the MMSE-based CE scheme at 20 dB, thereby significantly improving the estimation accuracy of CE. In addition, relative to the existing schemes, the improvement of the proposed scheme presents its robustness against the impacts of parameter variations.

eess.SP