SearcharxivSearch

arXiv subjects

Chan Li

Publications and source records attributed to Chan Li.

17 recordsLinked to original sources

Broken Ergodicity and the Violation of the Fluctuation-Dissipation Theorem Lead to Generalization Beyond Overfitting in Machine Learning

The remarkable ability of modern neural networks to generalize improves with increasing network capacity, even when the number of model parameters or effective degrees of freedom exceeds the number of training data points. This phenomenon is all the more surprising given that generalization error diverges when the number of model parameters approaches a critical value from below. Here we use dynamical mean field theory to show that this so-called "double descent" behavior is the outcome of a phase transition in the stochastic field theory describing the training process. We calculate the critical exponents and scaling function of the double descent phase transition, and show that it is marked by a breakdown of the fluctuation-dissipation theorem associated with broken ergodicity. The corresponding response function has the same functional form as the simple London model of the superconducting transition, with the rigidity of the wave function corresponding to the neural network's ability to generalize accurately.

cond-mat.dis-nn

Discrete-phase-randomized mode-pairing quantum key distribution

Mode-pairing quantum key distribution (MP-QKD) protocol achieves performance beyond the repeaterless rate-transmittance bound and exhibits excellent practicality by avoiding the requirement for difficult global phase locking. However, the source side of MP-QKD still relies on the assumption of continuous phase randomization, an experimentally infeasible requirement in practice. Therefore, the practical security of the protocol cannot be fully guaranteed. In this work, we propose a discrete-phase-randomized mode-pairing quantum key distribution (DPR-MP-QKD) protocol and analyze the basis-dependence of the source side. Then, we introduce a concrete discrete version of the decoy state method that ensures the security of the DPR-MP-QKD protocol. Finally, simulation results indicate that as the number of discrete phases increases, the key rate performance of DPR-MP-QKD progressively approaches that of the continuous case, with convergence achieved at approximately 14 discrete phases. Moreover, our approach drastically lowers the demand for randomness. While conventional continuous phase randomization demands an unlimited supply of random bits, we show that merely a few bits (e.g., 4) are adequate.

quant-ph

Intermediate-state Coulomb-corrected strong-field approximation for rescattering processes

We analytically derive the all-order strong-field S-matrix series incorporating intermediate-state Coulomb-Volkov corrections (ICSFA). Focusing on rescattering processes described by the second-order term, we systematically investigate the impact of intermediate-state Coulomb interactions on above-threshold ionization (ATI) spectra of atomic hydrogen in linearly polarized laser fields. Crucially, ICSFA spectra demonstrate superior agreement with the results obtained by numerically solving the time-dependent Schr\"{o}dinger equation compared to the standard strong-field approximation (SFA) and final-state Coulomb-corrected SFA (FCSFA). Our analysis reveals that intermediate-state Coulomb corrections enhance the yield of the third- and fourth-return-recollision trajectories while modifying interference patterns in the energy spectrum. The observed enhancement of the multi-return-recollision trajectories can be attributed to modifications of the ionization yield and scattering cross-section, which are induced by intermediate-state Coulomb effects. These effects are equivalent to the so-called Coulomb focusing effect.

physics.atom-ph

30-meter Land Surface Temperature from Landsat via Progressive Self-Training Downscaling

Land surface temperature (LST) is a critical parameter for characterizing surface energy balance and hydrothermal processes. While Landsat provides invaluable LST observations at medium spatial resolution for over 40 years, its native spatial resolution of thermal bands (e.g., 100 m) remains insufficient compared to its 30 m optical bands, failing to meet the demands of fine-scale studies. To address this issues, this study proposes a progressive self-training framework for downscaling Landsat LST to 30 m without relying on fine-scale ground truth, while maintaining minimal data dependence. The framework progressively optimizes a cross-modal fusion network to refine thermal details in a coarse-to-fine manner, characterized by one pre-training and two fine-tuning stages. Spatial validation against SDGSAT-1 30 m LST and temporal validation using in situ measurements confirm its reliability and accuracy, with both station-averaged MAE and RMSE outperforming the official cubic product by approximately 0.4 K. Further performance comparison experiments demonstrate that the proposed framework consistently reconstructs coherent fine-scale thermal patterns while preserving spatial heterogeneity. Multi spatial resolution evaluations and ablation studies verify the effectiveness of the proposed strategy and network design. Overall, the framework provides a stable pathway for enhancing the spatial resolution of Landsat LST, providing fine-resolution data support for fine-scale surface process studies and localized environmental monitoring.

physics.ao-ph

SwinYNet: A Transformer-based Multi-Task Model for Accurate and Efficient FRB Search

In this study, we present a transformer-based multi-task model for Fast Radio Burst (FRB) detection, signal segmentation, and parameter estimation directly from time-frequency data, without requiring computationally expensive de-dispersion preprocessing. To overcome the scarcity of labeled observational data, we develop an FRB simulator and a rule-based automatic annotation pipeline, enabling training exclusively on simulated data. Evaluations on the FAST-FREX dataset show that our model achieves an F1 score of 97.8%, recall of 95.7%, and precision of 100%, outperforming both conventional tools (e.g., PRESTO, Heimdall) and recent AI-based baselines (e.g., RaSPDAM, DRAFTS) in both accuracy and inference speed. The model supports pixel-level signal segmentation and yields reliable estimates for dispersion measure (DM) and time of arrival (ToA). Large-scale blind searches on CRAFTS data further demonstrate robustness, with an average false positive rate of 0.28% and minimal human verification required. This search has already led to the identification of two pulsar candidates, both confirmed as known pulsars. Processing benchmarks indicate that the model enables real-time searches on a single consumer-grade GPU, making petabyte-scale blind searches feasible. The code is publicly available on GitHub, and the model can be easily integrated with existing tools to automate and streamline radio data analysis beyond FRB or pulsar searches.

astro-ph.GA

Multipartite steering verification with imprecise measurements

Quantum steering is a fundamental quantum correlation that plays a pivotal role in quantum technologies, but its verification crucially relies on precise measurements -- an assumption often undermined by practical imperfections. Here, we investigate multipartite steering verification under imprecise measurements and develop a quantitative method that effectively eliminates false positives induced by measurement imprecision. A comparison with a device-independent approach demonstrates that our method accurately delineates the scope of valid verification. In a special case, our method also enables the verification of multipartite entanglement under nonideal conditions. These results substantially enhance the robustness of multipartite steering and entanglement verification against measurement imprecision, thereby promoting their applicability in realistic quantum technologies.

quant-ph

A Mechanism-Coupled Split Window Network for Medium- to High-Resolution Land Surface Temperature Retrieval

Land surface temperature (LST) is a fundamental physical variable in land-atmosphere interactions, surface energy budgets, and climate processes. LST derived from medium- to high-resolution thermal infrared (TIR) observations effectively reveals thermal environmental disparities across distinct landscape units. However, achieving accurate, robust, and globally generalizable LST retrieval remains challenging under complex atmospheric conditions and diverse land cover types. Traditional split window (SW) algorithms heavily rely on empirical parameterizations, whose fixed coefficients fail to adapt to complex scenarios such as high surface temperatures and high atmospheric water vapor content. Concurrently, conventional data-driven models exhibit limited generalizability to out-of-distribution (OOD) samples due to the absence of explicit physical structure constraints. To address these issues, this study proposes a Parallel Component Decoupled Neural Network (PCD-Net) framework, which reformulates SW retrieval as a dynamic learning problem of physical component coefficients. Using the SW equation as the physical backbone, the framework constructs parallel subnetworks to adaptively learn the dynamic coefficients corresponding to the constant, first-order, and second-order brightness temperature difference terms; meanwhile, a residual branch is incorporated to supplement the nonlinear coupling corrections induced by the joint effects of surface emissivity and atmospheric water vapor. Through this component-level decoupled modeling, PCD-Net explicitly characterizes the dynamic response relationships between land surface emissivity, atmospheric water vapor content, and different SW physical components.

physics.ao-ph

Information entropy of complex probability

Probability theory is fundamental for modeling uncertainty, with traditional probabilities being real and non-negative. Complex probability extends this concept by allowing complex-valued probabilities, opening new avenues for analysis in various fields. This paper explores the information-theoretic aspects of complex probability, focusing on its definition, properties, and applications. We extend Shannon entropy to complex probability and examine key properties, including maximum entropy, joint entropy, conditional entropy, equilibration, and cross entropy. These results offer a framework for understanding entropy in complex probability spaces and have potential applications in fields such as statistical mechanics and information theory.

cs.IT

Asymmetric mode-pairing quantum key distribution

Mode-pairing quantum key distribution (MP-QKD) can surpass the repeaterless rate-transmittance bound (Pirandola-Laurenza-Ottaviani-Banchi bound) without requiring global phase locking, exhibiting remarkable flexibility. However, MP-QKD necessitates equal communication distances in two channels, which is a challenging requirement in practical applications. To address this limitation, we extend the original MP-QKD to asymmetric cases. Our decoy-state estimation confirms that asymmetric channel transmittances and asymmetric intensities do not compromise the security of the protocol. We focus on the pulse-intensity relationship, a key factor for optimizing the performance of asymmetric MP-QKD. Unlike previous asymmetric protocols, the intensities of different bases in asymmetric MP-QKD cannot be decoupled. We introduce an optimal-pulse-intensity method, adaptable to various scenarios, to enhance key rates by calculating ideal pulse intensities. Simulation results in various representative scenarios indicate that our method effectively reduces the impact of asymmetric channel distances on MP-QKD performance, enhancing its practical applicability.

quant-ph

Meta predictive learning model of languages in neural circuits

Large language models based on self-attention mechanisms have achieved astonishing performances not only in natural language itself, but also in a variety of tasks of different nature. However, regarding processing language, our human brain may not operate using the same principle. Then, a debate is established on the connection between brain computation and artificial self-supervision adopted in large language models. One of most influential hypothesis in brain computation is the predictive coding framework, which proposes to minimize the prediction error by local learning. However, the role of predictive coding and the associated credit assignment in language processing remains unknown. Here, we propose a mean-field learning model within the predictive coding framework, assuming that the synaptic weight of each connection follows a spike and slab distribution, and only the distribution, rather than specific weights, is trained. This meta predictive learning is successfully validated on classifying handwritten digits where pixels are input to the network in sequence, and moreover on the toy and real language corpus. Our model reveals that most of the connections become deterministic after learning, while the output connections have a higher level of variability. The performance of the resulting network ensemble changes continuously with data load, further improving with more training data, in analogy with the emergent behavior of large language models. Therefore, our model provides a starting point to investigate the connection among brain computation, next-token prediction and general intelligence.

cs.CL

Statistical mechanics of continual learning: variational principle and mean-field potential

An obstacle to artificial general intelligence is set by continual learning of multiple tasks of different nature. Recently, various heuristic tricks, both from machine learning and from neuroscience angles, were proposed, but they lack a unified theory ground. Here, we focus on continual learning in single-layered and multi-layered neural networks of binary weights. A variational Bayesian learning setting is thus proposed, where the neural networks are trained in a field-space, rather than gradient-ill-defined discrete-weight space, and furthermore, weight uncertainty is naturally incorporated, and modulates synaptic resources among tasks. From a physics perspective, we translate the variational continual learning into Franz-Parisi thermodynamic potential framework, where previous task knowledge acts as a prior and a reference as well. We thus interpret the continual learning of the binary perceptron in a teacher-student setting as a Franz-Parisi potential computation. The learning performance can then be analytically studied with mean-field order parameters, whose predictions coincide with numerical experiments using stochastic gradient descent methods. Based on the variational principle and Gaussian field approximation of internal preactivations in hidden layers, we also derive the learning algorithm considering weight uncertainty, which solves the continual learning with binary weights using multi-layered neural networks, and performs better than the currently available metaplasticity algorithm. Our proposed principled frameworks also connect to elastic weight consolidation, weight-uncertainty modulated learning, and neuroscience inspired metaplasticity, providing a theory-grounded method for the real-world multi-task learning with deep networks.

cond-mat.stat-mech

Emergence of hierarchical modes from deep learning

Large-scale deep neural networks consume expensive training costs, but the training results in less-interpretable weight matrices constructing the networks. Here, we propose a mode decomposition learning that can interpret the weight matrices as a hierarchy of latent modes. These modes are akin to patterns in physics studies of memory networks, but the least number of modes increases only logarithmically with the network width, and becomes even a constant when the width further grows. The mode decomposition learning not only saves a significant large amount of training costs, but also explains the network performance with the leading modes, displaying a striking piecewise power-law behavior. The modes specify a progressively compact latent space across the network hierarchy, making a more disentangled subspaces compared to standard training. Our mode decomposition learning is also studied in an analytic on-line learning setting, which reveals multi-stage of learning dynamics with a continuous specialization of hidden nodes. Therefore, the proposed mode decomposition learning points to a cheap and interpretable route towards the magical deep learning.

cs.LG

Recollision of excited electron in below-threshold nonsequential double ionization

Consensus has been reached that recollision, as the most important post-tunneling process, is responsible for nonsequential double ionization process in intense infrared laser field, however, its effect has been restricted to interaction between the first ionized electron and the residual univalent ion so far. Here we identify the key role of recollision between the second ionized electron and the divalent ion in the below-threshold nonsequential double ionization process by introducing a Coulomb-corrected quantum-trajectories method, which enables us to well reproduce the experimentally observed cross-shaped and anti-correlated patterns in correlated two-electron momentum distributions, and also the transition between these two patterns. Being significantly enhanced relatively by the recapture process, recolliding trajectories of the second electron excited by the first- or third-return recolliding trajectories of the first electron produce the cross-shaped or anti-correlated distributions, respectively. And the transition is induced by the increasing contribution of the third return with increasing pulse duration. Our work provides new insight into atomic ionization dynamics and paves the new way to imaging of ultrafast dynamics of atoms and molecules in intense laser field.

physics.atom-ph

Ensemble perspective for understanding temporal credit assignment

Recurrent neural networks are widely used for modeling spatio-temporal sequences in both nature language processing and neural population dynamics. However, understanding the temporal credit assignment is hard. Here, we propose that each individual connection in the recurrent computation is modeled by a spike and slab distribution, rather than a precise weight value. We then derive the mean-field algorithm to train the network at the ensemble level. The method is then applied to classify handwritten digits when pixels are read in sequence, and to the multisensory integration task that is a fundamental cognitive function of animals. Our model reveals important connections that determine the overall performance of the network. The model also shows how spatio-temporal information is processed through the hyperparameters of the distribution, and moreover reveals distinct types of emergent neural selectivity. To provide a mechanistic analysis of the ensemble learning, we first derive an analytic solution of the learning at the infinitely-large-network limit. We then carry out a low-dimensional projection of both neural and synaptic dynamics, analyze symmetry breaking in the parameter space, and finally demonstrate the role of stochastic plasticity in the recurrent computation. Therefore, our study sheds light on mechanisms of how weight uncertainty impacts the temporal credit assignment in recurrent neural networks from the ensemble perspective.

cond-mat.dis-nn

Estimation of transmitted wavefronts at defocused positions in a broad bandwidth range

Wavefront aberrations can reflect the imaging quality of high-performance optical systems better than geometric aberrations. Although laser interferometers have emerged as the main tool for measurement of transmitted wavefronts, their application is greatly limited, as they are typically designed for operation at specific wavelengths. In a previous study, we proposed a method for determining the wavefront transmitted by an optical system at any wavelength in a certain band. Although this method works well for most monochromatic systems, where the image plane is at the focal point for the transmission wavelength, for general multi-color systems, it is more practical to measure the wavefront at the defocused image plane. Hence, in this paper, we have developed a complete method for determining transmitted wavefronts in a broad bandwidth at any defocused position, enabling wavefront measurements for multi-color systems. Here, we assume that in small ranges, the Zernike coefficients have a linear relationship with position, such that Zernike coefficients at defocused positions can be derived from measurements performed at the focal point. We conducted experiments to verify these assumptions, validating the new method. The experimental setup has been improved so that it can handle multi-color systems, and a detailed experimental process is summarized. With this technique, application of broadband transmission wavefront measurement can be extended to most general optical systems, which is of great significance for characterization of achromatic and apochromatic optical lenses.

physics.optics

Learning credit assignment

Deep learning has achieved impressive prediction accuracies in a variety of scientific and industrial domains. However, the nested non-linear feature of deep learning makes the learning highly non-transparent, i.e., it is still unknown how the learning coordinates a huge number of parameters to achieve a decision making. To explain this hierarchical credit assignment, we propose a mean-field learning model by assuming that an ensemble of sub-networks, rather than a single network, are trained for a classification task. Surprisingly, our model reveals that apart from some deterministic synaptic weights connecting two neurons at neighboring layers, there exist a large number of connections that can be absent, and other connections can allow for a broad distribution of their weight values. Therefore, synaptic connections can be classified into three categories: very important ones, unimportant ones, and those of variability that may partially encode nuisance factors. Therefore, our model learns the credit assignment leading to the decision, and predicts an ensemble of sub-networks that can accomplish the same task, thereby providing insights toward understanding the macroscopic behavior of deep learning through the lens of distinct roles of synaptic weights.

cs.LG

The effect of Coulomb field on laser-induced ultrafast imaging methods

By performing a joint theoretical and experimental investigation on the high-order above-threshold ionization (HATI) spectrum, the dominant role of the 3rd-return-recollision trajectories in the region near the cutoff due to the ionic Coulomb field is identified. This invalidates the key assumption adopted in the conventional laser-induced electron diffraction (LIED) approach that the 1st-returnrecollision trajectories dominate the spectrum according to strong field approximation (SFA). Our results show that the incident (return) electron beams produced by the 1st and 3rd returns possess distinct characteristics of beam energy, beam diameter and temporal evolution law due to the influence of Coulomb field, and therefore the extracted results in the LIED will be altered if the significance of the 3rd-return-recollision trajectories is properly considered in the analysis. Such Coulomb field effect should be taken into account in all kinds of laser-induced imaging schemes based on recollision.

physics.atom-ph