SearcharxivSearch

arXiv subjects

Pei Sun

Publications and source records attributed to Pei Sun.

At least 19 recordsLinked to original sources

Tensor-Network Analysis of Root Patterns in the XXX Model with Open Boundaries

The string hypothesis of Bethe roots is a cornerstone in the thermodynamic analysis of quantum integrable systems, since it connects root configurations with physical quantities such as the ground-state energy, surface energy and excitation spectra. For integrable models with \(U(1)\) symmetry, this connection is well established. When the \(U(1)\) symmetry is broken by generic non-diagonal boundary fields, however, the off-diagonal Bethe Ansatz leads to an inhomogeneous \(T\text{--}Q\) relation whose Bethe roots have highly nontrivial distributions. This raises two fundamental questions: whether the zero roots and the ODBA Bethe roots still possess regular and classifiable structures in the large-size limit, and whether such structures can be used to extract physical quantities. In this work, we address these two questions for the isotropic Heisenberg spin chain with non-diagonal open boundaries. By combining tensor-network algorithms with Bethe-Ansatz techniques, we determine the zero-root and Bethe-root configurations associated with the \(\Lambda\text{--}\theta\) relation and the inhomogeneous Bethe Ansatz equations for large system sizes, up to \(N\simeq 60\) and \(100\). We find that, despite the absence of \(U(1)\) symmetry, the roots exhibit well-organized patterns. The zero roots form bulk strings, boundary strings and additional roots, while the ODBA Bethe roots split into four geometric classes: regular roots, line roots, arc roots and paired-line roots.

math-ph

Fluid-network relations: decay laws meet with spatial self-similarity, scale-invariance, and control scaling

Diverse implicit structures of fluids are discovered lately, providing opportunities to study the physics of fluids applying network analysis. Although considerable works devote to identifying informative network structures of fluids, we have limited understanding about the information these networks convey about fluids. To analyze how fluid mechanics is embodied in network topology or vice versa, we reveal a set of fluid-network relations that quantify the interactions between fundamental fluid properties (e.g., kinetic energy and enstrophy decay laws) and defining network characteristics (e.g., spatial self-similarity, scale-invariance, and control scaling). By analyzing spatial self-similarity in classic and generalized contexts, we first assess the self-similarity of vortical interactions in fluid flows. Deviations from self-similarity in networks exhibit power-law scaling behaviors with respect to fluid properties, suggesting the diversity among vortex as essential to self-similar fluid flows. Then, the same paradigm is adopted to investigate scale-invariance using renormalization groups, which reveals that the breaking extents of scale-invariance in networks, similar to those of spatial self-similarity, also scale with fluid properties in power-law manners. Furthermore, we define a control problem on networks to study the propagation of perturbations through vortical interactions over different ranges. The minimum cost of controlling vortical networks exponentially scales with range diameters (i.e., control distances), whose growth rates experiences temporal decays. We show that this temporal decay speed is fully determined by fluid properties in power-law scaling behaviours. In summary, these fluid-network relations enable a deeper understanding of implicit fluid structures and their interactions with fluid dynamics.

physics.flu-dyn

Exact solution of a quantum integrable system associated with the $G_2$ exceptional Lie algebra

A quantum integrable spin chain model associated with the $G_2$ exceptional Lie algebra is studied. By using the fusion technique, the closed recursive relations among the fused transfer matrices are obtained. These identities allow us to derive the exact energy spectrum and Bethe ansatz equations of the system based on polynomial analysis. The present method provides a unified treatment to investigate the Bethe ansatz solutions for both periodic and non-diagonal open boundary conditions associated with exceptional Lie algebras.

math-ph

PVTransformer: Point-to-Voxel Transformer for Scalable 3D Object Detection

3D object detectors for point clouds often rely on a pooling-based PointNet to encode sparse points into grid-like voxels or pillars. In this paper, we identify that the common PointNet design introduces an information bottleneck that limits 3D object detection accuracy and scalability. To address this limitation, we propose PVTransformer: a transformer-based point-to-voxel architecture for 3D detection. Our key idea is to replace the PointNet pooling operation with an attention module, leading to a better point-to-voxel aggregation function. Our design respects the permutation invariance of sparse 3D points while being more expressive than the pooling-based PointNet. Experimental results show our PVTransformer achieves much better performance compared to the latest 3D object detectors. On the widely used Waymo Open Dataset, our PVTransformer achieves state-of-the-art 76.5 mAPH L2, outperforming the prior art of SWFormer by +1.7 mAPH L2.

cs.CV

Fast renormalizing the structures and dynamics of ultra-large systems via random renormalization group

Criticality and symmetry, studied by the renormalization groups, lie at the heart of modern physics theories of matters and complex systems. However, surveying these properties with massive experimental data is bottlenecked by the intolerable costs of computing renormalization groups on real systems. Here, we develop a time- and memory-efficient framework, termed as the random renormalization group, for renormalizing ultra-large systems (e.g., with millions of units) within minutes. This framework is based on random projections, hashing techniques, and kernel representations, which support the renormalization governed by linear and non-linear correlations. For system structures, it exploits the correlations among local topology in kernel spaces to unfold the connectivity of units, identify intrinsic system scales, and verify the existences of symmetries under scale transformation. For system dynamics, it renormalizes units into correlated clusters to analyze scaling behaviours, validate scaling relations, and investigate potential criticality. Benefiting from hashing-function-based designs, our framework significantly reduces computational complexity compared with classic renormalization groups, realizing a single-step acceleration of two orders of magnitude. Meanwhile, the efficient representation of different kinds of correlations in kernel spaces realized by random projections ensures the capacity of our framework to capture diverse unit relations. As shown by our experiments, the random renormalization group helps identify non-equilibrium phase transitions, criticality, and symmetry in diverse large-scale genetic, neural, material, social, and cosmological systems.

cond-mat.stat-mech

Exact surface energy of the Hubbard model with nonparallel boundary magnetic fields

In this study, we explore the precise physical quantities in the thermodynamic limit of the one-dimensional Hubbard model with nonparallel boundary magnetic fields based on the off-diagonal Bethe ansatz solution. A particular emphasis is placed on the half-filling condition to investigate the distinct patterns of Bethe roots in the reduced Bethe ansatz equations for different boundary parameters. The ground state of the system can be divided into five regions according to the distribution of Bethe roots. By analyzing these patterns, we calculate the densities of states, ground-state energy density, and surface energy. The results reveal the existence of stableboundary bound states, which are dependent on specific constraints regarding the boundary magnetic fields.

math-ph

Gemini: A Family of Highly Capable Multimodal Models

This report introduces a new family of multimodal models, Gemini, that exhibit remarkable capabilities across image, audio, video, and text understanding. The Gemini family consists of Ultra, Pro, and Nano sizes, suitable for applications ranging from complex reasoning tasks to on-device memory-constrained use-cases. Evaluation on a broad range of benchmarks shows that our most-capable Gemini Ultra model advances the state of the art in 30 of 32 of these benchmarks - notably being the first model to achieve human-expert performance on the well-studied exam benchmark MMLU, and improving the state of the art in every one of the 20 multimodal benchmarks we examined. We believe that the new capabilities of the Gemini family in cross-modal reasoning and language understanding will enable a wide variety of use cases. We discuss our approach toward post-training and deploying Gemini models responsibly to users through services including Gemini, Gemini Advanced, Google AI Studio, and Cloud Vertex AI.

cs.CL

LEF: Late-to-Early Temporal Fusion for LiDAR 3D Object Detection

We propose a late-to-early recurrent feature fusion scheme for 3D object detection using temporal LiDAR point clouds. Our main motivation is fusing object-aware latent embeddings into the early stages of a 3D object detector. This feature fusion strategy enables the model to better capture the shapes and poses for challenging objects, compared with learning from raw points directly. Our method conducts late-to-early feature fusion in a recurrent manner. This is achieved by enforcing window-based attention blocks upon temporally calibrated and aligned sparse pillar tokens. Leveraging bird's eye view foreground pillar segmentation, we reduce the number of sparse history features that our model needs to fuse into its current frame by 10$\times$. We also propose a stochastic-length FrameDrop training technique, which generalizes the model to variable frame lengths at inference for improved performance without retraining. We evaluate our method on the widely adopted Waymo Open Dataset and demonstrate improvement on 3D object detection against the baseline model, especially for the challenging category of large objects.

cs.CV

Theoretical foundations of studying criticality in the brain

Criticality is hypothesized as a physical mechanism underlying efficient transitions between cortical states and remarkable information processing capacities in the brain. While considerable evidence generally supports this hypothesis, non-negligible controversies persist regarding the ubiquity of criticality in neural dynamics and its role in information processing. Validity issues frequently arise during identifying potential brain criticality from empirical data. Moreover, the functional benefits implied by brain criticality are frequently misconceived or unduly generalized. These problems stem from the non-triviality and immaturity of the physical theories that analytically derive brain criticality and the statistic techniques that estimate brain criticality from empirical data. To help solve these problems, we present a systematic review and reformulate the foundations of studying brain criticality, i.e., ordinary criticality (OC), quasi-criticality (qC), self-organized criticality (SOC), and self-organized quasi-criticality (SOqC), using the terminology of neuroscience. We offer accessible explanations of the physical theories and statistic techniques of brain criticality, providing step-by-step derivations to characterize neural dynamics as a physical system with avalanches. We summarize error-prone details and existing limitations in brain criticality analysis and suggest possible solutions. Moreover, we present a forward-looking perspective on how optimizing the foundations of studying brain criticality can deepen our understanding of various neuroscience questions.

q-bio.NC

Simplex path integral and simplex renormalization group for high-order interactions

Modern theories of phase transitions and scale-invariance are rooted in path integral formulation and renormalization group (RG). Despite the applicability of these approaches on simple systems with only pairwise interactions, they are less effective on complex systems with un-decomposable high-order interactions (i.e., interactions among arbitrary sets of units). To precisely characterize the universality of high-order interacting systems, we propose simplex path integral and simplex renormalization group (SRG) as the generalizations of classic approaches to arbitrary high-order and heterogeneous interactions. We first formalize the trajectories of units governed by high-order interactions to define path integrals on corresponding simplices based on a high-order propagator. Then we develop a method to integrate out short-range high-order interactions in the momentum space, accompanied by a coarse graining procedure functioning on the simplex structure generated by high-order interactions. The proposed SRG, equipped with a divide-and-conquer framework, can deal with the absence of ergodicity arised from the sparse distribution of high-order interactions and renormalize a system with intertwined high-order interactions on the $p$-order according to its properties on the $q$-order ($p\leq q$). The associated scaling relation and its corollaries support to differentiate among scale-invariant, weakly scale-invariant, and scale-dependent systems across different orders. We have validated our theory in multi-order scale-invariance verification, topological invariance discovery, organizational structure identification, and information bottleneck analysis. These experiments demonstrate the capacity of our theory for identifying intrinsic statistical and topological properties of high-order interacting systems during system reduction.

cond-mat.stat-mech

WOMD-LiDAR: Raw Sensor Dataset Benchmark for Motion Forecasting

Widely adopted motion forecasting datasets substitute the observed sensory inputs with higher-level abstractions such as 3D boxes and polylines. These sparse shapes are inferred through annotating the original scenes with perception systems' predictions. Such intermediate representations tie the quality of the motion forecasting models to the performance of computer vision models. Moreover, the human-designed explicit interfaces between perception and motion forecasting typically pass only a subset of the semantic information present in the original sensory input. To study the effect of these modular approaches, design new paradigms that mitigate these limitations, and accelerate the development of end-to-end motion forecasting models, we augment the Waymo Open Motion Dataset (WOMD) with large-scale, high-quality, diverse LiDAR data for the motion forecasting task. The new augmented dataset WOMD-LiDAR consists of over 100,000 scenes that each spans 20 seconds, consisting of well-synchronized and calibrated high quality LiDAR point clouds captured across a range of urban and suburban geographies (https://waymo.com/open/data/motion/). Compared to Waymo Open Dataset (WOD), WOMD-LiDAR dataset contains 100x more scenes. Furthermore, we integrate the LiDAR data into the motion forecasting model training and provide a strong baseline. Experiments show that the LiDAR data brings improvement in the motion forecasting task. We hope that WOMD-LiDAR will provide new opportunities for boosting end-to-end motion forecasting models.

cs.CV

Laplacian dynamics of convergent and divergent swarm behaviors

Swarming phenomena are ubiquitous in various physical, biological, and social systems, where simple local interactions between individual units lead to complex global patterns. A common feature of diverse swarming phenomena is that the units exhibit either convergent or divergent evolution in their behaviors, i.e., becoming increasingly similar or distinct, respectively. The associated dynamics changes across time, leading to complex consequences on a global scale. In this study, we propose a generalized Laplacian dynamics model to describe both convergent and divergent swarm behaviors, where the trends of convergence and divergence compete with each other and jointly determine the evolution of global patterns. We empirically observe non-trivial phase-transition-like phenomena between the convergent and divergent evolution phases, which are controlled by local interaction properties. We also propose a conjecture regarding the underlying phase transition mechanisms and outline the main theoretical difficulties for testing this conjecture. Overall, our framework may serve as a minimal model of swarm behaviors and their intricate phase transition dynamics.

cond-mat.stat-mech

Thermodynamics of percolation in interacting systems

Interacting systems can be studied as the networks where nodes are system units and edges denote correlated interactions. Although percolation on network is a unified way to model the emergence and propagation of correlated behaviours, it remains unknown how the dynamics characterized by percolation is related to the thermodynamics of phase transitions. It is non-trivial to formalize thermodynamics for most complex systems, not to mention calculating thermodynamic quantities and verifying scaling relations during percolation. In this work, we develop a formalism to quantify the thermodynamics of percolation in interacting systems, which is rooted in a discovery that percolation transition is a process for the system to lose the freedom degrees associated with ground state configurations. We derive asymptotic formulas to accurately calculate entropy and specific heat under our framework, which enables us to detect phase transitions and demonstrate the Rushbrooke equality (i.e., $\alpha+2\beta+\gamma=2$) in six representative complex systems (e.g., Bernoulli and bootstrap percolation, classical and quantum synchronization, non-linear oscillations with damping, and cellular morphogenesis). These results suggest the general applicability of our framework in analyzing diverse interacting systems and percolation processes.

cond-mat.stat-mech

Koopman neural operator as a mesh-free solver of non-linear partial differential equations

The lacking of analytic solutions of diverse partial differential equations (PDEs) gives birth to a series of computational techniques for numerical solutions. Although numerous latest advances are accomplished in developing neural operators, a kind of neural-network-based PDE solver, these solvers become less accurate and explainable while learning long-term behaviors of non-linear PDE families. In this paper, we propose the Koopman neural operator (KNO), a new neural operator, to overcome these challenges. With the same objective of learning an infinite-dimensional mapping between Banach spaces that serves as the solution operator of the target PDE family, our approach differs from existing models by formulating a non-linear dynamic system of equation solution. By approximating the Koopman operator, an infinite-dimensional operator governing all possible observations of the dynamic system, to act on the flow mapping of the dynamic system, we can equivalently learn the solution of a non-linear PDE family by solving simple linear prediction problems. We validate the KNO in mesh-independent, long-term, and5zero-shot predictions on five representative PDEs (e.g., the Navier-Stokes equation and the Rayleigh-B{\'e}nard convection) and three real dynamic systems (e.g., global water vapor patterns and western boundary currents). In these experiments, the KNO exhibits notable advantages compared with previous state-of-the-art models, suggesting the potential of the KNO in supporting diverse science and engineering applications (e.g., PDE solving, turbulence modelling, and precipitation forecasting).

cs.LG

KoopmanLab: machine learning for solving complex physics equations

Numerous physics theories are rooted in partial differential equations (PDEs). However, the increasingly intricate physics equations, especially those that lack analytic solutions or closed forms, have impeded the further development of physics. Computationally solving PDEs by classic numerical approaches suffers from the trade-off between accuracy and efficiency and is not applicable to the empirical data generated by unknown latent PDEs. To overcome this challenge, we present KoopmanLab, an efficient module of the Koopman neural operator family, for learning PDEs without analytic solutions or closed forms. Our module consists of multiple variants of the Koopman neural operator (KNO), a kind of mesh-independent neural-network-based PDE solvers developed following dynamic system theory. The compact variants of KNO can accurately solve PDEs with small model sizes while the large variants of KNO are more competitive in predicting highly complicated dynamic systems govern by unknown, high-dimensional, and non-linear PDEs. All variants are validated by mesh-independent and long-term prediction experiments implemented on representative PDEs (e.g., the Navier-Stokes equation and the Bateman-Burgers equation in fluid mechanics) and ERA5 (i.e., one of the largest high-resolution global-scale climate data sets in earth physics). These demonstrations suggest the potential of KoopmanLab to be a fundamental tool in diverse physics studies related to equations or dynamic systems.

cs.LG

Statistical Physics of Deep Neural Networks: Initialization toward Optimal Channels

In deep learning, neural networks serve as noisy channels between input data and its representation. This perspective naturally relates deep learning with the pursuit of constructing channels with optimal performance in information transmission and representation. While considerable efforts are concentrated on realizing optimal channel properties during network optimization, we study a frequently overlooked possibility that neural networks can be initialized toward optimal channels. Our theory, consistent with experimental validation, identifies primary mechanics underlying this unknown possibility and suggests intrinsic connections between statistical physics and deep learning. Unlike the conventional theories that characterize neural networks applying the classic mean-filed approximation, we offer analytic proof that this extensively applied simplification scheme is not valid in studying neural networks as information channels. To fill this gap, we develop a corrected mean-field framework applicable for characterizing the limiting behaviors of information propagation in neural networks without strong assumptions on inputs. Based on it, we propose an analytic theory to prove that mutual information maximization is realized between inputs and propagated signals when neural networks are initialized at dynamic isometry, a case where information transmits via norm-preserving mappings. These theoretical predictions are validated by experiments on real neural networks, suggesting the robustness of our theory against finite-size effects. Finally, we analyze our findings with information bottleneck theory to confirm the precise relations among dynamic isometry, mutual information maximization, and optimal channel properties in deep learning.

cs.LG

LidarAugment: Searching for Scalable 3D LiDAR Data Augmentations

Data augmentations are important in training high-performance 3D object detectors for point clouds. Despite recent efforts on designing new data augmentations, perhaps surprisingly, most state-of-the-art 3D detectors only use a few simple data augmentations. In particular, different from 2D image data augmentations, 3D data augmentations need to account for different representations of input data and require being customized for different models, which introduces significant overhead. In this paper, we resort to a search-based approach, and propose LidarAugment, a practical and effective data augmentation strategy for 3D object detection. Unlike previous approaches where all augmentation policies are tuned in an exponentially large search space, we propose to factorize and align the search space of each data augmentation, which cuts down the 20+ hyperparameters to 2, and significantly reduces the search complexity. We show LidarAugment can be customized for different model architectures with different input representations by a simple 2D grid search, and consistently improve both convolution-based UPillars/StarNet/RSN and transformer-based SWFormer. Furthermore, LidarAugment mitigates overfitting and allows us to scale up 3D detectors to much larger capacity. In particular, by combining with latest 3D detectors, our LidarAugment achieves a new state-of-the-art 74.8 mAPH L2 on Waymo Open Dataset.

cs.CV