SearcharxivSearch

arXiv subjects

Zhenyu Zhu

Publications and source records attributed to Zhenyu Zhu.

At least 19 recordsLinked to original sources

Constraining the Equation of State of Neutron Stars with third-generation Gravitational Wave detectors

We investigated the impact of the number of binary neutron star merger events and neutron star (NS) mass distributions on constraining the equation of state (EoS), tidal deformability and radius of NSs, as well as the nuclear parameters, using binary neutron star inspiral gravitational wave signals with third-generation detectors. We generate simulated gravitational wave signals and compute the Fisher information matrix for each event after the number of events and mass distribution models are given. The covariance of EoS parameters is obtained by holding the non-EoS waveform parameters fixed at their injected values and accumulating the Fisher matrix over all events. Finally, the posterior samples of EoS, tidal deformability, radius and nuclear parameters are generated based on this covariance matrix. We find that larger number of events lead to tighter constraints due to more observed events and data accumulation, as expected given the increased number of detections. Meanwhile, we note that the mass distribution plays a more important role in constraining the EoS. We compare a realistic bimodal Gaussian distribution, a uniform distribution and another uniform including sub-solar mass NSs. The results show that the uniform distribution yields tighter constraints than bimodal models because it includes more low-mass and massive NSs. This also implies that events near $1.4\,\Msun$ provide partly redundant information about the EoS. Additionally, including sub-solar mass NSs can further improve the constraints by significantly reducing the EoS uncertainties at sub-saturation densities and inner-crust, highlighting the importance of sub-saturation density EoS.

astro-ph.HE

Detectability of bulk viscosity effects on post-merger gravitational wave signals from binary neutron star mergers

We investigate bulk viscosity (BV) effects on post-merger gravitational wave (PMGW) signals from binary neutron star (BNS) mergers and their detectability using realistic equation of state (EoS) posteriors constrained by recent progress in gravitational-wave (GW) and X-ray observations, as well as neutron skin thickness measurements. A linear fitting formula based on recent BNS simulations with BV is used to estimate the peak frequency shift of the PMGWs caused by BV effects. Subsequently, we evaluate this peak frequency shift for our set of EoS posterior samples. We use the Fisher information matrix and the dataset of simulated observable events from the ET and CE detector network in an optimistic scenario to estimate the measurement accuracy of the peak frequency. We find that BV effects on PMGWs may be detectable only for EoS models with large values of the symmetry energy slope ($L_{\rm sym}$), as favored by neutron-skin thickness measurements. Thus, the detection of BV effects on PMGWs itself can provide abundant information about the symmetry energy and its slope. However, BV effects on PMGWs are generally weak and may be observed only in some extreme and optimistic cases. Therefore, observations of PMGWs and their analyses may not be able to provide very accurate measurements of the symmetry energy through BV effects.

astro-ph.HE

Limitations in constraining neutron star radii and nuclear properties from inspiral gravitational wave detections

We investigate the constraints on the neutron star equation of state (EoS) and nuclear properties achievable with third-generation gravitational wave detectors using the Fisher information matrix approach within the relativistic mean field (RMF) theory. Assuming an optimistic binary neutron star (BNS) merger rate, we generate simulated inspiral gravitational wave (GW) signals corresponding to one year of observation. From these simulated data, we compute the covariance matrix and posterior distributions for nuclear properties and EoS. Our results show that the EoS can be tightly constrained, particularly in the density range between one and four times nuclear saturation density. However, due to the scarcity of low-mass neutron stars in the GW sample, the EoS at sub-saturation densities remains poorly constrained. Thus, in turn, leads to weaker constraints on neutron star radii, as the radii are sensitive to the low-density EoS. Additionally, we present the expected correlations among nuclear parameters in general and plots of the inferred symmetry energy in particular, which represent degeneracies in their influence on the EoS and make them difficult to be constrained through GW observations alone. These highlights inherent limitations of inspiral GW signals in probing dense matter properties. Therefore, precise radius measurements, post-merger GW observations, and supplementary constraints from terrestrial nuclear experiments remain essential for a comprehensive understanding of dense matter.

astro-ph.HE

Reconstruction of fast-rotating neutron star observables with the neural network

Rotation can significantly affect neutron-star (NS) properties, but accurate modeling of rapidly rotating NSs requires solving a two-dimensional, axially symmetric system, making traditional calculations too expensive for inference analyses that demand a large amount of model evaluations. We develop a causal convolutional neural networks that preserve the chronological-like dependence of NS properties on the equation of state (EoS) and rapidly reconstruct observables for static, Keplerian, and rotating configurations. Using \texttt{RNS}, we generate a dataset of NS observables and use it to train our networks. We validate our networks with three representative EoS (SFHo, SLy4, and DD2) and find that the they accurately reproduce the \texttt{RNS} results. The trained networks evaluate NS configurations for a single EoS in $\sim 50$ms, providing a substantial speedup over typical \texttt{RNS} runtimes of $\sim 30$ min and enabling efficient inference analyses involving rapidly rotating NSs.

astro-ph.HE

Subgrid Mean-field Dynamo Model with Dynamical Quenching in General Relativistic Magnetohydrodynamic Simulations

Large-scale magnetic fields are relevant for a number of dynamical processes in accretion disks, including driving turbulence, reconnection events, and launching outflows. Numerical simulations have indicated that the initial strengths and configurations of the large-scale magnetic fields have a direct imprint on the outcome of an accretion disk evolution. To facilitate future self-consistent simulations that include intrinsic dynamo processes, we derive and implement a subgrid model of a helical large-scale dynamo with dynamical quenching in general-relativistic resistive magnetohydrodynamical simulations of geometrically thin accretion disks. By incorporating previous numerical and analytical results of helical dynamos, our model features only one input parameter, the viscosity parameter $α_\text{SS}$. We demonstrate that our model can reproduce butterfly diagrams seen in previous local and global simulations. With rather aggressive parameter choice of $α_\text{SS}=0.02$ and black hole spin $a_\text{BH}=0.9375$, our thin-disk model launches weak collimated polar outflows with Lorentz factor $\simeq 1.2$, but no polar outflow is present with less vigorous turbulence or less positive $a_\text{BH}$. With negative $a_\text{BH}$, we find the field configurations to appear more similar to Newtonian cases, whereas for positive $a_\text{BH}$, the poloidal field loops become distorted and the cycle period becomes sporadic or even disappears. Moreover, we demonstrate how $α_\text{SS}$ can avoid to be prescribed and instead be determined by the local plasma beta. Such a fully dynamical subgrid dynamo allows for self-consistent amplification of the large-scale magnetic fields.

astro-ph.HE

Three-dimensional equation of state extension of quark matter in Fermi-liquid theory

The cold, dense matter equation of state (EoS) determines crucial global properties of neutron stars (NSs), including the mass, radius and tidal deformability. However, a one-dimensional (1D), cold, and $β$-equilibrated EoS is insufficient to fully describe the interactions or capture the dynamical processes of dense matter as realized in binary neutron star (BNS) mergers or core-collapse supernovae (CCSNe), where thermal and out-of-equilibrium effects play important roles. We develop a method to self-consistently extend a 1D cold and $β$-equilibrated EoS of quark matter to a full three-dimensional (3D) version, accounting for density, temperature, and electron fraction dependencies, within the framework of Fermi-liquid theory (FLT), incorporating both thermal and out-of-equilibrium contributions. We compare our FLT-extended EoS with the original bag model and find that our approach successfully reproduces the contributions of thermal and compositional dependencies of the 3D EoS. Furthermore, we construct a 3D EoS with a first-order phase transition (PT) by matching our 3D FLT-extended quark matter EoS to the hadronic DD2 EoS under Maxwell construction, and test it through the GRHD simulations of the TOV-star and CCSN explosion. Both simulations produce consistent results with previous studies, demonstrating the effectiveness and robustness of our 3D EoS construction with PT.

astro-ph.HE

Equation of State of Decompressed Quark Matter, and Observational Signatures of Quark-Star Mergers

Quark stars are challenging to confirm or exclude observationally because they can have similar masses and radii as neutron stars. By performing the first calculation of the non-equilibrium equation of state of decompressed quark matter at finite temperature, we determine the properties of the ejecta from binary quark-star or quark star-black hole mergers. We account for all relevant physical processes during the ejecta evolution, including quark nugget evaporation and cooling, and weak interactions. We find that these merger ejecta can differ significantly from those in neutron star mergers, depending on the binding energy of quark matter. For relatively high binding energies, quark star mergers are unlikely to produce r-process elements and kilonova signals. We propose that future observations of binary mergers and kilonovae could impose stringent constraints on the binding energy of quark matter and the existence of quark stars.

astro-ph.HE

Training Deep Learning Models with Norm-Constrained LMOs

In this work, we study optimization methods that leverage the linear minimization oracle (LMO) over a norm-ball. We propose a new stochastic family of algorithms that uses the LMO to adapt to the geometry of the problem and, perhaps surprisingly, show that they can be applied to unconstrained problems. The resulting update rule unifies several existing optimization methods under a single framework. Furthermore, we propose an explicit choice of norm for deep architectures, which, as a side benefit, leads to the transferability of hyperparameters across model sizes. Experimentally, we demonstrate significant speedups on nanoGPT training using our algorithm, Scion, without any reliance on Adam. The proposed method is memory-efficient, requiring only one set of model weights and one set of gradients, which can be stored in half-precision. The code is available at https://github.com/LIONS-EPFL/scion .

cs.LG

Multi-Step Alignment as Markov Games: An Optimistic Online Gradient Descent Approach with Convergence Guarantees

Reinforcement Learning from Human Feedback (RLHF) has been highly successful in aligning large language models with human preferences. While prevalent methods like DPO have demonstrated strong performance, they frame interactions with the language model as a bandit problem, which limits their applicability in real-world scenarios where multi-turn conversations are common. Additionally, DPO relies on the Bradley-Terry model assumption, which does not adequately capture the non-transitive nature of human preferences. In this paper, we address these challenges by modeling the alignment problem as a two-player constant-sum Markov game, where each player seeks to maximize their winning rate against the other across all steps of the conversation. Our approach Optimistic Multi-step Preference Optimization (OMPO) is built upon the optimistic online mirror descent algorithm~\citep{rakhlin2013online,joulani17a}. Theoretically, we provide a rigorous analysis for the convergence of OMPO and show that OMPO requires $\mathcal{O}(ε^{-1})$ policy updates to converge to an $ε$-approximate Nash equilibrium. We also validate the effectiveness of our method on multi-turn conversations dataset and math reasoning dataset.

cs.LG

LREF: A Novel LLM-based Relevance Framework for E-commerce

Query and product relevance prediction is a critical component for ensuring a smooth user experience in e-commerce search. Traditional studies mainly focus on BERT-based models to assess the semantic relevance between queries and products. However, the discriminative paradigm and limited knowledge capacity of these approaches restrict their ability to comprehend the relevance between queries and products fully. With the rapid advancement of Large Language Models (LLMs), recent research has begun to explore their application to industrial search systems, as LLMs provide extensive world knowledge and flexible optimization for reasoning processes. Nonetheless, directly leveraging LLMs for relevance prediction tasks introduces new challenges, including a high demand for data quality, the necessity for meticulous optimization of reasoning processes, and an optimistic bias that can result in over-recall. To overcome the above problems, this paper proposes a novel framework called the LLM-based RElevance Framework (LREF) aimed at enhancing e-commerce search relevance. The framework comprises three main stages: supervised fine-tuning (SFT) with Data Selection, Multiple Chain of Thought (Multi-CoT) tuning, and Direct Preference Optimization (DPO) for de-biasing. We evaluate the performance of the framework through a series of offline experiments on large-scale real-world datasets, as well as online A/B testing. The results indicate significant improvements in both offline and online metrics. Ultimately, the model was deployed in a well-known e-commerce application, yielding substantial commercial benefits.

cs.IR

Generalization Properties of NAS under Activation and Skip Connection Search

Neural Architecture Search (NAS) has fostered the automatic discovery of state-of-the-art neural architectures. Despite the progress achieved with NAS, so far there is little attention to theoretical guarantees on NAS. In this work, we study the generalization properties of NAS under a unifying framework enabling (deep) layer skip connection search and activation function search. To this end, we derive the lower (and upper) bounds of the minimum eigenvalue of the Neural Tangent Kernel (NTK) under the (in)finite-width regime using a certain search space including mixed activation functions, fully connected, and residual neural networks. We use the minimum eigenvalue to establish generalization error bounds of NAS in the stochastic gradient descent training. Importantly, we theoretically and experimentally show how the derived results can guide NAS to select the top-performing architectures, even in the case without training, leading to a train-free algorithm based on our theory. Accordingly, our numerical validation shed light on the design of computationally efficient methods for NAS. Our analysis is non-trivial due to the coupling of various architectures and activation functions under the unifying framework and has its own interest in providing the lower bound of the minimum eigenvalue of NTK in deep learning theory.

cs.LG

Initialization Matters: Privacy-Utility Analysis of Overparameterized Neural Networks

We analytically investigate how over-parameterization of models in randomized machine learning algorithms impacts the information leakage about their training data. Specifically, we prove a privacy bound for the KL divergence between model distributions on worst-case neighboring datasets, and explore its dependence on the initialization, width, and depth of fully connected neural networks. We find that this KL privacy bound is largely determined by the expected squared gradient norm relative to model parameters during training. Notably, for the special setting of linearized network, our analysis indicates that the squared gradient norm (and therefore the escalation of privacy loss) is tied directly to the per-layer variance of the initialization distribution. By using this analysis, we demonstrate that privacy bound improves with increasing depth under certain initializations (LeCun and Xavier), while degrades with increasing depth under other initializations (He and NTK). Our work reveals a complex interplay between privacy and depth that depends on the chosen initialization distribution. We further prove excess empirical risk bounds under a fixed KL privacy budget, and show that the interplay between privacy utility trade-off and depth is similarly affected by the initialization.

stat.ML

Sample Complexity Bounds for Score-Matching: Causal Discovery and Generative Modeling

This paper provides statistical sample complexity bounds for score-matching and its applications in causal discovery. We demonstrate that accurate estimation of the score function is achievable by training a standard deep ReLU neural network using stochastic gradient descent. We establish bounds on the error rate of recovering causal relationships using the score-matching-based causal discovery method of Rolland et al. [2022], assuming a sufficiently good estimation of the score function. Finally, we analyze the upper bound of score-matching estimation within the score-based generative modeling, which has been applied for causal discovery but is also of independent interest within the domain of generative models.

cs.LG

Equation of state of nuclear matter and neutron stars: Quark mean-field model versus relativistic mean-field model

The equation of state of neutron-rich nuclear matter is of interest to both nuclear physics and astrophysics. We have demonstrated the consistency between laboratory and astrophysical nuclear matter in neutron stars by considering low-density nuclear physics constraints (from $^{208}$Pb neutron-skin thickness) and high-density astrophysical constraints (from neutron star global properties). We have used both quark-level and hadron-level models, taking the quark mean-field (QMF) model and the relativistic mean-field (RMF) model as examples, respectively. We have constrained the equation of states of neutron stars and some key nuclear matter parameters within the Bayesian statistical approach, using the first multi-messenger event GW170817/AT 2017gfo, as well as the mass-radius simultaneous measurements of PSR J0030+0451 and PSR J0740+6620 from NICER, and the neutron-skin thickness of $^{208}$Pb from both PREX-II measurement and ab initio calculations. Our results show that, compared with the RMF model, QMF model's direct coupling of quarks with mesons and gluons leads to the evolution of the in-medium nucleon mass with the quark mass correction. This feature enables QMF model a wider range of model applicability, as shown by a slow drop of the nucleon mass with density and a large value at saturation that is jointly constrained by nuclear physics and astronomy.

nucl-th

Robust Failure Diagnosis of Microservice System through Multimodal Data

Automatic failure diagnosis is crucial for large microservice systems. Currently, most failure diagnosis methods rely solely on single-modal data (i.e., using either metrics, logs, or traces). In this study, we conduct an empirical study using real-world failure cases to show that combining these sources of data (multimodal data) leads to a more accurate diagnosis. However, effectively representing these data and addressing imbalanced failures remain challenging. To tackle these issues, we propose DiagFusion, a robust failure diagnosis approach that uses multimodal data. It leverages embedding techniques and data augmentation to represent the multimodal data of service instances, combines deployment data and traces to build a dependency graph, and uses a graph neural network to localize the root cause instance and determine the failure type. Our evaluations using real-world datasets show that DiagFusion outperforms existing methods in terms of root cause instance localization (improving by 20.9% to 368%) and failure type determination (improving by 11.0% to 169%).

cs.SE

Benign Overfitting in Deep Neural Networks under Lazy Training

This paper focuses on over-parameterized deep neural networks (DNNs) with ReLU activation functions and proves that when the data distribution is well-separated, DNNs can achieve Bayes-optimal test error for classification while obtaining (nearly) zero-training error under the lazy training regime. For this purpose, we unify three interrelated concepts of overparameterization, benign overfitting, and the Lipschitz constant of DNNs. Our results indicate that interpolating with smoother functions leads to better generalization. Furthermore, we investigate the special case where interpolating smooth ground-truth functions is performed by DNNs under the Neural Tangent Kernel (NTK) regime for generalization. Our result demonstrates that the generalization error converges to a constant order that only depends on label noise and initialization noise, which theoretically verifies benign overfitting. Our analysis provides a tight lower bound on the normalized margin under non-smooth activation functions, as well as the minimum eigenvalue of NTK under high-dimensional settings, which has its own interest in learning theory.

cs.LG

Robustness in deep learning: The good (width), the bad (depth), and the ugly (initialization)

We study the average robustness notion in deep neural networks in (selected) wide and narrow, deep and shallow, as well as lazy and non-lazy training settings. We prove that in the under-parameterized setting, width has a negative effect while it improves robustness in the over-parameterized setting. The effect of depth closely depends on the initialization and the training mode. In particular, when initialized with LeCun initialization, depth helps robustness with the lazy training regime. In contrast, when initialized with Neural Tangent Kernel (NTK) and He-initialization, depth hurts the robustness. Moreover, under the non-lazy training regime, we demonstrate how the width of a two-layer ReLU network benefits robustness. Our theoretical developments improve the results by [Huang et al. NeurIPS21; Wu et al. NeurIPS21] and are consistent with [Bubeck and Sellke NeurIPS21; Bubeck et al. COLT21].

cs.LG

A Bayesian inference of relativistic mean-field model for neutron star matter from observation of NICER and GW170817/AT2017gfo

The observations of optical and near-infrared counterparts of binary neutron star mergers not only enrich our knowledge about the abundance of heavy elements in the Universe, or help reveal the remnant object just after the merger as generally known, but also can effectively constrain dense nuclear matter properties and the equation of state (EOS) in the interior of the merging stars. Following the relativistic mean-field description of nuclear matter, we perform the Bayesian inference of the EOS and the nuclear matter properties using the first multi-messenger event GW170817/AT2017gfo, together with the NICER mass-radius measurements of pulsars. The kilonova is described by a radiation-transfer model with the dynamical ejecta, and light curves connect with the EOS through the quasi-universal relations between the ejecta properties (the ejected mass, velocity, opacity or electron fraction) and binary parameters (the mass ratio and reduced tidal deformability). It is found that the posterior distributions of the reduced tidal deformability from the AT2017gfo analysis display a bimodal structure, with the first peak enhanced by the GW170817 data, leading to slightly softened posterior EOSs, while the second peak cannot be achieved by a nuclear EOS with saturation properties in their empirical ranges. The inclusion of NICER data in our analyses results in stiffened EOS posterior because of the massive pulsar PSR J0740+6620. We give results at nuclear saturation density for the nuclear incompressibility, the symmetry energy and its slope, as well as the nucleon effective mass, from our analysis of those observational data.

astro-ph.HE