Searcharxiv⌕ Search

arXiv subjects

Hui Sun

Publications and source records attributed to Hui Sun.

At least 91 records · Page 5Linked to original sources

Parameter Estimation for the Truncated KdV Model through a Direct Filter Method

In this work, we develop a computational method that to provide realtime detection for water bottom topography based on observations on surface measurements, and we design an inverse problem to achieve this task. The forward model that we use to describe the feature of water surface is the truncated KdV equation, and we formulate the inversion mechanism as an online parameter estimation problem, which is solved by a direct filter method. Numerical experiments are carried out to show that our method can effectively detect abrupt changes of water depth.

math.NA↗

STU-Net: Scalable and Transferable Medical Image Segmentation Models Empowered by Large-Scale Supervised Pre-training

Large-scale models pre-trained on large-scale datasets have profoundly advanced the development of deep learning. However, the state-of-the-art models for medical image segmentation are still small-scale, with their parameters only in the tens of millions. Further scaling them up to higher orders of magnitude is rarely explored. An overarching goal of exploring large-scale models is to train them on large-scale medical segmentation datasets for better transfer capacities. In this work, we design a series of Scalable and Transferable U-Net (STU-Net) models, with parameter sizes ranging from 14 million to 1.4 billion. Notably, the 1.4B STU-Net is the largest medical image segmentation model to date. Our STU-Net is based on nnU-Net framework due to its popularity and impressive performance. We first refine the default convolutional blocks in nnU-Net to make them scalable. Then, we empirically evaluate different scaling combinations of network depth and width, discovering that it is optimal to scale model depth and width together. We train our scalable STU-Net models on a large-scale TotalSegmentator dataset and find that increasing model size brings a stronger performance gain. This observation reveals that a large model is promising in medical image segmentation. Furthermore, we evaluate the transferability of our model on 14 downstream datasets for direct inference and 3 datasets for further fine-tuning, covering various modalities and segmentation targets. We observe good performance of our pre-trained model in both direct inference and fine-tuning. The code and pre-trained models are available at https://github.com/Ziyan-Huang/STU-Net.

cs.CV↗

Clustered Embedding Learning for Recommender Systems

In recent years, recommender systems have advanced rapidly, where embedding learning for users and items plays a critical role. A standard method learns a unique embedding vector for each user and item. However, such a method has two important limitations in real-world applications: 1) it is hard to learn embeddings that generalize well for users and items with rare interactions on their own; and 2) it may incur unbearably high memory costs when the number of users and items scales up. Existing approaches either can only address one of the limitations or have flawed overall performances. In this paper, we propose Clustered Embedding Learning (CEL) as an integrated solution to these two problems. CEL is a plug-and-play embedding learning framework that can be combined with any differentiable feature interaction model. It is capable of achieving improved performance, especially for cold users and items, with reduced memory cost. CEL enables automatic and dynamic clustering of users and items in a top-down fashion, where clustered entities jointly learn a shared embedding. The accelerated version of CEL has an optimal time complexity, which supports efficient online updates. Theoretically, we prove the identifiability and the existence of a unique optimal number of clusters for CEL in the context of nonnegative matrix factorization. Empirically, we validate the effectiveness of CEL on three public datasets and one business dataset, showing its consistently superior performance against current state-of-the-art methods. In particular, when incorporating CEL into the business model, it brings an improvement of $+0.6\%$ in AUC, which translates into a significant revenue gain; meanwhile, the size of the embedding table gets $2650$ times smaller.

cs.AI↗

Deterministic All-versus-nothing Proofs of Bell Nonlocality Induced from Qudit Non-stabilizer States

Recently, a kind of deterministic all-versus-nothing proof of Bell nonlocality induced from the qubit non-stabilizer state was proposed, breaking the tradition that deterministic all-versus-nothing proofs are always derived from stabilizer states. A trivial generalization to the qudit (d is even) version is by using a special basis map, but such a proof can still be reduced to the qubit version. So far, whether high dimensional non-stabilizer states can induce nontrivial deterministic all-versus-nothing proofs of Bell nonlocality remains unknown. Here we present an example induced from a specific four-qudit non-stabilizer state (with d = 4), showing that such proofs can be constructed in high dimensional scenarios as well.

quant-ph↗

Convergence Analysis for Training Stochastic Neural Networks via Stochastic Gradient Descent

In this paper, we carry out numerical analysis to prove convergence of a novel sample-wise back-propagation method for training a class of stochastic neural networks (SNNs). The structure of the SNN is formulated as discretization of a stochastic differential equation (SDE). A stochastic optimal control framework is introduced to model the training procedure, and a sample-wise approximation scheme for the adjoint backward SDE is applied to improve the efficiency of the stochastic optimal control solver, which is equivalent to the back-propagation for training the SNN. The convergence analysis is derived with and without convexity assumption for optimization of the SNN parameters. Especially, our analysis indicates that the number of SNN training steps should be proportional to the square of the number of layers in the convex optimization case. Numerical experiments are carried out to validate the analysis results, and the performance of the sample-wise back-propagation method for training SNNs is examined by benchmark machine learning examples.

math.NA↗

Functional building blocks for scalable multipartite entanglement in optical lattices

Featuring excellent coherence and operated parallelly, ultracold atoms in optical lattices form a competitive candidate for quantum computation. For this, a massive number of parallel entangled atom pairs have been realized in superlattices. However, the more formidable challenge is to scale-up and detect multipartite entanglement due to the lack of manipulations over local atomic spins in retro-reflected bichromatic superlattices. Here we developed a new architecture based on a cross-angle spin-dependent superlattice for implementing layers of quantum gates over moderately-separated atoms incorporated with a quantum gas microscope for single-atom manipulation. We created and verified functional building blocks for scalable multipartite entanglement by connecting Bell pairs to one-dimensional 10-atom chains and two-dimensional plaquettes of $2\times4$ atoms. This offers a new platform towards scalable quantum computation and simulation.

cond-mat.quant-gas↗

Strings And Colorings Of Topological Coding Towards Asymmetric Topology Cryptography

We, for anti-quantum computing, will discuss various number-based strings, such as number-based super-strings, parameterized strings, set-based strings, graph-based strings, integer-partitioned and integer-decomposed strings, Hanzi-based strings, as well as algebraic operations based on number-based strings. Moreover, we introduce number-based string-colorings, magic-constraint colorings, and vector-colorings and set-colorings related with strings. For the technique of encrypting the entire network at once, we propose graphic lattices related with number-based strings, Hanzi-graphic lattices, string groups, all-tree-graphic lattices. We study some topics of asymmetric topology cryptography, such as topological signatures, Key-pair graphs, Key-pair strings, one-encryption one-time and self-certification algorithms. Part of topological techniques and algorithms introduced here are closely related with NP-complete problems or NP-hard problems.

cs.IT↗

Long-duration Gamma-ray Burst and Associated Kilonova Emission from Fast-spinning Black Hole--Neutron Star Mergers

Here we collect three unique bursts, GRBs\,060614, 211211A and 211227A, all characterized by a long-duration main emission (ME) phase and a rebrightening extended emission (EE) phase, to study their observed properties and the potential origin as neutron star-black hole (NSBH) mergers. NS-first-born (BH-first-born) NSBH mergers tend to contain fast-spinning (non-spinning) BHs that more easily (hardly) allow tidal disruption to happen with (without) forming electromagnetic signals. We find that NS-first-born NSBH mergers can well interpret the origins of these three GRBs, supported by that: (1) Their X-ray MEs and EEs show unambiguous fall-back accretion signatures, decreasing as $\propto{t}^{-5/3}$, which might account for their long duration. The EEs can result from the fall-back accretion of $r$-process heating materials, predicted to occur after NSBH mergers. (2) The beaming-corrected local event rate density for this type of merger-origin long-duration GRBs is $\mathcal{R}_0\sim2.4^{+2.3}_{-1.3}\,{\rm{Gpc}}^{-3}\,{\rm{yr}}^{-1}$, consistent with that of NS-first-born NSBH mergers. (3) Our detailed analysis on the EE, afterglow and kilonova of the recently high-impact event GRB\,211211A reveals it could be a merger between a $\sim1.23^{+0.06}_{-0.07}\,M_\odot$ NS and a $\sim8.21^{+0.77}_{-0.75}\,M_\odot$ BH with an aligned-spin of $χ_{\rm{BH}}\sim0.62^{+0.06}_{-0.07}$, supporting an NS-first-born NSBH formation channel. Long-duration burst with rebrightening fall-back accretion signature after ME, and bright kilonova might be commonly observed features for on-axis NSBHs. We estimate the multimessenger detection rate between gravitational waves, GRBs and kilonovae from NSBH mergers in O4 (O5) is $\sim0.1\,{\rm{yr}}^{-1}$ ($\sim1\,{\rm{yr}}^{-1}$).

astro-ph.HE↗

Thermalization dynamics of a gauge theory on a quantum simulator

Gauge theories form the foundation of modern physics, with applications ranging from elementary particle physics and early-universe cosmology to condensed matter systems. We perform quantum simulations of the unitary dynamics of a U(1) symmetric gauge field theory and demonstrate emergent irreversible behavior. The highly constrained gauge theory dynamics is encoded in a one-dimensional Bose--Hubbard simulator, which couples fermionic matter fields through dynamical gauge fields. We investigate global quantum quenches and the equilibration to a steady state well approximated by a thermal ensemble. Our work may enable the investigation of elusive phenomena, such as Schwinger pair production and string-breaking, and paves the way for simulating more complex higher-dimensional gauge theories on quantum synthetic matter devices.

cond-mat.quant-gas↗

Parameterized Colorings And Labellings Of Graphs In Topological Coding

The coming quantum computation is forcing us to reexamine the cryptosystems people use. We are applying graph colorings of topological coding to modern information security and future cryptography against supercomputer and quantum computer attacks in the near future. Many of techniques introduced here are associated with many mathematical conjecture and NP-problems. We will introduce a group of W-constraint (k,d)-total colorings and algorithms for realizing these colorings in some kinds of graphs, which are used to make quickly public-keys and private-keys with anti-quantum computing, these (k,d)-total colorings are: graceful (k,d)-total colorings, harmonious (k,d)-total colorings, (k,d)-edge-magic total colorings, (k,d)-graceful-difference total colorings and (k,d)-felicitous-difference total colorings. One of useful tools we used is called Topcode-matrix with elements can be all sorts of things, for example, sets, graphs, number-based strings. Most of parameterized graphic colorings/labelings are defined by Topcode-matrix algebra here. From the application point of view, many of our coloring techniques are given by algorithms and easily converted into programs.

cs.IT↗

Probing the progenitor of high-$z$ short-duration GRB 201221D and its possible bulk acceleration in prompt emission

The growing observed evidence shows that the long- and short-duration gamma-ray bursts (GRBs) originate from massive star core-collapse and the merger of compact stars, respectively. GRB 201221D is a short-duration GRB lasting $\sim 0.1$ s without extended emission (EE) at high redshift $z=1.046$. By analyzing data observed with the Swift/BAT and Fermi/GBM, we find that a cutoff power-law model can adequately fit the spectrum with a soft $E_{\rm p}=113^{+9}_{-7}$ keV, and isotropic energy $E_{γ,iso} =1.36^{+0.17}_{-0.14}\times 10^{51}~\rm erg$. In order to reveal the possible physical origin of GRB 201221D, we adopted multi-wavelength criteria (e.g., Amati relation, $\varepsilon$-parameter, amplitude parameter, local event rate density, luminosity function, and properties of the host galaxy), and find that most of the observations of GRB 201221D favor a compact star merger origin. Moreover, we find that $\hatα$ is larger than $2+\hatβ$ in the prompt emission phase which suggests that the emission region is possibly undergoing acceleration during the prompt emission phase with a Poynting-flux-dominated jet.

astro-ph.HE↗

Thermodynamic modeling with uncertainty quantification in the Nb-Ni system using the upgraded PyCalphad and ESPEI

The Nb-Ni system has been remodeled with uncertainty quantification (UQ) by using the presently upgraded software tools of PyCalphad and ESPEI that contain the new capability to model site occupancy of Wyckoff position for the phases of interest. Specifically, the five- and three-sublattice models are used to model the topologically close pack (TCP) phases of μ-Nb7Ni6 and δ-NbNi3, respectively, according to exactly their Wyckoff positions; where the inputs for CALPHAD-based modeling include the presently predicted thermochemical data as a function of temperature by density functional theory (DFT) based first-principles and phonon calculations together with both phase equilibrium and site occupancy data in the literature. Besides phase diagram and thermodynamic properties, the present CALPHAD predictions of site occupancies are also agreed well with experimental data such as the measured Nb sites in μ-Nb7Ni6. In addition, the predicted UQ values using the Markov Chain Monte Carlo (MCMC) method as implemented in ESPEI make it possible to quantify uncertainties in the Nb-Ni system, such as site occupancies in μ-Nb7Ni6 and enthalpy of mixing in liquid.

cond-mat.mtrl-sci↗

Rigorous criteria for anomalous waves induced by abrupt depth change using truncated KdV statistical mechanics

The truncated Korteweg-De Vries (TKdV) system, a shallow-water wave model with Hamiltonian structure that exhibits weakly turbulent dynamics, has been found to accurately predict the anomalous wave statistics observed in recent laboratory experiments. Majda et al. (2019) developed a TKdV statistical mechanics framework based on a mixed Gibbs measure that is supported on a surface of fixed energy (microcanonical) and takes the usual macroconical form in the Hamiltonian. This paper reports two rigorous results regarding the surface-displacement distributions implied by this ensemble, both in the limit of the cutoff wavenumber $Λ$ growing large. First, we prove that if the inverse temperature vanishes, microstate statistics converge to Gaussian as $Λ\to \infty$. Second, we prove that if nonlinearity is absent and the inverse-temperature satisfies a certain physically-motivated scaling law, then microstate statistics converge to Gaussian as $Λ\to \infty$. When the scaling law is not satisfied, simple numerical examples demonstrate symmetric, yet highly non-Gaussian, displacement statistics to emerge in the linear system, illustrating that nonlinearity is not a strict requirement for non-normality in the fixed-energy ensemble. The new results, taken together, imply necessary conditions for the anomalous wave statistics observed in previous numerical studies. In particular, non-vanishing inverse temperature and either the presence of nonlinearity or the violation of the scaling law are required for displacement statistics to deviate from Gaussian. The proof of this second theorem involves the construction of an approximating measure, which we find also elucidates the peculiar spectral decay observed in numerical studies and may open the door for improved sampling algorithms.

physics.flu-dyn↗

Luminosity function and event rate density of XMM-Newton-selected supernova shock-breakout candidates

A dozen X-ray supernova shock breakout (SN SBO) candidates were reported recently based on XMM-Newton archival data, which increased the X-ray selected SN SBO sample by an order of magnitude. Assuming they are genuine SN SBOs, we study the luminosity function (LF) by improving upon the method used in our previous work. The light curves and the spectra of the candidates were used to derive the maximum volume within which these objects could be detected with XMM-Newton by simulation. The results show that the SN SBO LF can be described by either a broken power law (BPL) with indices (at the 68$\%$ confidence level) of $0.48 \pm 0.28$ and $2.11 \pm 1.27$ before and after the break luminosity at $\log (L_b/\rm erg\,s^{-1})=$ $45.32 \pm 0.55$ or a single power law (SPL) with index of $0.80 \pm 0.16$. The local event rate densities of SN SBOs above $5\times 10^{42}$ $\rm erg\,s^{-1}$ are consistent for two models, i.e., $4.6 ^{+1.7}_{-1.3} \times 10^4$ and $4.9 ^{+1.9}_{-1.4} \times 10^4$ $\rm Gpc^{-3}\,yr^{-1}$ for BPL and SPL models, respectively. The number of fast X-ray transients of SN SBO origin can be significantly increased by the wide-field X-ray telescopes such as the Einstein Probe.

astro-ph.HE↗

Thermodynamic modeling of the Pd-Zn system with uncertainty quantification and its implication to tailor catalysts

Pd-Zn intermetallic catalysts show encouraging combinations of activity and selectivity on well-defined active site ensembles. Thermodynamic description of the Pd-Zn system, delineating phase boundaries, and enumerating site occupancies within intermediate alloy phases, are essential to determining the ensembles of Pd-Zn atoms as a function of composition and temperature. Combining the present extensive first-principles calculations based on density functional theory (DFT) and available experimental data, the Pd-Zn system was remodeled using the CALculation of PHAse Diagrams (CALPHAD) approach. High throughput modeling tools with uncertainty quantification, i.e., ESPEI and PyCalphad, were incorporated in the phase analysis. The site occupancies across the γ-phase composition region were given special attention. A four-sublattice model was used for the γ-phase owing to its four Wyckoff positions, i.e., the outer tetrahedral (OT) site 8c, the inner tetrahedral (IT) site 8c, the octahedral (OH) 12e, and the cuboctahedral (CO) site 24g. The site fractions of Pd and Zn calculated from the present thermodynamic model show the occupancy preference of Pd in the OT and OH sublattices in agreement with experimental observations. The force constants obtained from DFT-based phonon calculations further supports the tendency of Pd occupying the OH sublattice compared with the IT and CO sublattice. The catalytic assembles changing from Pd monomers (Pd1) to trimers (Pd3) on the surface of the γ-phase is attributed to the increase of Pd occupancy in the OH sublattice.

cond-mat.mtrl-sci↗

Adversarial Attacks Against Deep Generative Models on Data: A Survey

Deep generative models have gained much attention given their ability to generate data for applications as varied as healthcare to financial technology to surveillance, and many more - the most popular models being generative adversarial networks and variational auto-encoders. Yet, as with all machine learning models, ever is the concern over security breaches and privacy leaks and deep generative models are no exception. These models have advanced so rapidly in recent years that work on their security is still in its infancy. In an attempt to audit the current and future threats against these models, and to provide a roadmap for defense preparations in the short term, we prepared this comprehensive and specialized survey on the security and privacy preservation of GANs and VAEs. Our focus is on the inner connection between attacks and model architectures and, more specifically, on five components of deep generative models: the training data, the latent code, the generators/decoders of GANs/ VAEs, the discriminators/encoders of GANs/ VAEs, and the generated data. For each model, component and attack, we review the current research progress and identify the key challenges. The paper concludes with a discussion of possible future attacks and research directions in the field.

cs.CR↗

Annotation-efficient deep learning for automatic medical image segmentation

Automatic medical image segmentation plays a critical role in scientific research and medical care. Existing high-performance deep learning methods typically rely on large training datasets with high-quality manual annotations, which are difficult to obtain in many clinical applications. Here, we introduce Annotation-effIcient Deep lEarning (AIDE), an open-source framework to handle imperfect training datasets. Methodological analyses and empirical evaluations are conducted, and we demonstrate that AIDE surpasses conventional fully-supervised models by presenting better performance on open datasets possessing scarce or noisy annotations. We further test AIDE in a real-life case study for breast tumor segmentation. Three datasets containing 11,852 breast images from three medical centers are employed, and AIDE, utilizing 10% training annotations, consistently produces segmentation maps comparable to those generated by fully-supervised counterparts or provided by independent radiologists. The 10-fold enhanced efficiency in utilizing expert labels has the potential to promote a wide range of biomedical applications.

eess.IV↗

Generative deep learning as a tool for inverse design of high-entropy refractory alloys

Generative deep learning is powering a wave of new innovations in materials design. In this article, we discuss the basic operating principles of these methods and their advantages over rational design through the lens of a case study on refractory high-entropy alloys for ultra-high-temperature applications. We present our computational infrastructure and workflow for the inverse design of new alloys powered by these methods. Our preliminary results show that generative models can learn complex relationships in order to generate novelty on demand, making them a valuable tool for materials informatics.

cond-mat.mtrl-sci↗