Searcharxiv⌕ Search

arXiv subjects

Da Huang

Publications and source records attributed to Da Huang.

At least 37 records · Page 2Linked to original sources

Cosmic Birefringence from Neutrino and Dark Matter Asymmetries

In light of the recent measurement of the nonzero Cosmic Microwave Background (CMB) polarization rotation angle from the Planck 2018 data, we explore the possibility that such a cosmic birefringence effect is induced by coupling a fermionic current with photons via a Chern-Simons-like term. We begin our discussion by rederiving the general formulae of the cosmic birefringence angle with correcting a mistake in the previous study. We then identify the fermions in the current as the left-handed electron neutrinos and asymmetric dark matter (ADM) particles, since the rotation angle is sourced by the number density difference between particles and antiparticles. For the electron neutrino case, with the value of the degeneracy parameter $ξ_{ν_e}$ recently measured by the EMPRESS survey, we find a large parameter space which can explain the CMB photon polarization rotations. On the other hand, for the ADM solution, we consider two benchmark cases with $M_χ= 5$~GeV and 5~keV. The former is the natural value of the ADM mass if the observed ADM and baryon asymmetry in the Universe are produced by the same mechanism, while the latter provides a warm DM candidate. In addition, we explore the experimental constraints from the CMB power spectra and the DM direct detections.

astro-ph.CO↗

Unitarity bounds on extensions of Higgs sector

It is widely believed that extensions of the minimal Higgs sector is one of the promising directions for resolving many puzzles beyond the Standard Model (SM). In this work, we study the unitarity bounds on the models by extending the two-Higgs-doublet model with an additional real or complex Higgs triplet scalar. By noting that the SM gauge symmetries $SU(2)_L\times U(1)_Y$ are recovered at high energies, we can classify the two-body scattering states by decomposing the direct product of two scalar multiplets into their direct sum of irreducible representations of electroweak gauge groups. In such state bases, the s-wave amplitudes of two-body scalar scatterings can be written in the form of block-diagonalized scattering matrices. Then the application of the perturbative unitarity conditions on the eigenvalues of scattering matrices leads to the analytic constraints on the model parameters. Finally, we numerically investigate the complex triplet scalar extension of the two-Higgs-doublet model, finding that the perturbative unitarity places useful stringent bounds on the model parameter space.

hep-ph↗

Symbolic Discovery of Optimization Algorithms

We present a method to formulate algorithm discovery as program search, and apply it to discover optimization algorithms for deep neural network training. We leverage efficient search techniques to explore an infinite and sparse program space. To bridge the large generalization gap between proxy and target tasks, we also introduce program selection and simplification strategies. Our method discovers a simple and effective optimization algorithm, $\textbf{Lion}$ ($\textit{Evo$\textbf{L}$ved S$\textbf{i}$gn M$\textbf{o}$me$\textbf{n}$tum}$). It is more memory-efficient than Adam as it only keeps track of the momentum. Different from adaptive optimizers, its update has the same magnitude for each parameter calculated through the sign operation. We compare Lion with widely used optimizers, such as Adam and Adafactor, for training a variety of models on different tasks. On image classification, Lion boosts the accuracy of ViT by up to 2% on ImageNet and saves up to 5x the pre-training compute on JFT. On vision-language contrastive learning, we achieve 88.3% $\textit{zero-shot}$ and 91.1% $\textit{fine-tuning}$ accuracy on ImageNet, surpassing the previous best results by 2% and 0.1%, respectively. On diffusion models, Lion outperforms Adam by achieving a better FID score and reducing the training compute by up to 2.3x. For autoregressive, masked language modeling, and fine-tuning, Lion exhibits a similar or better performance compared to Adam. Our analysis of Lion reveals that its performance gain grows with the training batch size. It also requires a smaller learning rate than Adam due to the larger norm of the update produced by the sign function. Additionally, we examine the limitations of Lion and identify scenarios where its improvements are small or not statistically significant. Lion is also successfully deployed in production systems such as Google search ads CTR model.

cs.LG↗

Larger language models do in-context learning differently

We study how in-context learning (ICL) in language models is affected by semantic priors versus input-label mappings. We investigate two setups-ICL with flipped labels and ICL with semantically-unrelated labels-across various model families (GPT-3, InstructGPT, Codex, PaLM, and Flan-PaLM). First, experiments on ICL with flipped labels show that overriding semantic priors is an emergent ability of model scale. While small language models ignore flipped labels presented in-context and thus rely primarily on semantic priors from pretraining, large models can override semantic priors when presented with in-context exemplars that contradict priors, despite the stronger semantic priors that larger models may hold. We next study semantically-unrelated label ICL (SUL-ICL), in which labels are semantically unrelated to their inputs (e.g., foo/bar instead of negative/positive), thereby forcing language models to learn the input-label mappings shown in in-context exemplars in order to perform the task. The ability to do SUL-ICL also emerges primarily with scale, and large-enough language models can even perform linear classification in a SUL-ICL setting. Finally, we evaluate instruction-tuned models and find that instruction tuning strengthens both the use of semantic priors and the capacity to learn input-label mappings, but more of the former.

cs.CL↗

$W$-Boson Mass Anomaly from a General $SU(2)_{L}$ Scalar Multiplet

We explain the $W$-boson mass anomaly by introducing an $SU(2)_L$ scalar multiplet with general isospin and hypercharge in the case without its vacuum expectation value. It is shown that the dominant contribution from the scalar multiplet to the $W$-boson mass arises at one-loop level, which can be expressed in terms of the electroweak (EW) oblique parameters $T$ and $S$ at leading order. We firstly rederive the general formulae of $T$ and $S$ induced by a scalar multiplet of EW charges, confirming the results in the literature. We then study several specific examples of great phenomenological interest by applying these general expressions. As a result, it is found that the model with a scalar multiplet in an $SU(2)_L$ real representation with $Y=0$ cannot generate the required $M_W$ correction since it leads to vanishing values of $T$ and $S$. On the other hand, the cases with scalars in a complex representation under $SU(2)_L$ with a general hypercharge can explain the $M_W$ excess observed by CDF-\uppercase\expandafter{\romannumeral2} due to nonzero $T$ and $S$. We further take into account of the strong constraints from the perturbativity and the EW global fit of the precision data, and vary the isospin representation and hypercharge of the additional scalar multiplet, in order to assess the extent of the model to solve the $W$-boson mass anomaly. It turns out that these constraints play important roles in setting limits on the model parameter space. We also briefly describe the collider signatures of the extra scalar multiplet, especially when it contains long-lived heavy highly charged states.

hep-ph↗

Muon $g-2$ Anomaly from a Massive Spin-2 Particle

We investigate the possibility to interpret the muon $g-2$ anomaly in terms of a massive spin-2 particle, $G$, which can be identified as the first Kaluza-Klein graviton in the generalized Randall-Sundrum model. In particular, we obtain the leading-order contributions to the muon $g-2$ by calculating the relevant one-loop Feynman diagrams induced by $G$. The analytic expression is shown to keep the gauge invariance of the quantum electrodynamics and to be consistent with the expected UV divergence structure. Moreover, we impose the theoretical bounds from the perturbativity and the experimental constraints from LHC and LEP-II on our model. Especially, we derive novel perturbativity constraints on nonrenormalizable operators related to $G$, which are the natural generalization of the counterpart for the renormalizable operators. As a result, we show that there exists a substantial parameter space, which can accommodate the muon $g-2$ anomaly allowed by all constraints. Finally, we also make comments on the possible explanation of the electron $g-2$ anomalies with the massive spin-2 particle.

hep-ph↗

Contributions to the Muon $g-2$ from a Three-Form Field

We examine contributions to the muon dipole moment $g-2$ from a 3-form field $Ω$, which naturally arises from many fundamental theories, such as the string theory and the hyperunified field theory. In particular, by calculating the one-loop Feynman diagram, we have obtained the leading-order $Ω$-induced contribution to the muon $g-2$, which is found to be finite. Then we investigate the theoretical constraints from perturbativity and unitarity. Especially, the unitarity bounds are yielded by computing the tree-level $μ^+μ^-$ scattering amplitudes of various initial and final helicity configurations. As a result, despite the strong unitarity bounds imposed on this model of $Ω$, we have still found a substantial parameter space which can accommodates the muon $g-2$ data.

hep-ph↗

TabNAS: Rejection Sampling for Neural Architecture Search on Tabular Datasets

The best neural architecture for a given machine learning problem depends on many factors: not only the complexity and structure of the dataset, but also on resource constraints including latency, compute, energy consumption, etc. Neural architecture search (NAS) for tabular datasets is an important but under-explored problem. Previous NAS algorithms designed for image search spaces incorporate resource constraints directly into the reinforcement learning (RL) rewards. However, for NAS on tabular datasets, this protocol often discovers suboptimal architectures. This paper develops TabNAS, a new and more effective approach to handle resource constraints in tabular NAS using an RL controller motivated by the idea of rejection sampling. TabNAS immediately discards any architecture that violates the resource constraints without training or learning from that architecture. TabNAS uses a Monte-Carlo-based correction to the RL policy gradient update to account for this extra filtering step. Results on several tabular datasets demonstrate the superiority of TabNAS over previous reward-shaping methods: it finds better models that obey the constraints.

cs.LG↗

On the Factory Floor: ML Engineering for Industrial-Scale Ads Recommendation Models

For industrial-scale advertising systems, prediction of ad click-through rate (CTR) is a central problem. Ad clicks constitute a significant class of user engagements and are often used as the primary signal for the usefulness of ads to users. Additionally, in cost-per-click advertising systems where advertisers are charged per click, click rate expectations feed directly into value estimation. Accordingly, CTR model development is a significant investment for most Internet advertising companies. Engineering for such problems requires many machine learning (ML) techniques suited to online learning that go well beyond traditional accuracy improvements, especially concerning efficiency, reproducibility, calibration, credit attribution. We present a case study of practical techniques deployed in Google's search ads CTR model. This paper provides an industry case study highlighting important areas of current ML research and illustrating how impactful new ML methods are evaluated and made useful in a large-scale industrial setting.

cs.IR↗

Unitarity Bounds on the Massive Spin-2 Particle Explanation of Muon $g-2$ Anomaly

Motivated by the long-standing discrepancy between the Standard Model prediction and the experimental measurement of the muon magnetic dipole moment, we have recently proposed to interpret this muon $g-2$ anomaly in terms of the loop effect induced by a new massive spin-2 field $G$. In the present paper, we investigate the unitarity bounds on this scenario. We calculate the $s$-wave projected amplitudes for two-body elastic scatterings of charged leptons and photons mediated by $G$ at high energies for all possible initial and final helicity states. By imposing the condition of the perturbative unitarity, we obtain the analytic constraints on the charged-lepton-$G$ and photon-$G$ couplings. We then apply our results to constrain the parameter space relevant to the explanation of the muon $g-2$ anomaly.

hep-ph↗

Probing WIMPs in space-based gravitational wave experiments

Although searches for dark matter have lasted for decades, no convincing signal has been found without ambiguity in underground detections, cosmic ray observations, and collider experiments. We show by example that gravitational wave (GW) observations can be a supplement to dark matter detections if the production of dark matter follows a strong first-order cosmological phase transition. We explore this possibility in a complex singlet extension of the standard model with CP symmetry. We demonstrate three benchmarks in which the GW signals from the first-order phase transition are loud enough for future space-based GW observations, for example, BBO, U-DECIGO, LISA, Taiji, and TianQin. While satisfying the constraints from the XENON1T experiment and the Fermi-LAT gamma-ray observations, the dark matter candidate with its mass around $\sim 1$~TeV in these scenarios has a correct relic abundance obtained by the Planck observations of the cosmic microwave background radiation.

hep-ph↗

Impact of electroweak group representation in models for $B$ and $g-2$ anomalies from Dark Loops

We discuss two models which are part of a class providing a common explanation for lepton flavor universality violation in $b \to s l^+ l^- $ decays, the dark matter (DM) problem and the muon $(g-2)$ anomaly. The $B$ meson decays and the muon $(g-2)$ anomalies are explained by additional one-loop diagrams with DM candidates. The models have one extra fermion field and two extra scalar fields relative to the Standard Model (SM). The $SU(3)$ quantum numbers are fixed by the interaction with the SM fermions in a new Yukawa Lagrangian that connects the dark and the visible sectors. We compare two models, one where the fermion is a singlet and the scalars are doublets under $SU(2)_L$ and another one where the fermion is a doublet and the scalars are singlets under $SU(2)_L$. We conclude that both models can explain all new physics phenomena simultaneously, while satisfying all other flavor and DM constraints. However, there are crucial differences between how the DM constraints affect the two models leading to a noticeable difference in the allowed DM mass range.

hep-ph↗

Scalar Gravitational Wave Signals from Core Collapse in Massive Scalar-Tensor Gravity with Triple-Scalar Interactions

The spontaneous scalarization during the stellar core collapse in the massive scalar-tensor theories of gravity introduces extra polarizations (on top of the plus and cross modes) in gravitational waves, whose amplitudes are determined by several model parameters. Observations of such scalarization-induced gravitational waveforms therefore offer valuable probes into these theories of gravity. Considering a triple-scalar interactions in such theories, we find that the self-coupling effects suppress the magnitude of the scalarization and thus reduce the amplitude of the associated gravitational wave signals. In addition, the self-interacting effects in the gravitational waveform are shown to be negligible due to the dispersion throughout the astrophysically distant propagation. As a consequence, the gravitational waves observed on the Earth feature the characteristic inverse-chirp pattern. Although not with the on-going ground-based detectors, we illustrate that the scalarization-induced gravitational waves may be detectable at a signal-to-noise ratio level of ${\cal O}(100)$ with future detectors, such as Einstein Telescope and Cosmic Explorer.

gr-qc↗

PatentNet: A Large-Scale Incomplete Multiview, Multimodal, Multilabel Industrial Goods Image Database

In deep learning area, large-scale image datasets bring a breakthrough in the success of object recognition and retrieval. Nowadays, as the embodiment of innovation, the diversity of the industrial goods is significantly larger, in which the incomplete multiview, multimodal and multilabel are different from the traditional dataset. In this paper, we introduce an industrial goods dataset, namely PatentNet, with numerous highly diverse, accurate and detailed annotations of industrial goods images, and corresponding texts. In PatentNet, the images and texts are sourced from design patent. Within over 6M images and corresponding texts of industrial goods labeled manually checked by professionals, PatentNet is the first ongoing industrial goods image database whose varieties are wider than industrial goods datasets used previously for benchmarking. PatentNet organizes millions of images into 32 classes and 219 subclasses based on the Locarno Classification Agreement. Through extensive experiments on image classification, image retrieval and incomplete multiview clustering, we demonstrate that our PatentNet is much more diverse, complex, and challenging, enjoying higher potentials than existing industrial image datasets. Furthermore, the characteristics of incomplete multiview, multimodal and multilabel in PatentNet are able to offer unparalleled opportunities in the artificial intelligence community and beyond.

cs.CV↗

Rethinking Co-design of Neural Architectures and Hardware Accelerators

Neural architectures and hardware accelerators have been two driving forces for the progress in deep learning. Previous works typically attempt to optimize hardware given a fixed model architecture or model architecture given fixed hardware. And the dominant hardware architecture explored in this prior work is FPGAs. In our work, we target the optimization of hardware and software configurations on an industry-standard edge accelerator. We systematically study the importance and strategies of co-designing neural architectures and hardware accelerators. We make three observations: 1) the software search space has to be customized to fully leverage the targeted hardware architecture, 2) the search for the model architecture and hardware architecture should be done jointly to achieve the best of both worlds, and 3) different use cases lead to very different search outcomes. Our experiments show that the joint search method consistently outperforms previous platform-aware neural architecture search, manually crafted models, and the state-of-the-art EfficientNet on all latency targets by around 1% on ImageNet top-1 accuracy. Our method can reduce energy consumption of an edge accelerator by up to 2x under the same accuracy constraint, when co-adapting the model architecture and hardware accelerator configurations.

cs.LG↗

Probing the doubly-charged Higgs with Muonium to Antimuonium Conversion Experiment

The spontaneous muonium-to-antimuonium conversion is one of the interesting charged lepton flavor violation processes. MACE is the next generation experiment to probe such a phenomenon. In models with a triplet Higgs to generate neutrino masses, such as Type-II seesaw and its variant, this process can be induced by the doubly-charged Higgs contained in it. In this article, we study the prospect of MACE to probe these models via the muonium-to-antimuonium transitions. After considering the limits from $μ^+ \rightarrow e^+ γ$ and $μ^+ \rightarrow e^+ e^- e^+$, we find that MACE could probe a parameter space for the doubly-charged Higgs which is beyond the reach of LHC and other flavor experiments.

hep-ph↗

Electroweak phase transition confronted with dark matter detection constraints

We study the type-II first-order electroweak phase transition and dark matter (DM) phenomenology in both real and complex singlet extensions of SM. In the real singlet extension with a $\mathbb{Z}_2$ symmetry, we show that the parameter regions favored by the phase transition suffer from strong constraints from DM direct detection so that only a negligible fraction ($f_{X}\sim 10^{-4}-10^{-5}$) of DM composed of the real singlet scalar can survive the LUX and XENON1T constraints. In the complex singlet $S$ case, we impose a $CP$ symmetry $S\to S^{*}$ to the scalar potential. The real component of $S$ can mix with SM Higgs boson while the imaginary component becomes a DM candidate due to the protection of the $CP$ symmetry. By taking into account the current experimental constraints of invisible Higgs decays, Higgs signal strength measurements, and dark matter detections, we find that there exists a large parameter space for the type-II electroweak phase transition to occur while explaining all of the dark matter relic density. We identify a subset of parameter space that is promising for future experiments, including the di-Higgs and Higgs signal strength measurements at the HL-LHC and the dark matter direct detection in the XENONnT project.

hep-ph↗

Observational Constraints on the Cosmology with Holographic Dark Fluid

We consider the holographic Friedman-Robertson-Walker (hFRW) universe on the 4-dimensional membrane embedded in the 5-dimensional bulk spacetime and fit the parameters with the observational data. In order to fully account for the phenomenology of this scenario, we consider the models with the brane cosmological constant and the negative bulk cosmological constant. The contribution from the bulk is represented as the holographic dark fluid on the membrane. We derive the universal modified Friedmann equation by including all of these effects in both braneworld and holographic cutoff approaches. For three specific models, namely, the pure hFRW model, the one with the brane cosmological constant, and the one with the negative bulk cosmological constant, we compare the model predictions with the observations. The parameters in the considered hFRW models are constrained with observational data. In particular, it is shown that the model with the brane cosmological constant can fit data as well as the standard $Λ$CDM universe. We also find that the $σ_8$ tension observed in different large-structure experiments can be effectively relaxed in this holographic scenario.

astro-ph.CO↗