SearcharxivSearch

arXiv subjects

Michael Beyer

Publications and source records attributed to Michael Beyer.

16 recordsLinked to original sources

Why Does Post-Training Quantization Work?

Post-training quantization compresses large language models (LLMs) by storing their weights at reduced precision, and each quantized weight introduces an error into the hidden states. Naively, these errors should accumulate with depth and corrupt next-token prediction; randomly initialized models accumulate these discrepancies rapidly, whereas quantized pretrained models accumulate much less hidden-state error and largely maintain downstream task performance, even though they were never trained with quantization noise. This raises the question we address: why does post-training quantization work? Comparing full-precision and quantized forward passes, we identify two mechanisms that characterize pretrained quantization robustness. First, the error a layer newly introduces tends to oppose the error it inherits from the layer's input. The two cancel partially such that the discrepancy between full-precision and quantized passes grows slowly. This counteracting residual interaction develops during pretraining. Our quantitative analysis identifies it as a major factor slowing hidden-error growth. Second, LM-head geometry preferentially preserves the scores and probabilities of high-ranked tokens, which typically represent the model's most confident predictions. Together, these mechanisms explain why quantization error that passes through numerous layers can still produce only small output changes, and we verify the findings across models and quantization settings.

cs.LG

TetraJet-v2: Accurate NVFP4 Training for Large Language Models with Oscillation Suppression and Outlier Control

Large Language Models (LLMs) training is prohibitively expensive, driving interest in low-precision fully-quantized training (FQT). While novel 4-bit formats like NVFP4 offer substantial efficiency gains, achieving near-lossless training at such low precision remains challenging. We introduce TetraJet-v2, an end-to-end 4-bit FQT method that leverages NVFP4 for activations, weights, and gradients in all linear layers. We identify two critical issues hindering low-precision LLM training: weight oscillation and outliers. To address these, we propose: 1) an unbiased double-block quantization method for NVFP4 linear layers with practically optimal convergence in LLM training, 2) OsciReset, the first effective algorithm to suppress LLMs' weight oscillation bottleneck, and 3) OutControl, a mix-precision algorithm to retain outlier accuracy. TetraJet-v2 outperforms prior methods on FP4 pre-training for LLMs across models up to 370M parameters trained up to 212B tokens, reducing the performance gap to BF16 by an average of 51.3% while enabling an 1.67x end-to-end speedup over FP8. The code is available at https://github.com/thu-ml/TetraJet-v2-NVFP4Training.

cs.LG

Preemptive Two-stage Goal-Programming Formulation of a Strict Version of the Unbounded Knapsack Problem with Bounded Weights

The unbounded knapsack problem with bounded weights is a variant of the well-studied variant of the traditional binary knapsack problem; key changes being the relaxation of the binary constraint and allowing the unit weights of each item to fall within a range. In this paper, we formulate a variant of this problem, which we call the strict unbounded knapsack problem with bounded weights, by replacing the inequality constraint on the total weight with an equality. We show that this problem can be decomposed into a two-stage, pre-emptive goal programming problem, with the first stage being a 2-dimensional knapsack problem and the second being either a linear feasibility program (per canonical formulation) or simply a linearly-constrained program in the general case. This reformulation is shown to be equivalent to the original formulation but allows the use of well-studied, efficient algorithms for multidimensional knapsack problems. In addition, it separates the modeling effort around what to put in the knapsack from considerations around what unit weight one should assign to each item type, providing substantially more flexibility to the modeler without adding complexity to the choice of knapsack configuration. Finally, we show that for the feasibility version of the second stage, one can immediately get a feasible solution to the first stage solution.

cs.DS

Simulation of the flow of an explosive atmosphere exposed to a hot surface

The accidental ignition of combustible atmospheres by hot surfaces is of great concern for chemical and process plant safety. In this paper, we present our research regarding the evolution of thermal plumes originating from hot hemispheres and discs. In particular, we focus on the effect of the orientation of the surface on the ignition process. The auto-ignition temperatures and ignition locations were studied experimentally. To get further insight, we conducted detailed numerical simulations and validated them with measurements. Three-dimensional simulations were performed on hot hemispheres and hot discs for different orientations ranging from 0{\deg} to 180{\deg}. The solver employs a transient, implicit scheme which is based on the coupled heat transfer and flow equations. The mesh in the vicinity of the hot surfaces is refined to resolve the steep temperature gradients and to capture the boundary layer separation. The influence of the orientation on critical hot spots in the gas mixture is analysed by examining the flow structures and the temperature evolution of the buoyancy-driven flow. Using the obtained results, we discuss the change of the onset and location of the ignition.

physics.flu-dyn

Fault Injectors for TensorFlow: Evaluation of the Impact of Random Hardware Faults on Deep CNNs

Today, Deep Learning (DL) enhances almost every industrial sector, including safety-critical areas. The next generation of safety standards will define appropriate verification techniques for DL-based applications and propose adequate fault tolerance mechanisms. DL-based applications, like any other software, are susceptible to common random hardware faults such as bit flips, which occur in RAM and CPU registers. Such faults can lead to silent data corruption. Therefore, it is crucial to develop methods and tools that help to evaluate how DL components operate under the presence of such faults. In this paper, we introduce two new Fault Injection (FI) frameworks InjectTF and InjectTF2 for TensorFlow 1 and TensorFlow 2, respectively. Both frameworks are available on GitHub and allow the configurable injection of random faults into Neural Networks (NN). In order to demonstrate the feasibility of the frameworks, we also present the results of FI experiments conducted on four VGG-based Convolutional NNs using two image sets. The results demonstrate how random bit flips in the output of particular mathematical operations and layers of NNs affect the classification accuracy. These results help to identify the most critical operations and layers, compare the reliability characteristics of functionally similar NNs, and introduce selective fault tolerance mechanisms.

cs.LG

Anomalous Quartic Gauge Couplings from Six Quark Production

The absence of a light Higgs boson causes vector boson couplings to become strong at 1 TeV. A general framework for a systematic and consistent treatment is provided by effective theories of electroweak symmetry breaking. Already in next-to-leading order there appear quartic gauge couplings that go beyond the standard model and are hence called anomalous. We investigate intermediate three gauge boson states $W^+W^- Z$ and $ZZZ$ occurring in six quark production in electron positron collisions under the conditions of the International Linear Collider. We perform a sensitivity analysis of the relevant anomalous quartic gauge couplings presenting their expected limits.

hep-ph

Linear square-mass trajectories of radially and orbitally excited hadrons in holographic QCD

We consider a new approach towards constructing approximate holographic duals of QCD from experimental hadron properties. This framework allows us to derive a gravity dual which reproduces the empirically found linear square-mass trajectories of universal slope for radially and orbitally excited hadrons. Conformal symmetry breaking in the bulk is exclusively due to infrared deformations of the anti-de Sitter metric and governed by one free mass scale proportional to Lambda_QCD. The resulting background geometry exhibits dual signatures of confinement and provides the first examples of holographically generated linear trajectories in the baryon sector. The predictions for the light hadron spectrum include new relations between trajectory slopes and ground state masses and are in good overall agreement with experiment.

hep-ph

Linear meson and baryon trajectories in AdS/QCD

An approximate holographic dual of QCD is constructed and shown to reproduce the empirical linear trajectories of universal slope on which the square masses of radially and orbitally excited hadrons join. Conformal symmetry breaking and other IR effects are described exclusively by deformations of the anti-de Sitter background metric. The predictions for the light hadron spectrum include new relations between ground state masses and trajectory slopes and are in good overall agreement with experimental data.

hep-ph

Light-front field theory of hot and dense quark matter

Extending the concepts of light-front field theory to quantum statistics provides a novel approach towards nuclear matter under extreme conditions. Such conditions exist, e.g., in neutron stars or in the early stage of our universe. They are experimentally expected to occur in heavy ion collisions, e.g., at RHIC and accelerators to be build at GSI and CERN. Light-front field theory is particularly suited, since it is based on a relativistic Hamiltonian approach. It allows us to treat the perturbative as well as the nonperturbative regime of QCD and also correlations that emerge as a field of few-body physics and is important for hadronization. Last but not least the Hamiltonian approach is useful for nonequilibrium processes by utilizing, e.g., the formalism of nonequilibrium statistical operators.

nucl-th

Clusters and condensates in Fermi systems

Superconductivity, superfluidity, condensation, cluster formation, etc. are phenomena that might occur in many-particle systems. These are due to residual interactions between the particles. To explain these phenomena consistently in a microscopic approach, at some point, one needs to solve few-body equations that are modified because of the Pauli principle and the interactions of the many particles around.

nucl-th

Few-body correlations in the QCD phase diagram

From the viewpoint of statistical physics, nuclear matter is a strongly correlated many-particle system. Several regimes of the QCD phase diagram should exhibit strong correlations. Here I focus on three- and four-body correlations that might be important in the phase diagram.

nucl-th

Few nucleon dynamics in a nuclear medium

Few body methods are used in many particle physics to describe correlations, bound states, and reactions in strongly correlated quantum systems. Although this has already been recognized earlier, rigorous attempts to treat three-body collisions have only been done recently. In this talk I shall give examples and areas where few-body methods have been and might be of use in the future.

nucl-th

Effective T-odd P-even hadronic interactions from quark models

Tests of time reversal symmetry at low and medium energies may be analyzed in the framework of effective hadronic interactions. Here, we consider the quark structure of hadrons to make a connection to the more fundamental degrees of freedom. It turns out that for P-even T-odd interactions hadronic matrix elements evaluated in terms of quark models give rise to factors of 2 to 5. Also, it is possible to relate the strength of the anomalous part of the effective rho-type T-odd P-even tensor coupling to quark structure effects.

nucl-th

Quark Structure and Weak Decays of Heavy Mesons

We investigate the quark structure of D and B mesons in the framework of a constituent quark model. To this end, we assume a scalar confining and a one gluon exchange (OGE) potential. The parameters of the model are adopted to reproduce the meson mass spectrum. From a fit to ARGUS and CLEO data on B->D*lv semileptonic decay we find for the Cabbibo Kobayashi Maskawa matrix element Vcb=0.036+-0.003. We compare our form factors to the pole dominance hypothesis and the heavy quark limit. For non-leptonic decays we utilize factorization and for B->D*X decays we find a1 = 0.96+-0.05, and a2=0.31+-0.03.

hep-ph

Model Analysis of Time Reversal Symmetry Test in the Caltech Fe-57 Gamma-Transition Experiment

The CALTECH gamma-transition experiment testing time reversal symmetry via the E2/M1 mulipole mixing ratio of the 122 keV gamma-line in Fe-57 has already been performed in 1977. Extending an earlier analysis in terms of an effective one-body potential, this experiment is now analyzed in terms of effective one boson exchange T-odd P-even nucleon nucleon potentials. Within the model space considered for the Fe-57 nucleus no contribution from isovector rho-type exchange is possible. The bound on the coupling strength phi_A from effective short range axial-vector type exchange induced by the experimental bound on sin(eta) leads to phi_A < 10^{-2}.

nucl-th

Test of Time Reversal Invariance in Proton Deuteron System

Internal target experiments with high quality proton beams allow for a new class of experiments providing null tests of time reversal symmetry in forward scattering. This could yield more stringent limits on T-odd P-even observables. A excellent candidate for such experiments is the proton deuteron system. This system is analyzed in terms of effective T-violating P-conserving nucleon-nucleon interactions and bounds on coupling strengths that might be expected are given.

nucl-th