SearcharxivSearch

arXiv subjects

Eric Sather

Publications and source records attributed to Eric Sather.

10 recordsLinked to original sources

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation

Speculative decoding (SD) has proven to be an effective technique for accelerating autoregressive generation in large language models (LLMs) however, its application to vision-language models (VLMs) remains relatively unexplored. We propose~\textit{DREAM-S}, a novel SD framework designed specifically for fast and efficient decoding in VLMs. DREAM-S leverages a neural architecture search (NAS) framework with target-aware supernet training to automatically identify both the optimal interaction strategy between the draft and target models, and the most suitable draft model architecture for the underlying hardware implementation platform. DREAM-S additionally incorporates adaptive intermediate feature distillation, guided by attention entropy, to enable efficient draft training. Experiments on a range of well-established VLMs show that DREAM-S achieves up to a $3.85\times$ speedup compared to standard decoding approaches and significantly outperforms existing SD baselines. The code is publicly available at: https://github.com/SAI-Lab-NYU/DREAM-S .

cs.LG

DREAM-R: Multimodal Speculative Reasoning with RL-Based Refined Drafting, Precise Verification, and Fully Parallel Execution

Speculative reasoning has recently been proposed as a means to accelerate reasoning-intensive generation in large multimodal models, but its effectiveness is often constrained by misalignment between speculative drafts and target-verified reasoning. In this work, we introduce DREAM-R, a framework that substantially improves the performance of speculative reasoning. At its core, DREAM-R employs Speculative Alignment Policy Optimization (SAPO), a reinforcement-learning objective that trains draft models to generate reasoning steps that are both faithful to target trajectories and concise. We further propose a Threshold-based Verification Mechanism (TBVM) that uses a ratio-based criterion to provide stable and interpretable acceptance of speculative steps only when positive evidence clearly dominates, thereby preventing error propagation. Building on these components, we develop a Fully Parallel Speculative Reasoning (FPSR) framework that parallelizes draft generation, target-side reasoning, and verification across multi-step reasoning, enabling early stopping and clean fallback. Experiments on reasoning-heavy benchmarks demonstrate up to speedup while preserving target-model accuracy, yielding substantial efficiency gains without compromising reasoning quality.

cs.AI

CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts

Outliers have emerged as a fundamental bottleneck in preserving accuracy for low-precision large models, particularly within Mixture-of-Experts (MoE) architectures that are increasingly central to large-scale language modeling. Under post-training quantization (PTQ), these outliers induce substantial quantization errors, leading to severe accuracy degradation. While recent rotation-based smoothing techniques alleviate the problem by redistributing outlier magnitudes, residual errors remain and continue to impede reliable low-precision deployment. In this work, we tackle this challenge by introducing CodeQuant, a unified quantization-and-clustering scheme that contains smoothing activation outliers via learnable rotation and absorbing weight outliers into fine-tuned cluster centroids for MoE. This design reduces the influence of extreme values by fitting them within cluster centroids, thereby lowering quantization error while maintaining expressive capacity. Coupled with a dedicated kernel design for GPU and CPU, CodeQuant achieves up to $4.15\times$ speedup while delivering significantly higher accuracy than state-of-the-art quantization approaches across diverse MoE models. Our results highlight CodeQuant as a promising direction for efficient and accurate deployment of MoE-based large language models under low-precision constraints. Our code is available at https://github.com/SAI-Lab-NYU/CodeQuant.

cs.LG

DREAM: Drafting with Refined Target Features and Entropy-Adaptive Cross-Attention Fusion for Multimodal Speculative Decoding

Speculative decoding (SD) has emerged as a powerful method for accelerating autoregressive generation in large language models (LLMs), yet its integration into vision-language models (VLMs) remains underexplored. We introduce DREAM, a novel speculative decoding framework tailored for VLMs that combines three key innovations: (1) a cross-attention-based mechanism to inject intermediate features from the target model into the draft model for improved alignment, (2) adaptive intermediate feature selection based on attention entropy to guide efficient draft model training, and (3) visual token compression to reduce draft model latency. DREAM enables efficient, accurate, and parallel multimodal decoding with significant throughput improvement. Experiments across a diverse set of recent popular VLMs, including LLaVA, Pixtral, SmolVLM and Gemma3, demonstrate up to 3.6x speedup over conventional decoding and significantly outperform prior SD baselines in both inference throughput and speculative draft acceptance length across a broad range of multimodal benchmarks. The code is publicly available at: https://github.com/SAI-Lab-NYU/DREAM.git

cs.CL

Constraining the Strongly-Coupled Standard Model Including a W' Isotriplet

We consider the Strongly-Coupled Standard Model (Abbott-Farhi model) including an isotriplet of W' vector bosons. First we calculate the corrections to the low-energy theory, which can be effectively summarized in terms of the parameters S, T and U. Then we use high- precision electroweak measurements to constrain the mass and couplings of the W'. The W' couplings are restricted to be unnaturally small, and we conclude that this model is no longer compelling as a theory of the electroweak interactions.

hep-ph

Electroweak Baryogenesis and Standard Model CP Violation

We analyze the mechanism of electroweak baryogenesis proposed by Farrar and Shaposhnikov in which the phase of the CKM mixing matrix is the only source of $CP$ violation. This mechanism is based on a phase separation of baryons via the scattering of quasiparticles by the wall of an expanding bubble produced at the electroweak phase transition. In agreement with the recent work of Gavela, Hernández, Orloff and Pène, we conclude that QCD damping effects reduce the asymmetry produced to a negligible amount. We interpret the damping as quantum decoherence. We compute the asymmetry analytically. Our analysis reflects the observation that only a thin, outer layer of the bubble contributes to the coherent scattering of the quasiparticles. The generality of our arguments rules out any mechanism of electroweak baryogenesis that does not make use of a new source of $CP$ violation.

hep-ph

The Rate for $B\bbar$ Production Accompanied by a Single Pion

We study the rate for the production of ${B B^\pm π^\mp}$, where the sign of the charged pion tags the flavor content of the neutral $B$ meson. We estimate this branching ratio, employing the heavy meson chiral effective theory. We find that at center of mass energy of approximately 12 GeV, a $B$ meson pair should be produced as often with and without an accompanying charged pion. We also calculate two pion production at this center of mass energy, and find that it is negligible, as is the rate for rho production. We consider the implications for CP violation studies. (Invited Talk, 1993 Electroweak Rencontres de Moriond)

hep-ph

The Rate for $e^+e^-\to B B^\pm π^\mp$ and its Implications for the Study of CP Violation, $B_s$ Identification, and the Study of $B$ Meson Chiral Perturbation Theory

H.~Yamamoto has proposed employing $B$ mesons produced in conjunction with a single charged pion at an $Υ$ resonance for studies of CP violation in the neutral $B$ meson system at a symmetric $e^+$-$e^-$ collider. The sign of the charged pion would tag the neutral $B$ meson. We estimate this branching ratio, employing the heavy meson chiral effective field theory. We find a negligible branching ratio to $B B^{\pm} π^{\mp}$ at the $Υ$(5S) and a branching ratio of only a few percent at the $Υ$(6S). However, if nonresonant studies of neutral $B$ mesons should prove feasible, Yamamoto's proposal could be a good method for tagging neutral $B$'s for the study of CP violation at a symmetric collider. We also explore the possibility of studying $B_s$ at the $Υ$(5S). The rate is low but depends sensitively on the precise value of the mass of the $B_s$. The background we compute is comparable to the rate at the largest allowed value of the $B_s$ mass. Finally, we discuss the extraction of the axial pion coupling to $B$ mesons from measurement of the $B\bbarπ$ branching fraction in a restricted region of phase space, where chiral perturbation theory should work well.

hep-ph

Heavy Meson Hyperfine Splittings: A Puzzle for Heavy Quark Chiral Perturbation Theory

We show that there is a large discrepancy between the expected light flavor dependence of the heavy pseudoscalar--vector mass splittings and the measured values. We demonstrate that the one--loop calculation is unreliable. Moreover, agreement with experiment requires the leading dependence on SU(3) symmetry breaking to be nearly cancelled, so that the heavy quark mass dependence is unknown and the expected dependence on the light quark mass is not realized.

hep-ph

The QCD Scale in the Heavy Quark Expansion

We argue that consistency of the combined heavy quark and chiral effective lagrangian requires the QCD scale which multiplies $1/M$ in the heavy quark expansion to be the chiral symmetry breaking scale, $Λ_{CSB}$, rather than the QCD scale, $Λ_{QCD}$. This means that either there is large uncertainty in the accuracy with which the heavy quark effective theory can be applied to $c$ quarks or the cutoff scale of the heavy quark chiral effective theory is lower than has been assumed.

hep-ph