SearcharxivSearch

arXiv subjects

Fangyuan Yu

Publications and source records attributed to Fangyuan Yu.

16 recordsLinked to original sources

Hybrid Verified Decoding: Learning to Allocate Verification in Speculative Decoding

Large Language Model (LLM) generation remains expensive because autoregressive decoding calls the model once for each new token. Speculative decoding reduces this cost by drafting multiple tokens and verifying them with the target model in one step, but its speedup depends on how many drafted tokens are accepted. Parameter-free draft sources can propose long continuations at low cost in structured and agentic workloads, yet a cache match that looks promising at one generation step may have low payoff at the next. We propose Hybrid Verified Decoding, which predicts the accepted length of a cache draft before verification and uses this payoff estimate to choose between cache verification and a model-based drafter. Across three LLMs and sixteen datasets, Hybrid Verified Decoding is especially effective on agentic workflows, where it outperforms EAGLE3 in every setting with a 2.73x average speedup. Our analysis shows how prompt structure creates cache opportunities, how high-payoff cache drafts concentrate in a small part of the draft space, and how payoff-guided selection reduces sequential decoding work, pointing to runtime draft selection as a promising direction for speculative decoding.

cs.CL

Dynamic Latent Routing

We investigate the temporal concatenation of sub-policies in Markov Decision Processes (MDP) with time-varying reward functions. We introduce General Dijkstra Search (GDS), and prove that globally optimal goal-reaching policies can be recovered through temporal composition of intermediate optimal sub-policies. Motivated by the "search, select, update" principle underlying GDS, we propose Dynamic Latent Routing (DLR), a language-model post-training method that jointly learns discrete latent codes, routing policies, and model parameters through dynamic search in a single training stage. In low-data fine-tuning settings, DLR matches or outperforms supervised fine-tuning across four datasets and six models, achieving a mean gain of +6.6 percentage points, while prior discrete-latent baselines consistently underperform SFT. Mechanistic analyses and targeted code ablations show that DLR learns structured routing behaviors with distinct causal roles.

cs.LG

Multi-wavelength ALMA Imaging of HD 34282: Dust-trapping Signatures of a Vortex Candidate

Azimuthal arcs in millimeter continuum emission from protoplanetary disks are often attributed to dust-trapping vortices, but definitive observational confirmation of vortices remains lacking. We present sub-0.1" resolution ALMA continuum observations of the HD 34282 disk at 0.9, 1.3, 2.1, and 3.1 mm. These observations resolve a bright azimuthal arc superposed on a compact double-gap, triple-ring morphology, most clearly at shorter wavelengths, and enable us to probe the physical origin of the arc. It exhibits a lower spectral index than the surrounding rings, consistent with enhanced grain growth and/or higher dust surface density of a dust-trapping vortex. Its azimuthal width decreases with increasing wavelength, consistent with tighter confinement of larger grains, or lower optical depths at longer wavelengths. These observations probe dust with Stokes numbers St < 0.03. Vortex models predict negligible peak shifts in this regime, consistent with the 1.3 to 3.1 mm data. At 0.9 mm, however, the arc peak is offset by 15 +/- 4 degree in the direction of disk rotation relative to longer wavelengths, and the near-side ring emission is locally dimmer compared to the far-side, likely reflecting optical-depth or temperature effects. These observations are consistent with azimuthal dust trapping, potentially associated with a vortex-induced pressure maximum.

astro-ph.EP

Memorization-Compression Cycles Improve Generalization

We prove theoretically that generalization improves not only through data scaling but also by compressing internal representations. To operationalize this insight, we introduce the Information Bottleneck Language Modeling (IBLM) objective, which reframes language modeling as a constrained optimization problem: minimizing representation entropy subject to optimal prediction performance. Empirically, we observe an emergent memorization-compression cycle during LLM pretraining, evidenced by oscillation positive/negative gradient alignment between cross-entropy and Matrix-Based Entropy (MBE), a measure of representation entropy. This pattern closely mirrors the predictive-compressive trade-off prescribed by IBLM and also parallels the biological alternation between awake learning and sleep consolidation. Motivated by this observation, we propose Gated Phase Transition (GAPT), a training algorithm that adaptively switches between memorization and compression phases. When applied to GPT-2 pretraining on FineWeb dataset, GAPT reduces MBE by 50% and improves cross-entropy by 4.8%. GAPT improves OOD generalizatino by 35% in a pretraining task on arithmetic multiplication. In a setting designed to simulate catastrophic forgetting, GAPT reduces interference by compressing and separating representations, achieving a 97% improvement in separation - paralleling the functional role of sleep consolidation.

cs.LG

Binary Stars Approaching Supermassive Black Holes: Hydrodynamics of Stellar Collisions, Mass Fallback and Partial TDEs

When binaries are injected into low-angular-momentum orbits around a central supermassive black hole (SMBH), various outcomes can occur, including binary tidal breakup, double stellar disruptions and stellar collision. We use hydrodynamical simulations to study stellar collisions triggered by binary-SMBH encounters, examining both head-on and grazing collisions in deep ($\beta_b=5$) and gentle ($\beta_b=0.6$) encounters, where $\beta_b$ is the ratio of the binary tidal disruption radius to the binary pericenter distance to the SMBH. Head-on collisions consistently result in appreciable mass loss ($\sim 5\%$) and a single merger remnant. Grazing collisions have varied outcomes. In gentle encounters, multiple collisions typically form a single remnant with minimal mass loss ($\lesssim 1 \%$). For deep encounters, the result depends on the specific collision parameters and stellar structure: $\gamma=5/3$ polytropic stars in our simulation produced two disturbed remnants, while solar-type stars (modeled with MESA) in our deep-grazing run formed a single merger remnant in a low-velocity collision. All merger remnants feature extended envelopes, making them susceptible to partial tidal disruptions when they return to the SMBH. The morphology and orbital energy distribution of collision-induced debris differ significantly from those of tidal disruption event (TDE) debris of single stars. Approximately half of the collision-generated debris falls back onto the SMBH, exhibiting a distinct time evolution of the fallback rate. We suggest that such mass loss and fallback can generate electromagnetic flares that mimic weak TDEs.

astro-ph.HE

Leaky Dust Traps in Planet-Embedded Protoplanetary Disks

From the survival of dust disks for a few Myr to the establishment of chemical dichotomy, dust traps are expected to play a pivotal role in sculpting protoplanetary disks and the early planet formation process. These traps however may not be perfect as evidenced by the detection of gas and dust inside the gaps and cavities of structured disks. Using two-fluid hydrodynamic global simulations in both two-dimensions (2D) and three-dimensions (3D), we directly compute the dynamics of dust grains as they aerodynamically interact with the disk gas that is being perturbed by an embedded planet of varying mass. In both 2D and 3D, we find the dust trap to be more leaky for lower mass planet and for higher turbulent $\alpha$. More crucially, we find the fraction of the dust mass that remain trapped within the pressure bump can be up to an order of magnitude more reduced in 3D vs. 2D with all else equal. Our simulations show a complex behavior of dust radial motion that is both azimuthally and poloidally non-uniform, with the overall dynamics dominated by the dust coupling to the gas flow even for relatively high St = 0.1. The leaky traps we find suggest pebble isolation mass is likely not truly isolating and that gap-opening planets do not establish as an unconditional impermeable barrier. Our findings have implications for recent JWST MINDS results, which show that volatiles, including water, are present in the inner regions of disks hosting outer dust rings.

astro-ph.EP

Scaling LLM Pre-training with Vocabulary Curriculum

Modern language models rely on static vocabularies, fixed before pretraining, in contrast to the adaptive vocabulary acquisition observed in human language learning. To bridge this gap, we introduce vocabulary curriculum learning, an approach that improves pretraining efficiency with log-linear scaling gains relative to vocabulary size. Our method alternates between entropy-guided vocabulary expansion and model optimization, enabling models to learn transferable representations across diverse tokenization granularities. This approach naturally gives rise to an optimal computation allocation pattern: longer tokens capture predictable content, while shorter tokens focus on more complex, harder-to-predict contexts. Experiments on small-scale GPT models demonstrate improved scaling efficiency, reinforcing the effectiveness of dynamic tokenization. We release our code to support further research and plan to extend our experiments to larger models and diverse domains.

cs.CL

Binary Stars Approaching Supermassive Black Holes: Tidal Break-up, Double Stellar Disruptions and Stellar Collision

In galactic centers, stars and binaries can be injected into low-angular-momentum orbits, resulting in close encounters with the central supermassive black hole (SMBH). Previous works have shown that under different conditions, such close encounters can lead to the break-up of the binary, disruptions of both stars and collision between the stars. We use 3-body scattering experiments to characterize these different outcomes for a range of system parameters, such as $\beta_b$, the ratio of binary tidal radius to pericenter distance $r_p$ to the SMBH and the compactness of the binary. We focus on stellar collisions, which occur for a range of $\beta_b$'s, with a few to 10's percent probabilities (depending on the compactness of the binary). In gentle encounters ($\beta_b\lesssim 1$), stellar collisions occur after the pericenter passage, and the merger remnants are typically ejected from the SMBH at a small velocity. In deep encounters ($\beta_b\gtrsim 1$), collisions occur near the pericenter, with the impact velocity a few times the escape velocity of the star, and the merger remnants are typically bound to the SMBH. We suggest that stellar collisions induced by binary-SMBH encounters may produce exotic stars in galactic centers, trigger accretion flares onto the SMBH due to the mass loss, and result in bound merger remnants causing repeated partial TDEs.

astro-ph.HE

Iterative Graph Alignment

By compressing diverse narratives, LLMs go beyond memorization, achieving intelligence by capturing generalizable causal relationships. However, they suffer from local 'representation gaps' due to insufficient training data diversity, limiting their real-world utility, especially in tasks requiring strict alignment to rules. Traditional alignment methods relying on heavy human annotations are inefficient and unscalable. Recent self-alignment techniques also fall short, as they often depend on self-selection based prompting and memorization-based learning. To address these issues, we introduce Iterative Graph Alignment (IGA), an annotation-free rule-based alignment algorithm. A teacher model (VLM) employs Iterative Graph Prompting (IGP) to create logical graphs and reference answers. The student model (LLM) identifies local knowledge gaps by attempting to align its responses with these references, collaborating with helper models to generate diverse answers. These aligned responses are then used for iterative supervised fine-tuning (SFT). Our evaluations across five rule-based scenarios demonstrate IGP's effectiveness, with a 73.12\% alignment improvement in Claude Sonnet 3.5, and Llama3-8B-Instruct achieving an 86.20\% improvement, outperforming Claude Sonnet 3.5 in rule-based alignment.

cs.LG

Free-Floating Planets, Survivor Planets, Captured Planets and Binary Planets from Stellar Flybys

In star clusters, close stellar encounters can strongly impact the architecture of a planetary system or even destroy it. We present a systematic study on the effects of stellar flybys on two-planet systems. When such a system experiences flybys, one or both planets can be ejected, forming free-floating planets (FFPs), captured planets (CPs) around the flyby star, and free-floating binary planets (BPs); the remaining single-surviving-planets (SSPs) can have their orbital radii and eccentricities greatly changed. Through numerical experiments, we calculate the formation fractions (or branching ratios) of FFPs, SSPs, CPs and BPs as a function of the pericenter distance of the flyby, and use them to derive analytical expressions for the formation rates of FFPs, SSPs, CPs and BPs in general cluster environments. We find that the production rates of FFPs and SSPs are similar (for initial planet semi-major axis ratio $a_1/a_2=0.6-0.8$), while the rate for CPs is a few times smaller. The formation fraction of BPs depends strongly on $a_1/a_2$ and on the planet masses. For Jupiter-mass planets, the formation fraction of BPs is always less than $1\%$ (for $a_1/a_2=0.8$) and typically much smaller ($\lesssim 0.2\%$ for $a_1/a_2\lesssim 0.7$). The fraction remains less than $1\%$ when considering $4M_{\rm J}$ planets. Overall, when averaging over all flybys, the production rate of BPs is less than $0.1\%$ of that for FFPs. We also derive the velocity distribution of FFPs produced by stellar flybys, and the orbital parameter distributions of SSPs, CPs and BPs. These results can be used in future studies of exotic planets (including FFPs) and planetary systems.

astro-ph.EP

Unbiased Estimation of the Hessian for Partially Observed Diffusions

In this article we consider the development of unbiased estimators of the Hessian, of the log-likelihood function with respect to parameters, for partially observed diffusion processes. These processes arise in numerous applications, where such diffusions require derivative information, either through the Jacobian or Hessian matrix. As time-discretizations of diffusions induce a bias, we provide an unbiased estimator of the Hessian. This is based on using Girsanov's Theorem and randomization schemes developed through Mcleish [2011] and Rhee & Glynn [2015]. We demonstrate our developed estimator of the Hessian is unbiased, and one of finite variance. We numerically test and verify this by comparing the methodology here to that of a newly proposed particle filtering methodology. We test this on a range of diffusion models, which include different Ornstein--Uhlenbeck processes and the Fitzhugh--Nagumo model, arising in neuroscience.

stat.ME

Randomized multilevel Monte Carlo for embarrassingly parallel inference

This position paper summarizes a recently developed research program focused on inference in the context of data centric science and engineering applications, and forecasts its trajectory forward over the next decade. Often one endeavours in this context to learn complex systems in order to make more informed predictions and high stakes decisions under uncertainty. Some key challenges which must be met in this context are robustness, generalizability, and interpretability. The Bayesian framework addresses these three challenges elegantly, while bringing with it a fourth, undesirable feature: it is typically far more expensive than its deterministic counterparts. In the 21st century, and increasingly over the past decade, a growing number of methods have emerged which allow one to leverage cheap low-fidelity models in order to precondition algorithms for performing inference with more expensive models and make Bayesian inference tractable in the context of high-dimensional and expensive models. Notable examples are multilevel Monte Carlo (MLMC), multi-index Monte Carlo (MIMC), and their randomized counterparts (rMLMC), which are able to provably achieve a dimension-independent (including $\infty-$dimension) canonical complexity rate with respect to mean squared error (MSE) of $1/$MSE. Some parallelizability is typically lost in an inference context, but recently this has been largely recovered via novel double randomization approaches. Such an approach delivers i.i.d. samples of quantities of interest which are unbiased with respect to the infinite resolution target distribution. Over the coming decade, this family of algorithms has the potential to transform data centric science and engineering, as well as classical machine learning applications such as deep learning, by scaling up and scaling out fully Bayesian inference.

stat.CO

Multilevel Ensemble Kalman-Bucy Filters

In this article we consider the linear filtering problem in continuous-time. We develop and apply multilevel Monte Carlo (MLMC) strategies for ensemble Kalman-Bucy filters (EnKBFs). These filters can be viewed as approximations of conditional McKean-Vlasov-type diffusion processes. They are also interpreted as the continuous-time analogue of the \textit{ensemble Kalman filter}, which has proven to be successful due to its applicability and computational cost. We prove that an ideal version of our multilevel EnKBF can achieve a mean square error (MSE) of $\mathcal{O}(\epsilon^2), \ \epsilon>0$ with a cost of order $\mathcal{O}(\epsilon^{-2}\log(\epsilon)^2)$. In order to prove this result we provide a Monte Carlo convergence and approximation bounds associated to time-discretized EnKBFs. This implies a reduction in cost compared to the (single level) EnKBF which requires a cost of $\mathcal{O}(\epsilon^{-3})$ to achieve an MSE of $\mathcal{O}(\epsilon^2)$. We test our theory on a linear problem, which we motivate through high-dimensional examples of order $\sim \mathcal{O}(10^4)$ and $\mathcal{O}(10^5)$.

math.NA

Unbiased Filtering of a Class of Partially Observed Diffusions

In this article we consider a Monte Carlo-based method to filter partially observed diffusions observed at regular and discrete times. Given access only to Euler discretizations of the diffusion process, we present a new procedure which can return online estimates of the filtering distribution with no discretization bias and finite variance. Our approach is based upon a novel double application of the randomization methods of Rhee & Glynn (2015) along with the multilevel particle filter (MLPF) approach of Jasra et al (2017). A numerical comparison of our new approach with the MLPF, on a single processor, shows that similar errors are possible for a mild increase in computational cost. However, the new method scales strongly to arbitrarily many processors.

math.NA

Multilevel Particle Filters for the Non-Linear Filtering Problem in Continuous Time

In the following article we consider the numerical approximation of the non-linear filter in continuous-time, where the observations and signal follow diffusion processes. Given access to high-frequency, but discrete-time observations, we resort to a first order time discretization of the non-linear filter, followed by an Euler discretization of the signal dynamics. In order to approximate the associated discretized non-linear filter, one can use a particle filter (PF). Under assumptions, this can achieve a mean square error of $\mathcal{O}(\epsilon^2)$, for $\epsilon>0$ arbitrary, such that the associated cost is $\mathcal{O}(\epsilon^{-4})$. We prove, under assumptions, that the multilevel particle filter (MLPF) of Jasra et al (2017) can achieve a mean square error of $\mathcal{O}(\epsilon^2)$, for cost $\mathcal{O}(\epsilon^{-3})$. This is supported by numerical simulations in several examples.

math.NA

Central Limit Theorems for Coupled Particle Filters

In this article we prove a new central limit theorem (CLT) for coupled particle filters (CPFs). CPFs are used for the sequential estimation of the difference of expectations w.r.t. filters which are in some sense close. Examples include the estimation of the filtering distribution associated to different parameters (finite difference estimation) and filters associated to partially observed discretized diffusion processes (PODDP) and the implementation of the multilevel Monte Carlo (MLMC) identity. We develop new theory for CPFs and based upon several results, we propose a new CPF which approximates the maximal coupling (MCPF) of a pair of predictor distributions. In the context of ML estimation associated to PODDP with discretization $\Delta_l$ we show that the MCPF and the approach in Jasra et al. (2018) have, under assumptions, an asymptotic variance that is upper-bounded by an expression that is (almost) $\mathcal{O}(\Delta_l)$, uniformly in time. The $\mathcal{O}(\Delta_l)$ rate preserves the so-called forward rate of the diffusion in some scenarios which is not the case for the CPF in Jasra et al (2017).

math.ST