SearcharxivSearch

arXiv subjects

Maximilian Jeblick

Publications and source records attributed to Maximilian Jeblick.

9 recordsLinked to original sources

LLM Router: Rethinking Routing with Prefill Activations

Existing routers rely on semantic query features or handcrafted features, which often fail to capture model-specific failures or intrinsic task difficulty. We instead route using internal LLM activations, specifically the residual stream. Our key idea, Encoder-Target Decoupling, separates the model that produces the predictive signal (the Encoder) from the model whose correctness is being estimated (the Target), allowing open-weight encoders to predict the performance of closed-source target models. We evaluate layerwise geometric probes, finding that Fisher Separability ($J$) effectively identifies informative layers, supported by Effective Dimensionality ($d_{\mathrm{eff}}$) diagnostics. We then utilize a SharedTrunkNet, a joint multi-output MLP that predicts simultaneous correctness probabilities across candidate models using concatenated prefill features. In our experiments, SharedTrunkNet consistently outperforms semantic baselines. At its best, SharedTrunkNet closes 45.58% of the gap between the strongest standalone model and the oracle while achieving 74.31% cost savings relative to the most expensive model. These results demonstrate that prefill activations provide a robust routing signal, establishing activation-based routing as a high-performance alternative to purely semantic selection.

cs.CL

KVzap: Fast, Adaptive, and Faithful KV Cache Pruning

Growing context lengths in transformer-based language models have made the key-value (KV) cache a critical inference bottleneck. While many KV cache pruning methods have been proposed, they have not yet been adopted in major inference engines due to speed--accuracy trade-offs. We introduce KVzap, a fast, input-adaptive approximation of KVzip that works in both prefilling and decoding. On Qwen3-8B, Llama-3.1-8B-Instruct, and Qwen3-32B across long-context and reasoning tasks, KVzap achieves $2$--$4\times$ KV cache compression with negligible accuracy loss and achieves state-of-the-art performance on the KVpress leaderboard. Code and models are available at https://github.com/NVIDIA/kvpress.

cs.LG

Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution

Memory consumption of the Key-Value (KV) cache represents a major bottleneck for efficient large language model inference. While attention-score-based KV cache pruning shows promise, it faces critical practical limitations: attention scores from future tokens are unavailable during compression, and modern implementations like Flash Attention do not materialize the full attention matrix, making past scores inaccessible. To overcome these challenges, we introduce $\textbf{Expected Attention, a training-free compression method}$ that estimates KV pairs importance by predicting how future queries will attend to them. Our approach leverages the distributional properties of LLM activations to compute expected attention scores in closed form for each KV pair. These scores enable principled ranking and pruning of KV pairs with minimal impact on the residual stream, achieving effective compression without performance degradation. Importantly, our method operates seamlessly across both prefilling and decoding phases, consistently outperforming state-of-the-art baselines in both scenarios. Finally, $\textbf{we release KVPress, a comprehensive library to enable researchers to implement and benchmark KV cache compression methods, already including more than 20 techniques}$.

cs.AI

H2O-Danube-1.8B Technical Report

We present H2O-Danube, a series of small 1.8B language models consisting of H2O-Danube-1.8B, trained on 1T tokens, and the incremental improved H2O-Danube2-1.8B trained on an additional 2T tokens. Our models exhibit highly competitive metrics across a multitude of benchmarks and, as of the time of this writing, H2O-Danube2-1.8B achieves the top ranking on Open LLM Leaderboard for all models below the 2B parameter range. The models follow core principles of LLama 2 and Mistral, and we leverage and refine various techniques for pre-training large language models. We additionally release chat models trained with supervised fine-tuning followed by direct preference optimization. We make all models openly available under Apache 2.0 license further democratizing LLMs to a wider audience economically.

cs.CL

H2O Open Ecosystem for State-of-the-art Large Language Models

Large Language Models (LLMs) represent a revolution in AI. However, they also pose many significant risks, such as the presence of biased, private, copyrighted or harmful text. For this reason we need open, transparent and safe solutions. We introduce a complete open-source ecosystem for developing and testing LLMs. The goal of this project is to boost open alternatives to closed-source approaches. We release h2oGPT, a family of fine-tuned LLMs of diverse sizes. We also introduce H2O LLM Studio, a framework and no-code GUI designed for efficient fine-tuning, evaluation, and deployment of LLMs using the most recent state-of-the-art techniques. Our code and models are fully open-source. We believe this work helps to boost AI development and make it more accessible, efficient and trustworthy. The demo is available at: https://gpt.h2o.ai/

cs.CL

h2oGPT: Democratizing Large Language Models

Applications built on top of Large Language Models (LLMs) such as GPT-4 represent a revolution in AI due to their human-level capabilities in natural language processing. However, they also pose many significant risks such as the presence of biased, private, or harmful text, and the unauthorized inclusion of copyrighted material. We introduce h2oGPT, a suite of open-source code repositories for the creation and use of LLMs based on Generative Pretrained Transformers (GPTs). The goal of this project is to create the world's best truly open-source alternative to closed-source approaches. In collaboration with and as part of the incredible and unstoppable open-source community, we open-source several fine-tuned h2oGPT models from 7 to 40 Billion parameters, ready for commercial use under fully permissive Apache 2.0 licenses. Included in our release is 100\% private document search using natural language. Open-source language models help boost AI development and make it more accessible and trustworthy. They lower entry hurdles, allowing people and groups to tailor these models to their needs. This openness increases innovation, transparency, and fairness. An open-source strategy is needed to share AI benefits fairly, and H2O.ai will continue to democratize AI and LLMs.

cs.CL

Derivation of the Time Dependent Gross-Pitaevskii Equation in Two Dimensions

We present a microscopic derivation of the defocusing two-dimensional cubic nonlinear Schrödinger equation as a mean field equation starting from an interacting $N$-particle system of Bosons. We consider the interaction potential to be given either by $W_β(x)=N^{-1+2 β}W(N^βx)$, for any $β>0$, or to be given by $V_N(x)=e^{2N} V(e^N x)$, for some spherical symmetric, positive and compactly supported $W,V \in L^\infty(\mathbb{R}^2,\mathbb{R})$. In both cases we prove the convergence of the reduced density matrix corresponding to the exact time evolution to the projector onto the solution of the corresponding nonlinear Schrödinger equation in trace norm. For the latter potential $V_N$ we show that it is crucial to take the microscopic structure of the condensate into account in order to obtain the correct dynamics.

math-ph

Derivation of the time dependent Gross-Pitaevskii equation for a class of non purely positive potentials

We present a microscopic derivation of the time-dependent Gross-Pitaevskii equation starting from an interacting N-particle system of Bosons. We prove convergence of the reduced density matrix corresponding to the exact time evolution to the projector onto the solution of the respective Gross-Pitaevskii equation. Our work extends a previous result by one of us (P.P.[44]) to interaction potentials which need not to be nonnegative, but may have a sufficiently small negative part. One key estimate in our proof is an operator inequality which was first proven by Jun Yin, see [49].

math-ph

Free time evolution of a tracer particle coupled to a Fermi gas in the high-density limit

The dynamics of a particle coupled to a dense and homogeneous ideal Fermi gas in two spatial dimensions is studied. We analyze the model for coupling parameter g=1 (i.e., not in the weak coupling regime), and prove closeness of the time evolution to an effective dynamics for large densities of the gas and for long time scales of the order of some power of the density. The effective dynamics is generated by the free Hamiltonian with a large but constant energy shift which is given at leading order by the spatially homogeneous mean field potential of the gas particles. Here, the mean field approximation turns out to be accurate although the fluctuations of the potential around its mean value can be arbitrarily large. Our result is in contrast to a dense bosonic gas in which the free motion of a tracer particle would be disturbed already on a very short time scale. The proof is based on the use of strong phase cancellations in the deviations of the microscopic dynamics from the mean field time evolution.

math-ph