SearcharxivSearch

arXiv subjects

Olabode M. Sule

Publications and source records attributed to Olabode M. Sule.

4 recordsLinked to original sources

Zamba2-VL Technical Report

We present Zamba2-VL, a suite of vision-language models built on Zamba2, a hybrid language-model architecture combining Mamba2 state-space layers with a small number of shared transformer blocks. Across a broad range of image understanding, reasoning, OCR, grounding, and counting benchmarks, Zamba2-VL is competitive with leading Transformer-based open-weight VLMs of comparable scale, including the Molmo2, Qwen3-VL, and InternVL3.5 families, and substantially outperforms prior SSM-based and hybrid VLMs such as VL-Mamba, Cobra, and mmMamba. Inheriting the near-linear prefill compute and small, near-constant recurrent state of its Zamba2 backbone, Zamba2-VL delivers roughly an order of magnitude lower time-to-first-token (TTFT) than these Transformer baselines at matched parameter scale, with the efficiency gap most pronounced at the smaller 1.2B and 2.7B scales most relevant to on-device and edge deployment. We release three models -- 1.2B, 2.7B, and 7B -- together with inference code at https://huggingface.co/collections/Zyphra/zamba2-vl.

cs.CV

ZAYA1-VL-8B Technical Report

We present ZAYA1-VL-8B, a compact mixture-of-experts vision-language model built upon our in-house language model, ZAYA1-8B. Despite its compact size, ZAYA1-VL achieves performance competitive with leading base models such as Molmo2-4B and InternVL3.5-4B, while surpassing models including Qwen2.5-VL-3B, PLM-3B, and MolmoE-1B across a range of image understanding, reasoning, and counting benchmarks. The architecture incorporates two key innovations: (1) vision-specific LoRA adapters integrated into the LLM to increase modality-specific capacity without increasing the number of experts, and (2) bidirectional attention over image tokens within the LLM to enhance visual understanding. We detail the full training pipeline including data composition at each stage, sequence packing, and the attention masking scheme. The model comprises 9.2B total parameters, with 1.4B active parameters including the vision encoder, and is publicly available at https://huggingface.co/Zyphra/ZAYA1-VL.

cs.CV

Determination of Tomonaga-Luttinger parameters for a two-component liquid

We provide evidence for the mapping of critical spin-1 chains, in particular the SU(3) symmetric bilinear-biquadratic model with additional interactions, to free boson theories using exact diagonalization and the density matrix renormalization group algorithm. Using the correspondence with a conformal field theory with central charge c=2, we determine the analytic formulae for the scaling dimensions in terms of four Tomonaga-Luttinger liquid parameters. By matching the lowest scaling dimensions, we numerically calculate these field-theoretic parameters and track their evolution as a function of the parameters of the lattice model.

cond-mat.str-el

Symmetry-protected topological phases and orbifolds: Generalized Laughlin's argument

We consider non-chiral symmetry-protected topological phases of matter in two spatial dimensions protected by a discrete symmetry such as $\mathbb{Z}_K$ or $\mathbb Z_K \times \mathbb Z_K $ symmetry. We argue that modular invariance/noninvariance of the partition function of the one-dimensional edge theory can be used to diagnose whether, by adding a suitable potential, the edge theory can be gapped or not without breaking the symmetry. By taking bosonic phases described by Chern-Simons K-matrix theories and fermionic phases relevant to topological superconductors as an example, we demonstrate explicitly that when the modular invariance is achieved, we can construct an interaction potential that is consistent with the symmetry and can completely gap out the edge.

cond-mat.str-el