SearcharxivSearch

arXiv subjects

Shuai Peng

Publications and source records attributed to Shuai Peng.

At least 19 recordsLinked to original sources

Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context

Long-context modeling is becoming a core capability of modern large vision-language models (LVLMs), enabling sustained context management across long-document understanding, video analysis, and multi-turn tool use in agentic workflows. Yet practical training recipes remain insufficiently explored, particularly for designing and balancing long-context data mixtures. In this work, we present a systematic study of long-context continued pre-training for LVLMs, extending a 7B model from 32K to 128K context with extensive ablations on long-document data. We first show that long-document VQA is substantially more effective than OCR transcription. Building on this observation, our ablations further yield three key findings: i) for sequence-length distribution, balanced data outperforms target-length-focused data (e.g., 128K), suggesting that long-context ability requires generalizable key-information retrieval across various lengths and positions; ii) retrieval remains the primary bottleneck, favoring retrieval-heavy mixtures with modest reasoning data for task diversity; and iii) pure long-document VQA largely preserves short-context capabilities, suggesting that instruction-formatted long data reduces the need for short-data mixing. Based on these findings, we introduce MMProLong, obtained by long-context continued pre-training from Qwen2.5-VL-7B with only a 5B-token budget. MMProLong improves long-document VQA scores by 7.1% and maintains strong performance at 256K and 512K contexts beyond its 128K training window, without additional training. It further generalizes to webpage-based multimodal needle retrieval, long-context vision-text compression, and long-video understanding without task-specific supervision. Overall, our study establishes a practical LongPT recipe and an empirical foundation for advancing long-context vision-language models.

cs.CV

Orbital-Dependent Dimensional Crossover of a $p$-Wave Feshbach Resonance

We report the observation of a dimensional crossover of a $p$-wave Feshbach resonance in an ultracold, spin-polarized $^6$Li Fermi gas confined by a one-dimensional optical lattice. Using high-resolution atom-loss spectroscopy, we resolve the orbital doublet associated with the $\ml=0$ and $|\ml|=1$ scattering channels over a wide range of lattice depths. In the weak-confinement regime, the atom loss signal associated with the $|\ml|=1$ branch is stronger, consistent with the twofold orbital degeneracy of the three-dimensional system. As the lattice confinement increases, the relative loss weight of the two orbital branches evolves continuously toward the quasi-two-dimensional limit, indicating a progressive suppression of relative motion along the lattice direction. In addition, we observe a systematic confinement dependence of the orbital splitting between the two resonance branches. These results provide an experimental characterization of orbital-dependent $p$-wave scattering in reduced dimensions and motivate future microscopic studies of confined anisotropic scattering.

cond-mat.quant-gas

Accelerating evaporative cooling of a strongly interacting Fermi gas by tilting the optical trap with a magnetic field gradient

We present a rapid evaporative cooling scheme for a strongly interacting $^{6}\mathrm{Li}$ Fermi gas in an optical dipole trap. The method uses a magnetic-field-gradient--induced tilt of the trapping potential to accelerate cooling in the unitarity-limited regime. In evaporation based only on lowering the optical trap depth, the unitarity-limited scattering cross section can support runaway cooling; however, the cooling rate slows around $T/T_F \simeq 0.5$, and the runaway behavior is no longer maintained. We improve on this approach by applying a magnetic-field gradient when the gas temperature reaches about half the Fermi temperature. The induced tilt opens an escape channel for energetic atoms while keeping the trap frequencies nearly unchanged. This modification increases the cooling speed and cools the gas below the superfluid transition temperature, reaching $T/T_F = 0.16$ on a timescale of $\sim 25\,\mathrm{ms}$. Our results provide a simple and robust route for rapidly cooling a strongly interacting Fermi gas into the superfluid regime, facilitating studies of the physics of unitary Fermi superfluids.

cond-mat.quant-gas

Orbital-resolved three-body recombination across a p-wave Feshbach resonance in ultracold $^6$Li

We report precision, orbital-resolved measurements of three-body recombination near the 159~G $p$-wave Feshbach resonance in an ultracold gas of $^{6}$Li atoms prepared in their lowest hyperfine state. Using a radio-frequency gated protocol that suppresses magnetic-field transients below the milligauss level, we resolve loss features associated with the $|m_\ell|=1$ and $m_\ell=0$ orbital projections. The measured three-body loss coefficient $L_3$ is well captured by a thermally averaged cascade-recombination model, enabling extraction of the resonance splitting $\delta B$ and effective-range parameter $k_e$. At the lowest temperature, we obtain $\delta B = 7.6(3)$~mG and $k_e = 0.151(6)\,a_0^{-1}$, both in quantitative agreement with coupled-channel theory. These results establish orbital-resolved three-body spectroscopy as a precision probe of $p$-wave scattering and provide a benchmark for microscopic models of resonant few-body loss.

cond-mat.quant-gas

Clair Obscur: an Illumination-Aware Method for Real-World Image Vectorization

Image vectorization aims to convert raster images into editable, scalable vector representations while preserving visual fidelity. Existing vectorization methods struggle to represent complex real-world images, often producing fragmented shapes at the cost of semantic conciseness. In this paper, we propose COVec, an illumination-aware vectorization framework inspired by the Clair-Obscur principle of light-shade contrast. COVec is the first to introduce intrinsic image decomposition in the vector domain, separating an image into albedo, shade, and light layers in a unified vector representation. A semantic-guided initialization and two-stage optimization refine these layers with differentiable rendering. Experiments on various datasets demonstrate that COVec achieves higher visual fidelity and significantly improved editability compared to existing methods. The code will be released at https://github.com/decade-de/COVec.

cs.CV

Uni-MuMER: Unified Multi-Task Fine-Tuning of Vision-Language Model for Handwritten Mathematical Expression Recognition

Handwritten Mathematical Expression Recognition (HMER) remains a persistent challenge in Optical Character Recognition (OCR) due to the inherent freedom of symbol layouts and variability in handwriting styles. Prior methods have faced performance bottlenecks by proposing isolated architectural modifications, making them difficult to integrate coherently into a unified framework. Meanwhile, recent advances in pretrained vision-language models (VLMs) have demonstrated strong cross-task generalization, offering a promising foundation for developing unified solutions. In this paper, we introduce Uni-MuMER, which fully fine-tunes a VLM for the HMER task without modifying its architecture, effectively injecting domain-specific knowledge into a generalist framework. Our method integrates three data-driven tasks: Tree-Aware Chain-of-Thought (Tree-CoT) for structured spatial reasoning, Error-Driven Learning (EDL) for reducing confusion among visually similar characters, and Symbol Counting (SC) for improving recognition consistency in long expressions. Experiments on the CROHME and HME100K datasets show that Uni-MuMER achieves super state-of-the-art performance, outperforming the best lightweight specialized model SSAN by 16.31\% and the top-performing VLM Gemini2.5-flash by 24.42\% under zero-shot setting. Our datasets, models, and code are open-sourced at: {https://github.com/BFlameSwift/Uni-MuMER

cs.CV

Precision Measurement of Spin-Dependent Dipolar Splitting in $^6$Li p-Wave Feshbach Resonances

The magnetic dipolar splitting of a p-wave Feshbach resonance is governed by the spin-orbital configuration of the valence electrons in the triplet molecular state. We perform high-resolution trap loss spectroscopy on ultracold 6Li atoms to resolve this splitting with sub-milligauss precision. By comparing spin-polarized (|mS| = 1) and spin-mixture (mS = 0) configurations of the triplet state, we observe a clear spin-dependent reversal in the splitting structure, confirmed via momentumresolved absorption imaging. This behavior directly reflects the interplay between electron spin projection mS and orbital angular momentum ml in the molecular states. Our results provide a stringent benchmark for dipole-dipole interaction models and lay the groundwork for controlling the pairing in p-wave superfluid systems.

cond-mat.quant-gas

Seed1.5-VL Technical Report

We present Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning. Seed1.5-VL is composed with a 532M-parameter vision encoder and a Mixture-of-Experts (MoE) LLM of 20B active parameters. Despite its relatively compact architecture, it delivers strong performance across a wide spectrum of public VLM benchmarks and internal evaluation suites, achieving the state-of-the-art performance on 38 out of 60 public benchmarks. Moreover, in agent-centric tasks such as GUI control and gameplay, Seed1.5-VL outperforms leading multimodal systems, including OpenAI CUA and Claude 3.7. Beyond visual and video understanding, it also demonstrates strong reasoning abilities, making it particularly effective for multimodal reasoning challenges such as visual puzzles. We believe these capabilities will empower broader applications across diverse tasks. In this report, we mainly provide a comprehensive review of our experiences in building Seed1.5-VL across model design, data construction, and training at various stages, hoping that this report can inspire further research. Seed1.5-VL is now accessible at https://www.volcengine.com/ (Volcano Engine Model ID: doubao-1-5-thinking-vision-pro-250428)

cs.CV

LogicPro: Improving Complex Logical Reasoning via Program-Guided Learning

In this paper, we propose a new data synthesis method called \textbf{LogicPro}, which leverages LeetCode-style algorithm \underline{Pro}blems and their corresponding \underline{Pro}gram solutions to synthesize Complex \underline{Logic}al Reasoning data in text format. First, we synthesize complex reasoning problems through source algorithm problems and test cases. Then, standard answers and intermediate variable outputs are obtained for each problem based on standard python solutions and test cases. Finally, with the guidance of code intermediate variables, we synthesize the text reasoning process for each reasoning problems. Through this method, we can synthesize data that is difficult, scalable, effective, and comes with golden standard answers and high-quality reasoning processes. As a result, with our 540K synthesized dataset constructed solely from 2,360 algorithm problems, our approach \footnote{Code and data are publicly available at https://github.com/jiangjin1999/LogicPro} achieves significant improvements in multiple models for the datasets \textit{BBH$^{27}$}, \textit{LogicBench}, \textit{DROP}, \textit{AR-LSAT}, and \textit{GSM8K}, etc. outperforming a wide range of existing reasoning datasets.

cs.CL

Vote&Mix: Plug-and-Play Token Reduction for Efficient Vision Transformer

Despite the remarkable success of Vision Transformers (ViTs) in various visual tasks, they are often hindered by substantial computational cost. In this work, we introduce Vote\&Mix (\textbf{VoMix}), a plug-and-play and parameter-free token reduction method, which can be readily applied to off-the-shelf ViT models \textit{without any training}. VoMix tackles the computational redundancy of ViTs by identifying tokens with high homogeneity through a layer-wise token similarity voting mechanism. Subsequently, the selected tokens are mixed into the retained set, thereby preserving visual information. Experiments demonstrate VoMix significantly improves the speed-accuracy tradeoff of ViTs on both images and videos. Without any training, VoMix achieves a 2$\times$ increase in throughput of existing ViT-H on ImageNet-1K and a 2.4$\times$ increase in throughput of existing ViT-L on Kinetics-400 video dataset, with a mere 0.3\% drop in top-1 accuracy.

cs.CV

MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models

The rapid development of large language models (LLMs) has spurred extensive research into their domain-specific capabilities, particularly mathematical reasoning. However, most open-source LLMs focus solely on mathematical reasoning, neglecting the integration with visual injection, despite the fact that many mathematical tasks rely on visual inputs such as geometric diagrams, charts, and function plots. To fill this gap, we introduce \textbf{MultiMath-7B}, a multimodal large language model that bridges the gap between math and vision. \textbf{MultiMath-7B} is trained through a four-stage process, focusing on vision-language alignment, visual and math instruction-tuning, and process-supervised reinforcement learning. We also construct a novel, diverse and comprehensive multimodal mathematical dataset, \textbf{MultiMath-300K}, which spans K-12 levels with image captions and step-wise solutions. MultiMath-7B achieves state-of-the-art (SOTA) performance among open-source models on existing multimodal mathematical benchmarks and also excels on text-only mathematical benchmarks. Our model and dataset are available at {\textcolor{blue}{\url{https://github.com/pengshuai-rin/MultiMath}}}.

cs.CL

SketchRef: a Multi-Task Evaluation Benchmark for Sketch Synthesis

Sketching is a powerful artistic technique for capturing essential visual information about real-world objects and has increasingly attracted attention in image synthesis research. However, the field lacks a unified benchmark to evaluate the performance of various synthesis methods. To address this, we propose SketchRef, the first comprehensive multi-task evaluation benchmark for sketch synthesis. SketchRef fully leverages the shared characteristics between sketches and reference photos. It introduces two primary tasks: category prediction and structural consistency estimation, the latter being largely overlooked in previous studies. These tasks are further divided into five sub-tasks across four domains: animals, common things, human body, and faces. Recognizing the inherent trade-off between recognizability and simplicity in sketches, we are the first to quantify this balance by introducing a recognizability calculation method constrained by simplicity, mRS, ensuring fair and meaningful evaluations. To validate our approach, we collected 7,920 responses from art enthusiasts, confirming the effectiveness of our proposed evaluation metrics. Additionally, we evaluate the performance of existing sketch synthesis methods on our benchmark, highlighting their strengths and weaknesses. We hope this study establishes a standardized benchmark and offers valuable insights for advancing sketch synthesis algorithms.

cs.CV

Observation of a broad state-to-state spin-exchange collision near a p-wave Feshbach resonances of $^6$Li atoms

The study of state-to-state spin-exchange collisions in the vicinity of $p$-wave Feshbach resonances offer great opportunities to explore many-body interactions and novel quantum phases. Here, we report the observation of a spin-exchange collision near a $p$-wave Feshbach resonance within a mixture of the lowest and third-lowest hyperfine states of $^6$Li atoms. The spin-exchange interaction is observed over a range of ten gausses and produces a pair of atoms in the second-lowest hyperfine states that are captured by a deep optical dipole trap. We apply a coupled-channel method to calculate the scattering properties of this system. We find that the $p$-wave resonance exhibits a low inelastic collision rate and a broad resonance profile, which is due to the modification by the accompanying spin-exchange collisions. These findings open up new possibilities for the creation of long-lived, strongly interacting $p$-wave Fermi gases.

cond-mat.quant-gas

Scaling law for three-body collisions near a narrow s-wave Feshbach resonance

Ultracold atomic gases provide a controllable system to study the inelastic processes for three-body systems, where the three-body recombination rate depends on the scattering length scaling. Such scalings have been confirmed in bosonic systems with various interaction strengths, but their existence with fermionic atoms remains elusive. In this work, we report on an experimental investigation of the scaling law for the three-body atomic loss rate $L_3$ in a two-component $^6$Li Fermi gas with the scattering length $a<0$. The scaling law is validated within a certain range of $a$ near the narrow $s$-wave Feshbach resonance, where $L_3\propto T|a|^{2.60(5)}$, and $T$ is the gas temperature. The scaling law is observed to have an upper and a lower bound in terms of the scattering length. For the upper bound, when $a\rightarrow \infty$, the power-law scaling is suppressed by the unitary behavior of the resonance caused by the strong three-body collisions. For the lower bound, $a\rightarrow 0$, the finite range effect modifies the scaling law by the effective scattering length $L_e$. These results indicate that the three-body recombination rate in a fermionic system could be characterized by the scaling law associated with the generalized Efimov physics.

cond-mat.quant-gas

Controllable Production of Degenerate Fermi Gases of $^6$Li Atoms in the 2D-3D Crossover

The many-body physics in the dimensional crossover regime attracts much attention in cold atom experiments, but yet to explore systematically. One of the technical difficulties existed in the experiments is the lack of the experimental technique to quantitatively tune the atom occupation ratio of the different lattice bands. In this letter, we report such techniques in a process of transferring a 3D Fermi gas into a 1D optical lattice, where the capability of tuning the occupation of the energy band is realized by varying the trapping potentials of the optical dipole trap (ODT) and the lattice, respectively. We could tune a Fermi gas with the occupation in the lowest band from unity to 50$\%$ quantitatively. This provides a route to experimentally study the dependence of many-body interaction on the dimensionality in a Fermi gas.

cond-mat.quant-gas

Cooling a Fermi gas with three-body recombination near a narrow Feshbach resonance

Three-body recombination is a phenomenon common in atomic and molecular collisions, producing heating in the system. However, we find the cooling effect of the three-body recombination of a 6Li Fermi gas near its s-wave narrow Feshbach resonance. Such counter-intuitive behavior is explained as follows, the threshold energy of the quasi-bounded Feshbach molecule acts as the knife of cooling, expelling the scattering atoms with selected kinetic energy from the trap. When the threshold energy happens to be larger than 3/2kBT, each lost atom in the three-body recombination process has more than 3kBT energy which results in cooling. The best cooling is found with the threshold energy set at about 3kBT, consistent with a theoretical model. The three-body recombination induced cooling raises potential applications for cooling complex atomic systems.

cond-mat.quant-gas

The first OGLE-discovered ultracompact X-ray binary is an intermediate polar

The variable source OGLE-UCXB-01 is the first OGLE-discovered ultracompact X-ray binary (UCXB). The 12-year long-term OGLE optical photometry of this source shows a period of P= 12.8 min and a fast period decreasing rate Pdot= -9.2E-11 s s^-1. At a luminosity of L_X ~ 4E33 erg s^-1, its X-ray emission is also variable and correlated with the optical variability. To determine the nature of this variable source, specifically the masses and types of its binary components, we consider first an attractive possibility that the optical variation is due to the secondary's ellipsoidal variation and a strong gravitational wave emission drives the orbital decay. However, we can not find an allowable solution to the secondary that satisfies simultaneously the three constraints: an ultra-tight orbit, the bright absolute magnitude, and the large amplitude of the brightness variation. Moreover, the inferred mass transfer rate is too high. This scenario is therefore ruled out. We then find the system is fully consistent with an "intermediate polar" model, in which the optical and X-ray emission comes from a magnetized white dwarf (WD) accreting from a low-mass (<~ 0.7 M_sun) main-sequence secondary. The observed period decay is the accretion-driven spin-up of the WD. The WD spin period is 12.8 min and the orbital period is shorter than 10 hr. The method presented here can be applied to other UCXB candidates or impostors with time-domain data available only.

astro-ph.HE

Handwritten Mathematical Expression Recognition with Bidirectionally Trained Transformer

Encoder-decoder models have made great progress on handwritten mathematical expression recognition recently. However, it is still a challenge for existing methods to assign attention to image features accurately. Moreover, those encoder-decoder models usually adopt RNN-based models in their decoder part, which makes them inefficient in processing long $\LaTeX{}$ sequences. In this paper, a transformer-based decoder is employed to replace RNN-based ones, which makes the whole model architecture very concise. Furthermore, a novel training strategy is introduced to fully exploit the potential of the transformer in bidirectional language modeling. Compared to several methods that do not use data augmentation, experiments demonstrate that our model improves the ExpRate of current state-of-the-art methods on CROHME 2014 by 2.23%. Similarly, on CROHME 2016 and CROHME 2019, we improve the ExpRate by 1.92% and 2.28% respectively.

cs.CV