SearcharxivSearch

arXiv subjects

Qiming Peng

Publications and source records attributed to Qiming Peng.

10 recordsLinked to original sources

Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs

Hybrid-thinking multimodal large language models (MLLMs) allow a single model to alternate between deliberative thinking and latency-efficient non-thinking inference. Although these modes differ in reasoning budget, their delivered responses should satisfy the same user-facing standard. Correctness alone may not characterize this response quality; we therefore evaluate task accuracy and response-pattern failures as complementary outcomes. We study this gap through \textbf{response-pattern alignment}: whether thinking and non-thinking interfaces preserve acceptable final-response behavior. We introduce \textbf{PatternEval}, a failure-enriched diagnostic benchmark comprising 2,415 multimodal prompts spanning visual perception and grounding, structured image understanding, and multimodal knowledge reasoning. PatternEval tests four recurrent failures: chain-of-thought leakage, response repetition, logical contradiction, and performative reasoning. Response-pattern failures are widespread across models from different providers, with non-thinking inference exhibiting substantially higher failure rates and thereby creating systematic misalignment between thinking and non-thinking interfaces. Motivated by this diagnosis, we develop \textbf{PatternRM}, a response-level reward model, and \textbf{PatternRL}, which introduces pattern-specific penalties during reinforcement learning. Experiments on Qwen3-VL-4B and Qwen3-VL-8B show that incorporating pattern-specific penalties into reinforcement learning can mitigate cross-mode misalignment while incurring a marginal task performance trade-off. Together, PatternEval and PatternRL provide an evaluation-and-training framework for aligning user-visible response patterns across hybrid-thinking interfaces.

cs.CV

Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models

Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in multimodal tasks. Despite their impressive performance, MLLMs suffer from the modality imbalance issue, where visual information is often underutilized compared to textual representations in deeper layers, leading to degraded visual performance or hallucinations. This issue stems from the predominant reliance on next-text-token-prediction during training, which fails to provide direct visual supervisory signals, resulting in progressive homogenization of visual representations throughout the layers. To this end, we propose Latent Visual Reconstruction (LaVer), a novel training framework that facilitates MLLMs in learning more discriminative visual representations via masked image modeling in the joint latent semantic space of LLM. Our method offers direct visual activation to MLLMs, which exhibit increased visual attention allocation, indicating enhanced utilization of visual information. Extensive experiments across diverse benchmarks prove the superiority of our approach in various scenarios, especially those requiring dense visual capabilities. Code of LaVer is available at https://github.com/Fir-lat/LaVer.

cs.CV

HunyuanOCR Technical Report

This paper presents HunyuanOCR, a commercial-grade, open-source, and lightweight (1B parameters) Vision-Language Model (VLM) dedicated to OCR tasks. The architecture comprises a Native Vision Transformer (ViT) and a lightweight LLM connected via an MLP adapter. HunyuanOCR demonstrates superior performance, outperforming commercial APIs, traditional pipelines, and larger models (e.g., Qwen3-VL-4B). Specifically, it surpasses current public solutions in perception tasks (Text Spotting, Parsing) and excels in semantic tasks (IE, Text Image Translation), securing first place in the ICDAR 2025 DIMT Challenge (Small Model Track). Furthermore, it achieves state-of-the-art (SOTA) results on OCRBench among VLMs with fewer than 3B parameters. HunyuanOCR achieves breakthroughs in three key aspects: 1) Unifying Versatility and Efficiency: We implement comprehensive support for core capabilities including spotting, parsing, IE, VQA, and translation within a lightweight framework. This addresses the limitations of narrow "OCR expert models" and inefficient "General VLMs". 2) Streamlined End-to-End Architecture: Adopting a pure end-to-end paradigm eliminates dependencies on pre-processing modules (e.g., layout analysis). This fundamentally resolves error propagation common in traditional pipelines and simplifies system deployment. 3) Data-Driven and RL Strategies: We confirm the critical role of high-quality data and, for the first time in the industry, demonstrate that Reinforcement Learning (RL) strategies yield significant performance gains in OCR tasks. HunyuanOCR is officially open-sourced on HuggingFace. We also provide a high-performance deployment solution based on vLLM, placing its production efficiency in the top tier. We hope this model will advance frontier research and provide a solid foundation for industrial applications.

cs.CV

ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding

Recent years have witnessed the rise and success of pre-training techniques in visually-rich document understanding. However, most existing methods lack the systematic mining and utilization of layout-centered knowledge, leading to sub-optimal performances. In this paper, we propose ERNIE-Layout, a novel document pre-training solution with layout knowledge enhancement in the whole workflow, to learn better representations that combine the features from text, layout, and image. Specifically, we first rearrange input sequences in the serialization stage, and then present a correlative pre-training task, reading order prediction, to learn the proper reading order of documents. To improve the layout awareness of the model, we integrate a spatial-aware disentangled attention into the multi-modal transformer and a replaced regions prediction task into the pre-training phase. Experimental results show that ERNIE-Layout achieves superior performance on various downstream tasks, setting new state-of-the-art on key information extraction, document image classification, and document question answering datasets. The code and models are publicly available at http://github.com/PaddlePaddle/PaddleNLP/tree/develop/model_zoo/ernie-layout.

cs.CL

ERNIE-mmLayout: Multi-grained MultiModal Transformer for Document Understanding

Recent efforts of multimodal Transformers have improved Visually Rich Document Understanding (VrDU) tasks via incorporating visual and textual information. However, existing approaches mainly focus on fine-grained elements such as words and document image patches, making it hard for them to learn from coarse-grained elements, including natural lexical units like phrases and salient visual regions like prominent image regions. In this paper, we attach more importance to coarse-grained elements containing high-density information and consistent semantics, which are valuable for document understanding. At first, a document graph is proposed to model complex relationships among multi-grained multimodal elements, in which salient visual regions are detected by a cluster-based method. Then, a multi-grained multimodal Transformer called mmLayout is proposed to incorporate coarse-grained information into existing pre-trained fine-grained multimodal Transformers based on the graph. In mmLayout, coarse-grained information is aggregated from fine-grained, and then, after further processing, is fused back into fine-grained for final prediction. Furthermore, common sense enhancement is introduced to exploit the semantic information of natural lexical units. Experimental results on four tasks, including information extraction and document question answering, show that our method can improve the performance of multimodal Transformers based on fine-grained elements and achieve better performance with fewer parameters. Qualitative analyses show that our method can capture consistent semantics in coarse-grained elements.

cs.CV

Triplet-Polaron Interaction Induced Upconversion from Triplet to Singlet: a New Way to Obtain Highly Efficient OLEDs

The triplet harvesting is a main challenge in organic light-emitting devices (OLEDs), due to the radiative decay of triplet is spin-forbidden. Here, we designed and synthesized two D-A type molecules, TPA-TAZ and TCP. The OLEDs based on them exhibit deep-blue emission and the singlet formation ratios are higher than the simple spin-statistics of 25 %. Specially, a TPA-TAZ-based OLED achieves a maximum EQE of 6.8 %, which is the largest value of the undoped OLEDs with CIE(y)< 0.06 (the EBU blue standard) up to date. Comprehensive experiments eliminate the triplet-harvesting processes of thermally activated delayed fluorescence and triplet-triplet annihilation. Instead, the triplet-polaron interaction induced upconversion from triplet to singlet through one-electron transfer mechanism is proposed, and proven by the magneto-current measurement and quantum chemistry computation. Our results may offer a new route to break through the 25 % upper limit of IQE of fluorescent OLEDs, especially, the deep-blue fluorescent OLEDs.

physics.chem-ph

Organic light-emitting diodes using open-shell molecule as emitter: the emission from doublet

We fabricate OLEDs using a stable neutral π radical, BDPA, as the emitter. There is only one electron in the singly occupied molecular orbital (SOMO) of this open-shell molecule. This feature makes the excited state of open-shell molecules be neither singlet nor triplet, but doublet. The key issue of how to harvest the triplet energy in an OLED is thus bypassed, due to the radiative decay of doublet is totally spin allowed. In the BDPA-based OLED, the emission was confirmed to be from the electronic transition from LUMO to SOMO, via the frontier molecular orbital analysis combined with the spectroscopy measurements. The maximum luminance of the OLEDs is 4879 cd/m2 which is comparable to the first reported Fluorescence-, Phosphorecence- and TADF-based OLEDs.

physics.chem-ph

Time-resolved study of the magnetic field effects on electroluminescence in tri-(8-hydroxyquinoline)- aluminum based organic light emitting devices

We investigated the magnetic field effects (MFEs) in organic light-emitting diodes (OLEDs) through the transient electroluminescence (EL) method. The time-resolved MFEs on the emission were obtained for the first time, which would be a useful method to clarify the underlying mechanisms of the MFEs. The fluorescent dye doped tri-(8-hydroxyquinoline)-aluminum (Alq3) based OLEDs were fabricated. Then, the transient EL was measured both with and without a magnetic field. To explore the time-resolved MFEs on the emission of the device, the excitons population dynamics in the device have been analyzed by a kinetic model. Our results suggest that both the intersystem crossing between the singlet and triplet electron-hole pairs and the triplet-triplet annihilation perturbed by the external magnetic field cause the time-resolved MFEs.

cond-mat.mtrl-sci

Driving conditions dependence of magneto-electroluminescence in tri-(8-hydroxyquinoline)-aluminum based organic light emitting diodes

we investigated the magneto-electroluminescence (MEL) in tri-(8-hydroxyquinoline)-aluminum based organic light-emitting diodes (OLEDs) through the steady-state and transient method simultaneously. The MELs show the great different behaviors when we turn the driving condition from a constant voltage to a pulse voltage. For devices driven by the constant voltage, the MELs are similar with the literature data; for devices driven by the pulse voltage, the MELs are quite different, they firstly increase to a maximum then decrease as the magnetic field increases continuously. Negative MELs can be seen when both the magnetic field and driving voltage are high enough.

physics.chem-ph

Investigation of the magnetic field effects on the electron mobility in tri-(8-hydroxyquinoline)-aluminum based light-emitting devices

We investigated the mganetic field effects (MFEs) on electron mobility in tri-(8-hydroxyquinoline)-aluminum based light-emitting devices by transient-electroluminescence method upon application of various offset voltages. It is found the rising edges of EL pulses are well overlapped and the falling edges of EL pulses are separated for the magnetic field on and off when Voffset=0 V and Voffset>Vturnon of the devices. The results suggest the bipolaron model and the triplet-polaron interaction model related to the carriers mobilities are not the dominant mechanisms to explain the MFEs under our experimental conditions, and the magnetic field affects the carriers recombination process is confirmed.

physics.optics