SearcharxivSearch

arXiv subjects

Dongsheng Xu

Publications and source records attributed to Dongsheng Xu.

5 recordsLinked to original sources

BlueLM-2.5-3B Technical Report

We present BlueLM-2.5-3B, a compact and unified dense Multimodal Large Language Model (MLLM) designed for efficient edge-device deployment, offering strong general-purpose and reasoning capabilities. To the best of our knowledge, this is the first 3B-scale MLLM to support both thinking and non-thinking modes, while also enabling explicit control over thinking token budget. BlueLM-2.5-3B is developed through diversified data curation, key data resampling, hybrid heterogeneous reinforcement learning, and a high-performance training infrastructure. Our model achieves superior multimodal capacity while preserving competitive pure-text performance with only 2.9 billion parameters. We conduct comprehensive evaluations across a broad range of multimodal and text-only benchmarks. In thinking mode, BlueLM-2.5-3B achieves comparable performance to Qwen3-4B on text-only benchmarks, and trails the larger Kimi-VL-A3B-16B by only about 5% on average across multimodal evaluations. In non-thinking mode, it outperforms Qwen2.5-VL-3B on the majority of multimodal benchmarks. Additionally, BlueLM-2.5-3B exhibits exceptional data efficiency. All of the aforementioned performance is achieved with substantially less total training data than Qwen2.5-VL-3B and Qwen3-4B. We hope our work contributes to the advancement of high-performance, on-device MLLMs and provides meaningful insights to the research community.

cs.AI

DEVICE: Depth and Visual Concepts Aware Transformer for OCR-based Image Captioning

OCR-based image captioning is an important but under-explored task, aiming to generate descriptions containing visual objects and scene text. Recent studies have made encouraging progress, but they are still suffering from a lack of overall understanding of scenes and generating inaccurate captions. One possible reason is that current studies mainly focus on constructing the plane-level geometric relationship of scene text without depth information. This leads to insufficient scene text relational reasoning so that models may describe scene text inaccurately. The other possible reason is that existing methods fail to generate fine-grained descriptions of some visual objects. In addition, they may ignore essential visual objects, leading to the scene text belonging to these ignored objects not being utilized. To address the above issues, we propose a Depth and Visual Concepts Aware Transformer (DEVICE) for OCR-based image captinong. Concretely, to construct three-dimensional geometric relations, we introduce depth information and propose a depth-enhanced feature updating module to ameliorate OCR token features. To generate more precise and comprehensive captions, we introduce semantic features of detected visual concepts as auxiliary information, and propose a semantic-guided alignment module to improve the model's ability to utilize visual concepts. Our DEVICE is capable of comprehending scenes more comprehensively and boosting the accuracy of described visual entities. Sufficient experiments demonstrate the effectiveness of our proposed DEVICE, which outperforms state-of-the-art models on the TextCaps test set.

cs.CV

Locking ssDNA in a Graphene-Terraces Nanopore and Steering Its Step-by-Step Transportation via Electric Trigger

This study demonstrates that the nanopore terraces constructed on a multilayer graphene sheet could be employed to con-trol the conformation and transportation of an ssDNA for nanopore sequencing. As adsorbed on a terraced graphene na-nopore, the ssDNA has no in-plane swing nearby the nanopore, and can be locked on graphene terraces in a stretched con-formation. Under biasing, the accumulated ions near the nanopore promote the translocation of the locked ssDNA, and also disturb the balance between the driven force and resistance force acted on the nucleotide in pore. A critical force is found to be necessary in trigging the kickoff of the ssDNA translocation, implying an inherent field effect of the terraced graphene nanopore. By changing the intensities of electric field as trigger signal, the stop and go of an ssDNA in the nanopore are manipulated at single nucleobase level. The velocity of ssDNA in the nanopore can also be regulated by the frequency of the electro-stimulations. As a result, a new scheme of controllable translocation of ssDNA in graphene nanopores is realized by introducing controllers and triggers, appealing more explorations in experiment.

physics.bio-ph

Interlayer Water Regulates the Bio-nano Interface of a \b{eta}-sheet Protein stacking on Graphene

Using molecular dynamics simulations, we investigated an integrated bio-nano interface consisting of a \b{eta}-sheet protein stacked onto graphene. We found that the stacking assembly of the model protein on graphene could be controlled by water molecules. The interlayer water filled within interstices of the bio-nano interface could suppress the molecular vibration of surface groups on protein, and could impair the CH...π interaction driving the attraction of the protein and graphene. The intermolecular coupling of interlayer water would be relaxed by the relative motion of protein upon graphene due to the interaction between water and protein surface. This effect reduced the hindrance of the interlayer water against the assembly of protein on graphene, resulting an appropriate adsorption status of protein on graphene with a deep free energy trap. Thereby, the confinement and the relative sliding between protein and graphene, the coupling of protein and water, and the interaction between graphene and water all have involved in the modulation of behaviors of water molecules within the bio-nano interface, governing the hindrance of interlayer water against the protein assembly on hydrophobic graphene. These results provide a deep insight into the fundamental mechanism of protein adsorption onto graphene surface in water.

physics.bio-ph

Near Preservation of Quadratic Invariants by Stochastic Runge-Kutta Methods

Based on the combinatory theory of rooted colored trees, we investigate the conditions for the explicit stochastic Runge-Kutta (SRK) methods to preserve quadratic invariants (QI) up to certain orders of accuracy. These conditions can supply a practical approach of constructing explicit nearly conservative SRK methods. Meanwhile, we estimate errors in the preservation of QI resulting from iterative implementation of implicit conservative SRK methods with fixed-point and Newton's iterations. Finally, numerical experiments are performed to test the behavior of the methods in preserving QI.

math.NA