SearcharxivSearch

arXiv subjects

Shasha Wang

Publications and source records attributed to Shasha Wang.

12 recordsLinked to original sources

MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale

Current document parsing methods advance primarily through model architecture innovation, while systematic engineering of training data remains underexplored. Yet state-of-the-art models spanning diverse architectures and parameter scales exhibit highly consistent failure patterns on the same set of hard samples, suggesting that the performance bottleneck stems from shared deficiencies in training data rather than from architectural differences. Building on this finding, we present MinerU2.5-Pro, which advances the state of the art purely through data engineering and training strategy design while retaining the 1.2B-parameter architecture of MinerU2.5 unchanged. At its core is a Data Engine co-designed around coverage, informativeness, and annotation accuracy: Diversity-and-Difficulty-Aware Sampling expands training data from under 10M to 65.5M samples while mitigating distribution shift; Cross-Model Consistency Verification leverages output consensus among heterogeneous models to assess sample difficulty and generate reliable annotations; the Judge-and-Refine pipeline improves annotation quality for hard samples through render-then-verify iterative correction. A three-stage progressive training strategy--large-scale pre-training, hard sample fine-tuning, and GRPO alignment--sequentially exploits these data at different quality tiers. On the evaluation front, we rectify element-matching biases in OmniDocBench v1.5 and introduce a Hard subset, establishing the more discriminative OmniDocBench v1.6 protocol. Without any architectural modification, MinerU2.5-Pro achieves 95.69 on OmniDocBench v1.6, improving over the same-architecture baseline by 2.71 points and surpassing all existing methods, including those based on models with over 200x more parameters.

cs.CV

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

We introduce MinerU2.5, a 1.2B-parameter document parsing vision-language model that achieves state-of-the-art recognition accuracy while maintaining exceptional computational efficiency. Our approach employs a coarse-to-fine, two-stage parsing strategy that decouples global layout analysis from local content recognition. In the first stage, the model performs efficient layout analysis on downsampled images to identify structural elements, circumventing the computational overhead of processing high-resolution inputs. In the second stage, guided by the global layout, it performs targeted content recognition on native-resolution crops extracted from the original image, preserving fine-grained details in dense text, complex formulas, and tables. To support this strategy, we developed a comprehensive data engine that generates diverse, large-scale training corpora for both pretraining and fine-tuning. Ultimately, MinerU2.5 demonstrates strong document parsing ability, achieving state-of-the-art performance on multiple benchmarks, surpassing both general-purpose and domain-specific models across various recognition tasks, while maintaining significantly lower computational overhead.

cs.CV

OpenHuEval: Evaluating Large Language Model on Hungarian Specifics

We introduce OpenHuEval, the first benchmark for LLMs focusing on the Hungarian language and specifics. OpenHuEval is constructed from a vast collection of Hungarian-specific materials sourced from multiple origins. In the construction, we incorporated the latest design principles for evaluating LLMs, such as using real user queries from the internet, emphasizing the assessment of LLMs' generative capabilities, and employing LLM-as-judge to enhance the multidimensionality and accuracy of evaluations. Ultimately, OpenHuEval encompasses eight Hungarian-specific dimensions, featuring five tasks and 3953 questions. Consequently, OpenHuEval provides the comprehensive, in-depth, and scientifically accurate assessment of LLM performance in the context of the Hungarian language and its specifics. We evaluated current mainstream LLMs, including both traditional LLMs and recently developed Large Reasoning Models. The results demonstrate the significant necessity for evaluation and model optimization tailored to the Hungarian language and specifics. We also established the framework for analyzing the thinking processes of LRMs with OpenHuEval, revealing intrinsic patterns and mechanisms of these models in non-English languages, with Hungarian serving as a representative example. We will release OpenHuEval at https://github.com/opendatalab/OpenHuEval .

cs.CL

Benchmarking and Boosting Multilingual Capabilities of LVLMs via OCR-Centric Reinforcement Learning

Evaluating the multilingual capabilities of Large Vision-Language Models (LVLMs) remains challenging because most benchmarks rely on non-parallel corpora, making it unclear whether cross-lingual performance gaps reflect model limitations or dataset inconsistencies. To address this, we introduce PM4Bench, the first multimodal, multilingual, multi-task benchmark built on a strictly parallel 10-language corpus, enabling fair, apples-to-apples cross-lingual comparison of model performance. We further introduce a vision setting that embeds textual inputs directly into images, better approximating deployment scenarios where LVLM-driven agents interact with virtual or physical environments through unified visual observations. Experiments with 10 LVLMs reveal that OCR is a key factor behind cross-lingual disparity when textual content is rendered visually. Motivated by this, we design an OCR-centric GRPO training strategy using fully synthesized, label-free OCR data, without expensive task-specific VQA supervision. The resulting model improves general multilingual VQA capability, reduces cross-lingual disparities under the vision setting, and transfers gains beyond PM4Bench. This methodology offers an efficient, label-free pathway toward more equitable multilingual deployment of LVLM-driven agents.

cs.CV

B\"acklund-Darboux Transformations for Super KdV Type Equations

By introducing a Miura transformation, we derive a generalized super modified Korteweg-de Vries (gsmKdV) equation from the generalized super KdV (gsKdV) equation. It is demonstrated that, while the gsKdV equation takes Kupershmidt's super KdV (sKdV) equation and Geng-Wu's sKdV equation as two distinct reductions, there are also two equations, namely Kupershmidt's super modified KdV (smKdV) equation and Hu's smKdV equation, which are associated with the gsmKdV equation. By analyzing the flows within the gsKdV and gsmKdV hierarchies, we specifically derive the first negative flows associated with both hierarchies.We then construct a number of B\"acklund-Darboux transformations (BDTs) for both the gsKdV and gsmKdV equations, elucidating the interrelationship between them. By proper reductions, we are able not only to recover the previously known BDTs for Kupershimdt's sKdV and smKdV equations, but also to obtain the BDTs for the Geng-Wu's sKdV/smKdV and Hu's smKdV equations. As applications, we construct some exact solutions for those equations. Since all flows of the sKdV or smKdV hierarchy share the same spatial parts of spectral problem, thus these Darboux matrices and spatial parts of BTs are applicable to any flow of those hierarchies.

nlin.SI

Beyond Pixel-Wise Supervision for Medical Image Segmentation: From Traditional Models to Foundation Models

Medical image segmentation plays an important role in many image-guided clinical approaches. However, existing segmentation algorithms mostly rely on the availability of fully annotated images with pixel-wise annotations for training, which can be both labor-intensive and expertise-demanding, especially in the medical imaging domain where only experts can provide reliable and accurate annotations. To alleviate this challenge, there has been a growing focus on developing segmentation methods that can train deep models with weak annotations, such as image-level, bounding boxes, scribbles, and points. The emergence of vision foundation models, notably the Segment Anything Model (SAM), has introduced innovative capabilities for segmentation tasks using weak annotations for promptable segmentation enabled by large-scale pre-training. Adopting foundation models together with traditional learning methods has increasingly gained recent interest research community and shown potential for real-world applications. In this paper, we present a comprehensive survey of recent progress on annotation-efficient learning for medical image segmentation utilizing weak annotations before and in the era of foundation models. Furthermore, we analyze and discuss several challenges of existing approaches, which we believe will provide valuable guidance for shaping the trajectory of foundational models to further advance the field of medical image segmentation.

cs.CV

Achieving High Yield of Perpendicular SOT-MTJ Manufactured on 300 mm Wafers

The large-scale fabrication of three-terminal magnetic tunnel junctions (MTJs) with high yield is becoming increasingly crucial, especially with the growing interest in spin-orbit torque (SOT) magnetic random access memory (MRAM) as the next generation of MRAM technology. To achieve high yield and consistent device performance in MTJs with perpendicular magnetic anisotropy, an integration flow has been developed that incorporates special MTJ etching technique and other CMOS-compatible processes on a 300 mm wafer manufacturing platform. Systematic studies have been conducted on device performance and statistical uniformity, encompassing magnetic properties, electrical switching behavior, and reliability. Achievements include a switching current of 680 uA at 2 ns, a TMR as high as 119%, ultra-high endurance (over 1012 cycles), and excellent uniformity in the fabricated SOT-MTJ devices, with a yield of up to 99.6%. The proposed integration process, featuring high yield, is anticipated to streamline the mass production of SOT-MRAM.

physics.app-ph

Competitive Ensembling Teacher-Student Framework for Semi-Supervised Left Atrium MRI Segmentation

Semi-supervised learning has greatly advanced medical image segmentation since it effectively alleviates the need of acquiring abundant annotations from experts and utilizes unlabeled data which is much easier to acquire. Among existing perturbed consistency learning methods, mean-teacher model serves as a standard baseline for semi-supervised medical image segmentation. In this paper, we present a simple yet efficient competitive ensembling teacher student framework for semi-supervised for left atrium segmentation from 3D MR images, in which two student models with different task-level disturbances are introduced to learn mutually, while a competitive ensembling strategy is performed to ensemble more reliable information to teacher model. Different from the one-way transfer between teacher and student models, our framework facilitates the collaborative learning procedure of different student models with the guidance of teacher model and motivates different training networks for a competitive learning and ensembling procedure to achieve better performance. We evaluate our proposed method on the public Left Atrium (LA) dataset and it obtains impressive performance gains by exploiting the unlabeled data effectively and outperforms several existing semi-supervised methods.

cs.CV

Direct imaging of a zero-field target skyrmion and its polarity switch in a chiral magnetic nanodisk

A target skyrmion is a flux-closed spin texture that has two-fold degeneracy and is promising as a binary state in next generation universal memories. Although its formation in nanopatterned chiral magnets has been predicted, its observation has remained challenging. Here, we use off-axis electron holography to record images of target skyrmions in a 160-nm-diameter nanodisk of the chiral magnet FeGe. We compare experimental measurements with numerical simulations, demonstrate switching between two stable degenerate target skyrmion ground states that have opposite polarities and rotation senses and discuss the observed switching mechanism.

cond-mat.mes-hall

Experimental observation of magnetic bobbers for a new concept of magnetic solid-state memory

The use of chiral skyrmions, which are nanoscale vortex-like spin textures, as movable data bit carriers forms the basis of a recently proposed concept for magnetic solid-state memory. In this concept, skyrmions are considered to be unique localized spin textures, which are used to encode data through the quantization of different distances between identical skyrmions on a guiding nanostripe. However, the conservation of distances between highly mobile and interacting skyrmions is difficult to implement in practice. Here, we report the direct observation of another type of theoretically-predicted localized magnetic state, which is referred to as a chiral bobber (ChB), using quantitative off-axis electron holography. We show that ChBs can coexist together with skyrmions. Our results suggest a novel approach for data encoding, whereby a stream of binary data representing a sequence of ones and zeros can be encoded via a sequence of skyrmions and bobbers. The need to maintain defined distances between data bit carriers is then not required. The proposed concept of data encoding promises to expedite the realization of a new generation of magnetic solid-state memory.

cond-mat.str-el

The effect of social welfare system based on the complex network

With the passage of time, the development of communication technology and transportation broke the isolation among people. Relationship tends to be complicated, pluralism, dynamism. In the network where interpersonal relationship and evolved complex net based on game theory work serve respectively as foundation architecture and theoretical model, with the combination of game theory and regard public welfare as influencing factor, we artificially initialize that closed network system. Through continual loop operation of the program, we summarize the changing rule of the cooperative behavior in the interpersonal relationship, so that we can analyze the policies about welfare system about whole network and the relationship of frequency of betrayal in cooperative behavior. Most analytical data come from some simple investigations and some estimates based on internet and environment and the study put emphasis on simulating social network and analyze influence of social welfare system on Cooperative Behavio.

cs.SI