SearcharxivSearch

arXiv subjects

David Setiawan

Publications and source records attributed to David Setiawan.

2 recordsLinked to original sources

MUDIDI: A Two-Stage Framework for Multilingual Dictionary Digitization with Language Models

Multilingual dictionaries are among the most valuable documentary resources for low-resource and endangered languages, yet many remain available only as scans. For many decades, their digitization and conversion into a machine-readable format was nearly impossible due to language-specific scripts, complex multi-column layouts full of entries with abbreviations and cross-references. Recent vision-language models offer a promising solution, but it is unclear how well they preserve characters, markup, and process lexicographic structure. We introduce MUDIDI, a two-stage framework for multi-lingual dictionary digitization. Stage One evaluates the quality of character recognition and markup preservation; Stage Two focuses on dictionary entry segmentation with subsequent mapping into a machine-readable lexicographic schema, SIL's Multi-Dictionary Formatter. We also release a dataset that consists of human-annotated lexicographic entries collected from 30 public-domain dictionaries featuring diverse writing systems, language families, and formats. We benchmark OCR systems, general-purpose Large Language Models (LLMs), and Vision Language Models (VLMs) on the dataset, demonstrating superior performance of LLMs across most writing systems and languages in both stages, and provide practical guidelines on improving the results for more challenging scenarios. Finally, we show that supplementing additional information, such as dictionary introduction, to the LLMs can improve the quality of the digitized dictionary. Github: https://github.com/naarm-nlp/mudidi

cs.CL

Enhancing Science Literacy through Cognitive Conflict-Based Generative Learning Model: An Experimental Study in Physics Learning

This experimental study investigates the effectiveness of the Cognitive Conflict-Based Generative Learning Model (GLBCC) in enhancing science literacy among high school physics students. The novelty of this research lies in the innovative integration of cognitive conflict strategies with generative learning principles through a six stage structured framework, specifically designed to address persistent misconceptions in physics education while systematically developing scientific literacy competencies. The research employed pretest-posttest control group design involving 167 Grade XI students from three schools. Students were randomly assigned to experimental groups (n = 83) that received GLBCC instruction and control groups (n = 84) that used the expository learning model. Science literacy was measured using validated instruments assessing scientific knowledge, inquiry processes, and application skills across six key indicators. Statistical analysis using ANOVA with Tukey HSD post-hoc tests revealed significant improvements in science literacy scores for students receiving GLBCC instruction compared to traditional methods (p < 0.001). This study makes a unique contribution to physics education by demonstrating how the deliberate creation of cognitive conflict, combined with authentic real-world physics phenomena, can effectively restructure students conceptual understanding and enhance their scientific thinking capabilities. Factor analysis identified four critical implementation factors: science literacy development components, learning stages and orientation, and objectives, and knowledge construction processes. The findings provide empirical evidence supporting the integration of cognitive conflict strategies with generative learning approaches in physics education, offering practical implications

physics.ed-ph