arXiv · 2505.11177
Low-Resource Language Processing: An OCR-Driven Summarization and Translation Pipeline
Abstract
This paper presents an end-to-end suite for multilingual information extraction and processing from image-based documents. The system uses Optical Character Recognition (Tesseract) to extract text in languages such as English, Hindi, and Tamil, and then a pipeline involving large language model APIs (Gemini) for cross-lingual translation, abstractive summarization, and re-translation into a target language. Additional modules add sentiment analysis (TensorFlow), topic classification (Transformers), and date extraction (Regex) for better document comprehension. Made available in an accessible Gradio interface, the current research shows a real-world application of libraries, models, and APIs to close the language gap and enhance access to information in image media across different linguistic environments
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Hrishit Madhavi, Jacob Cherian, Yuvraj Khamkar, Dhananjay Bhagat. 2025-05-16. Low-Resource Language Processing: An OCR-Driven Summarization and Translation Pipeline. https://arxiv.org/abs/2505.11177
Cite the original work for its findings. Save a collection to share your selection of sources.