SearcharxivSearch

arXiv subjects

Sachin Sharma

Publications and source records attributed to Sachin Sharma.

At least 19 recordsLinked to original sources

TrOCR for Medieval HTR: A Systematic Ablation Study with Cross-Dataset Validation

Fine-tuning transformer-based handwritten text recognition (HTR) models on medieval manuscripts is challenging because these models are pre-trained on modern text and must adapt to a very different visual domain. This paper studies how three controllable fine-tuning choices (contrast normalization, data augmentation, and layer freezing) affect recognition accuracy when adapting TrOCR to small historical datasets. We run controlled experiments on a 13th-century Italian manuscript (I-CT 91 "Cortonese") and replicate the same experimental grid on the public READ-16 benchmark as robustness evidence. On Cortonese, our best configuration achieves 8.03% character error rate (CER). Statistical comparisons across 13 configurations show that freezing up to three encoder layers or six decoder layers does not significantly harm accuracy, while deeper freezing becomes progressively detrimental. Removing contrast normalization (CLAHE) yields 7.84% CER, comparable to a domain-specialized baseline, suggesting strong optimization can reduce reliance on image preprocessing. Cross-dataset validation on READ-16 shows that decoder freezing thresholds transfer more robustly than encoder thresholds, and combined freezing strategies require dataset-specific re-validation. Finally, we use Grad-CAM gradient attributions and decoder cross-attention maps to diagnose error patterns and failure modes revealed by the ablations. Source code is available at https://github.com/LaudareProject/TrOCR-analysis

cs.CV

The Cognitive Kardashev Scale: Quantifying the Material Envelope of Civilisational Computation

How much thinking can a civilisation do? Kardashev ranked civilisations by the energy they command. This paper borrows his ladder and asks how much machine cognition each rung could support. The arithmetic is deliberately simple. A civilisation has some total power. Only a fraction of that can be spared for computing, and each joule spent buys computation at whatever efficiency the hardware of the day has reached. The product of the three sets a ceiling on machine thought. To keep the resulting quantities intelligible, I express them in units of the human brain's own processing rate, as a rough yardstick rather than a claim about minds. Calibrating the ceiling against present-day supercomputers and AI accelerators led me to two conclusions I did not expect at the outset. Even today's energy supply could support far more machine cognition than humanity actually uses, so physical capacity is not what binds. And whether energy or hardware efficiency becomes the constraint over the coming decade turns on engineering choices that have not yet been made. On the question that may matter most, who gets access to the cognition it describes, the scale is silent. That is a matter of political economy, and the calibration is offered as an input to that debate.

physics.soc-ph

Uber's Failover Architecture: Reconciling Reliability and Efficiency in Hyperscale Microservice Infrastructure

Operating a global, real-time platform at Uber's scale requires infrastructure that is both resilient and cost-efficient. Historically, reliability was ensured through a costly 2x capacity model--each service provisioned to handle global traffic independently across two regions--leaving half the fleet idle. We present Uber's Failover Architecture (UFA), which replaces the uniform 2x model with a differentiated architecture aligned to business criticality. Critical services retain failover guarantees, while non-critical services opportunistically use failover buffer capacity reserved for critical services during steady state. During rare "full-peak" failovers, non-critical services are selectively preempted and rapidly restored, with differentiated Service-Level Agreements (SLAs) using on-demand capacity. Automated safeguards, including dependency analysis and regression gates, ensure critical services continue to function even while non-critical services are unavailable. The quantitative impact is significant: UFA reduces steady-state provisioning from 2x to 1.3x, raising utilization from ~20% to ~30% while sustaining 99.97% availability. To date, UFA has hardened over 4,000 unsafe dependencies, eliminated over one million CPU cores from a baseline of about four million cores.

cs.DC

Pre-Hoc Predictions in AutoML: Leveraging LLMs to Enhance Model Selection and Benchmarking for Tabular datasets

The field of AutoML has made remarkable progress in post-hoc model selection, with libraries capable of automatically identifying the most performing models for a given dataset. Nevertheless, these methods often rely on exhaustive hyperparameter searches, where methods automatically train and test different types of models on the target dataset. Contrastingly, pre-hoc prediction emerges as a promising alternative, capable of bypassing exhaustive search through intelligent pre-selection of models. Despite its potential, pre-hoc prediction remains under-explored in the literature. This paper explores the intersection of AutoML and pre-hoc model selection by leveraging traditional models and Large Language Model (LLM) agents to reduce the search space of AutoML libraries. By relying on dataset descriptions and statistical information, we reduce the AutoML search space. Our methodology is applied to the AWS AutoGluon portfolio dataset, a state-of-the-art AutoML benchmark containing 175 tabular classification datasets available on OpenML. The proposed approach offers a shift in AutoML workflows, significantly reducing computational overhead, while still selecting the best model for the given dataset.

cs.LG

Shock wave bending around a dusty plasma void

We report on experimental observations of the bending of a dust acoustic shock wave around a dust void region. This phenomenon occurs as a planar shock wavefront encounters a compressible obstacle in the form of a void whose size is larger than the wavelength of the wave. As they collide, the central portion of the wavefront, that is the first to touch the void, is blocked while the rest of the front continues to propagate, resulting in an inward bending of the shock wave. The bent shock wave eventually collapses, leading to the transient trapping of dust particles in the void. Subsequently, a Coulomb explosion of the trapped particles generates a bow shock. The experiments have been carried out in a DC glow discharge plasma, where the shock wave and the void are simultaneously created as self-excited modes of a three-dimensional dust cloud. The salient features of this phenomenon are reproduced in molecular dynamics simulations, which provide valuable insights into the underlying dynamics of this interaction.

physics.plasm-ph

Experimenting active and sequential learning in a medieval music manuscript

Optical Music Recognition (OMR) is a cornerstone of music digitization initiatives in cultural heritage, yet it remains limited by the scarcity of annotated data and the complexity of historical manuscripts. In this paper, we present a preliminary study of Active Learning (AL) and Sequential Learning (SL) tailored for object detection and layout recognition in an old medieval music manuscript. Leveraging YOLOv8, our system selects samples with the highest uncertainty (lowest prediction confidence) for iterative labeling and retraining. Our approach starts with a single annotated image and successfully boosts performance while minimizing manual labeling. Experimental results indicate that comparable accuracy to fully supervised training can be achieved with significantly fewer labeled examples. We test the methodology as a preliminary investigation on a novel dataset offered to the community by the Anonymous project, which studies laude, a poetical-musical genre spread across Italy during the 12th-16th Century. We show that in the manuscript at-hand, uncertainty-based AL is not effective and advocates for more usable methods in data-scarcity scenarios.

cs.CV

Drafting the Landscape of Computational Musicology Tools: a Survey-Based Approach

Since the 60s, musicology has been increasingly impacted by computational tools in various ways, from systematic analysis approaches to modeling of creativity. This article presents a comprehensive assessment of the current state of Computational Musicology tools based on survey data collected from practitioners in the field. We gathered information on tool usage patterns, common analytical tasks, user satisfaction levels, data characteristics, and prioritized features across four distinct domains: symbolic music, music-related imagery, audio, and text. Our findings reveal significant gaps between current tooling capabilities and user needs, highlighting some limitations of these tools across all domains. This assessment contributes to the ongoing dialogue between tool developers and music scholars, aiming to enhance the effectiveness and accessibility of computational methods in musicological research.

cs.DL

AI-driven Web Application for Early Detection of Sudden Death Syndrome (SDS) in Soybean Leaves Using Hyperspectral Images and Genetic Algorithm

Sudden Death Syndrome (SDS), caused by Fusarium virguliforme, poses a significant threat to soybean production. This study presents an AI-driven web application for early detection of SDS on soybean leaves using hyperspectral imaging, enabling diagnosis prior to visible symptom onset. Leaf samples from healthy and inoculated plants were scanned using a portable hyperspectral imaging system (398-1011 nm), and a Genetic Algorithm was employed to select five informative wavelengths (505.4, 563.7, 712.2, 812.9, and 908.4 nm) critical for discriminating infection status. These selected bands were fed into a lightweight Convolutional Neural Network (CNN) to extract spatial-spectral features, which were subsequently classified using ten classical machine learning models. Ensemble classifiers (Random Forest, AdaBoost), Linear SVM, and Neural Net achieved the highest accuracy (>98%) and minimal error across all folds, as confirmed by confusion matrices and cross-validation metrics. Poor performance by Gaussian Process and QDA highlighted their unsuitability for this dataset. The trained models were deployed within a web application that enables users to upload hyperspectral leaf images, visualize spectral profiles, and receive real-time classification results. This system supports rapid and accessible plant disease diagnostics, contributing to precision agriculture practices. Future work will expand the training dataset to encompass diverse genotypes, field conditions, and disease stages, and will extend the system for multiclass disease classification and broader crop applicability.

cs.CV

YouLeQD: Decoding the Cognitive Complexity of Questions and Engagement in Online Educational Videos from Learners' Perspectives

Questioning is a fundamental aspect of education, as it helps assess students' understanding, promotes critical thinking, and encourages active engagement. With the rise of artificial intelligence in education, there is a growing interest in developing intelligent systems that can automatically generate and answer questions and facilitate interactions in both virtual and in-person education settings. However, to develop effective AI models for education, it is essential to have a fundamental understanding of questioning. In this study, we created the YouTube Learners' Questions on Bloom's Taxonomy Dataset (YouLeQD), which contains learner-posed questions from YouTube lecture video comments. Along with the dataset, we developed two RoBERTa-based classification models leveraging Large Language Models to detect questions and analyze their cognitive complexity using Bloom's Taxonomy. This dataset and our findings provide valuable insights into the cognitive complexity of learner-posed questions in educational videos and their relationship with interaction metrics. This can aid in the development of more effective AI models for education and improve the overall learning experience for students.

cs.CL

Observation of Kolmogorov turbulence due to multiscale vortices in dusty plasma experiments

We report the experimental observation of fully developed Kolmogorov turbulence originating from self-excited vortex flows in a three-dimensional (3D) dust cloud. The characteristic -5/3 scaling of three-dimensional Kolmogorov turbulence is universally observed in both the spatial and temporal power spectra. Additionally, the 2/3 scaling in the second-order structure function further confirms the presence of Kolmogorov turbulence. We also identified a slight deviation in the tails of the probability distribution functions for velocity gradients. The dust cloud formed in the diffused region away from the electrode and above the glass device surface in the glow discharge experiments. The dust rotation was observed in multiple experimental campaigns under different discharge conditions at different spatial locations and background plasma environments.

physics.plasm-ph

Analyzing LLM Usage in an Advanced Computing Class in India

This study examines the use of large language models (LLMs) by undergraduate and graduate students for programming assignments in advanced computing classes. Unlike existing research, which primarily focuses on introductory classes and lacks in-depth analysis of actual student-LLM interactions, our work fills this gap. We conducted a comprehensive analysis involving 411 students from a Distributed Systems class at an Indian university, where they completed three programming assignments and shared their experiences through Google Form surveys. Our findings reveal that students leveraged LLMs for a variety of tasks, including code generation, debugging, conceptual inquiries, and test case creation. They employed a spectrum of prompting strategies, ranging from basic contextual prompts to advanced techniques like chain-of-thought prompting and iterative refinement. While students generally viewed LLMs as beneficial for enhancing productivity and learning, we noted a concerning trend of over-reliance, with many students submitting entire assignment descriptions to obtain complete solutions. Given the increasing use of LLMs in the software industry, our study highlights the need to update undergraduate curricula to include training on effective prompting strategies and to raise awareness about the benefits and potential drawbacks of LLM usage in academic settings.

cs.HC

Precision-controlled ultrafast electron microscope platforms. A case study: Multiple-order coherent phonon dynamics in 1T-TaSe$_2$ probed at 50 femtosecond - 10 femtometer scales

We report on the first detailed beam test attesting the fundamental principle behind the development of high-current-efficiency ultrafast electron microscope systems where a radio-frequency cavity is incorporated as a condenser lens in the beam delivery system. To allow the experiment to be carried out with a sufficient resolution to probe the performance at the emittance floor, a new cascade loop RF controller system is developed to reduce the RF noise floor. Temporal resolution at 50 femtoseconds in full-width-at-half-maximum and detection sensitivity better than 1% are demonstrated on exfoliated 1T-TaSe$_2$ layers where the multi-order edge-mode coherent phonon excitation is employed as the standard candle to benchmark the performance. The high temporal resolution and the significant visibility to very low dynamical contrast in diffraction signals give strong support to the working principle of the high-brightness beam delivery via phase-space manipulation in the electron microscope system.

cond-mat.mtrl-sci

Global Message Ordering using Distributed Kafka Clusters

In contemporary distributed systems, logs are produced at an astounding rate, generating terabytes of data within mere seconds. These logs, containing pivotal details like system metrics, user actions, and diverse events, are foundational to the system's consistent and accurate operations. Precise log ordering becomes indispensable to avert potential ambiguities and discordances in system functionalities. Apache Kafka, a prevalent distributed message queue, offers significant solutions to various distributed log processing challenges. However, it presents an inherent limitation while Kafka ensures the in-order delivery of messages within a single partition to the consumer, it falls short in guaranteeing a global order for messages spanning multiple partitions. This research delves into innovative methodologies to achieve global ordering of messages within a Kafka topic, aiming to bolster the integrity and consistency of log processing in distributed systems. Our code is available on GitHub.

cs.DC

Ultrafast Hot-Carrier cooling in Quasi Freestanding Bilayer Graphene with Hydrogen Intercalated Atoms

We perform a femtosecond-THz optical pump-probe spectroscopy to investigate the cooling dynamics of hot carriers in quasi-free standing bilayer epitaxial graphene. We observed longer decay time constants, in the range of 2.6 to 6.4 ps, compared to previous studies on monolayer graphene, which increase nonlinearly with excitation intensity. The increased relaxation times are due to the decoupling of the graphene layer from the SiC substrate after hydrogen intercalation which increases the distance between graphene and substrate. Furthermore, our measurements do not show that the supercollision mechanism is related to the cooling process of the hot carriers, which is ultimately achieved by electron-optical phonon scattering.

cond-mat.mes-hall

Room-temperature photo-chromism of silicon vacancy centers in CVD diamond

The silicon-vacancy (SiV) center in diamond is typically found in three stable charge states, SiV0, SiV- and SiV2-, but studying the processes leading to their formation is challenging, especially at room temperature, due to their starkly different photo-luminescence rates. Here, we use confocal fluorescence microscopy to activate and probe charge interconversion between all three charge states under ambient conditions. In particular, we witness the formation of SiV0 via the two-step capture of diffusing, photo-generated holes, a process we expose both through direct SiV0 fluorescence measurements at low temperatures and confocal microscopy observations in the presence of externally applied electric fields. Further, we show that continuous red illumination induces the converse process, first transforming SiV0 into SiV-, then into SiV2-. Our results shed light on the charge dynamics of SiV and promise opportunities for nanoscale sensing and quantum information processing.

cond-mat.mtrl-sci

Characterization of Frequent Online Shoppers using Statistical Learning with Sparsity

Developing shopping experiences that delight the customer requires businesses to understand customer taste. This work reports a method to learn the shopping preferences of frequent shoppers to an online gift store by combining ideas from retail analytics and statistical learning with sparsity. Shopping activity is represented as a bipartite graph. This graph is refined by applying sparsity-based statistical learning methods. These methods are interpretable and reveal insights about customers' preferences as well as products driving revenue to the store.

cs.LG

Magnetically collected platinum/nickel alloy nanoparticles -- insight into low noble metal content catalysts for hydrogen evolution reaction

The hydrogen evolution reaction (HER) is a key process in electrochemical water splitting. To lower the cost and environmental impact of this process, it is highly motivated to develop electrocatalysts with low or no content of noble metals. Here we report on a novel and ingenious synthesis of hybrid PtxNi1-x electrocatalysts in the form of a nanoparticle-necklace structure named nanotrusses, with very low noble metal content. The nanotruss structure possesses important features, such as good conductivity, high surface area, strong interlinking and substrate adhesion, which renders for an excellent HER activity. Specifically, the best performing Pt0.05Ni0.95 sample, demonstrates a Tafel slope of 30 mV dec-1 in 0.5 M H2SO4, and an overpotential of 20 mV at a current density of 10 mA cm-2 with high stability. The impressive catalytic performance is further rationalized in a theoretical study, which provides insight into the mechanism for how such small platinum content can allow for close-to-optimal adsorption energies for hydrogen.

cond-mat.mtrl-sci

Boomerang: Rebounding the Consequences of Reputation Feedback on Crowdsourcing Platforms

Paid crowdsourcing platforms suffer from low-quality work and unfair rejections, but paradoxically, most workers and requesters have high reputation scores. These inflated scores, which make high-quality work and workers difficult to find, stem from social pressure to avoid giving negative feedback. We introduce Boomerang, a reputation system for crowdsourcing that elicits more accurate feedback by rebounding the consequences of feedback directly back onto the person who gave it. With Boomerang, requesters find that their highly-rated workers gain earliest access to their future tasks, and workers find tasks from their highly-rated requesters at the top of their task feed. Field experiments verify that Boomerang causes both workers and requesters to provide feedback that is more closely aligned with their private opinions. Inspired by a game-theoretic notion of incentive-compatibility, Boomerang opens opportunities for interaction design to incentivize honest reporting over strategic dishonesty.

cs.CY