SearcharxivSearch

arXiv subjects

Samopriya Basu

Publications and source records attributed to Samopriya Basu.

5 recordsLinked to original sources

Non-Parametric Model Calibration with Stochastic Control Parameters

We present a method for calibrating a computer model using non-parametric techniques where the inputs are stochastic but include calibration parameters whose distributions are unknown and control parameters whose distributions are specified. Our solution gives a distributional estimate over the input space that is consistent with observed field data, while also preserving the distribution of the known marginal of the control parameters. This property is desirable since stochastic inputs often include physical processes affecting the experimental conditions, and a scientifically plausible calibration estimate should preserve well-established distributional properties of these inputs. The method builds on recently developed non-parametric computer model calibration techniques based on the disintegration of measure and Bayesian inference.

stat.ME

GPU-accelerated Bayesian inference for block-cave geometry recovery via muon tomography

We describe a Bayesian framework for the inverse problem of geometry recovery of block caving via muon tomography. We work with a low dimensional surface-based representation of the geometry of the block cave, which dramatically reduces the computational requirements of the model while allowing realistic geometries. Adopting a Bayesian approach, we define a prior distribution on the space of geometries that favors realistic cave shapes. Pairing this prior with a likelihood based on the muon tomography forward model, we obtain a posterior distribution over cave geometries using Bayes rule. We obtain approximate samples from this posterior distribution using Markov chain Monte Carlo algorithms running on GPUs, resulting in fast and accurate sampling. We test the fidelity of our methodology by applying it to a simulated block caving scenario for which the ground truth is known. Results show that our method produces sensible geometries that are simultaneously compatible with the data.

stat.AP

Jambu: A historical linguistic database for South Asian languages

We introduce Jambu, a cognate database of South Asian languages which unifies dozens of previous sources in a structured and accessible format. The database includes 287k lemmata from 602 lects, grouped together in 23k sets of cognates. We outline the data wrangling necessary to compile the dataset and train neural models for reflex prediction on the Indo-Aryan subset of the data. We hope that Jambu is an invaluable resource for all historical linguists and Indologists, and look towards further improvement and expansion of the database.

cs.CL

Computational historical linguistics and language diversity in South Asia

South Asia is home to a plethora of languages, many of which severely lack access to new language technologies. This linguistic diversity also results in a research environment conducive to the study of comparative, contact, and historical linguistics -- fields which necessitate the gathering of extensive data from many languages. We claim that data scatteredness (rather than scarcity) is the primary obstacle in the development of South Asian language technology, and suggest that the study of language history is uniquely aligned with surmounting this obstacle. We review recent developments in and at the intersection of South Asian NLP and historical-comparative linguistics, describing our and others' current efforts in this area. We also offer new strategies towards breaking the data barrier.

cs.CL

Bhasacitra: Visualising the dialect geography of South Asia

We present Bhasacitra, a dialect mapping system for South Asia built on a database of linguistic studies of languages of the region annotated for topic and location data. We analyse language coverage and look towards applications to typology by visualising example datasets. The application is not only meant to be useful for feature mapping, but also serves as a new kind of interactive bibliography for linguists of South Asian languages.

cs.CL