Searcharxiv⌕ Search

arXiv subjects

Maximiliano Romero

Publications and source records attributed to Maximiliano Romero.

2 recordsLinked to original sources

Automated Detection and Structuring of Social Tipping Point Evidence in Climate related Documents: A Modular AI Framework

The climate literature has grown faster than review teams can read it. That gap matters most for a concept like the environmental social tipping point, the threshold at which a small change triggers rapid, self-reinforcing change in a social system. Evidence of this kind of shift is usually contained in one or two paragraphs within a longer document. As a result, existing text mining tools-which categorize entire documents by topic or highlight isolated claims-leave an expanding set of important evidence without any systematic method for discovery or organization. This paper presents an open and modular transformer-based framework that detects and structures social tipping point evidence at the passage level. The framework joins five components into a single deployable workflow: a DistilBERT boundary splitter for segmentation, an iteratively augmented RoBERTa classifier for detection, a Mistral 7B model that rewrites each detected passage for clarity, a LLaMA 3.2 3B model that rates the passage against five published social tipping point criteria, and a Milvus vector store for semantic retrieval. The system is wrapped in a Streamlit interface backed by MinIO object storage. Evaluated on a 163-passage benchmark labelled by GPT-4.1 and a 51-passage set reviewed by experts, the splitter surpassed three competing methods on a nine-metric composite score (6.137). The tuned RoBERTa model achieved 71.4 percent accuracy with a Cohen's kappa of 0.337 on the full benchmark, and 87.5 percent accuracy with a kappa of 0.742 on passages with labels, outperforming both a climate-focused model and untuned language models.

cs.CL↗

Partial Identification of Mean Achievement in ILSA Studies with Multi-Stage Stratified Sample Design and Student Non-Participation

International large-scale assessment (ILSA) studies collect information across education systems with the objective of learning about the population-wide distribution of student achievement in the assessment. In this article, we study one of the most fundamental threats that these studies face when justifying the conclusions reached about these distributions: the identification problem that arises from student non-participation during data collection. Recognizing that ILSA studies have traditionally employed a narrow range of strategies to address non-participation, we examine this problem using tools developed within the framework of partial identification of probability distributions. We tailor this framework to the problem of non-participation when data are collected using a multi-stage stratified random sample design, as in most ILSA studies. We demonstrate this approach with application to the International Computer and Information Literacy Study in 2018. We show how to use the framework to assess mean achievement under reasonable and credible sets of assumptions about the non-participating population. We also provide examples of how these results may be reported by agencies that administer ILSA studies. By doing so, we bring to the field of ILSA an alternative strategy for identification, estimation, and reporting of population parameters of interest.

econ.EM↗