SearcharxivSearch

arXiv subjects

Ozan Evkaya

Publications and source records attributed to Ozan Evkaya.

3 recordsLinked to original sources

Welcome to the Statverse: A Metaverse for Data Science

This paper introduces the Statverse, a Metaverse framework designed to revolutionize statistical education in the digital age. Our key goal is to report our progress and encourage others to integrate similar strategies into their programs. The proposed framework seamlessly integrates the physical and digital realms to provide an immersive environment for the nuanced representation of complex statistical concepts. Finally, we discuss the potential impact of Statverse on advancing Statistical Education, offering a transformative approach to teaching and learning in the digital age. Statverse is the outcome of an academic partnership between Universidad Técnica Federico Santa María (UTFSM) and the University of Edinburgh (UoE).

stat.OT

Using ChatGPT for Data Science Analyses

As a result of recent advancements in generative AI, the field of data science is prone to various changes. The way practitioners construct their data science workflows is now irreversibly shaped by recent advancements, particularly by tools like OpenAI's Data Analysis plugin. While it offers powerful support as a quantitative co-pilot, its limitations demand careful consideration in empirical analysis. This paper assesses the potential of ChatGPT for data science analyses, illustrating its capabilities for data exploration and visualization, as well as for commonly used supervised and unsupervised modeling tasks. While we focus here on how the Data Analysis plugin can serve as co-pilot for Data Science workflows, its broader potential for automation is implicit throughout.

cs.LG

Cluster-specific ranking and variable importance for Scottish regional deprivation via vine mixtures

Socioeconomic deprivation is a key determinant of public health, as highlighted by the Scottish Government's Scottish Index of Multiple Deprivation (SIMD). We propose an approach for clustering Scottish zones based on multiple deprivation indicators using vine mixture models. This framework uses the flexibility of vine copulas to capture tail dependent and asymmetric relationships among the indicators. From the fitted vine mixture model, we obtain posterior probabilities for each zone's membership in clusters. This allows the construction of a cluster-driven deprivation ranking by sorting zones according to their probability of belonging to the most deprived cluster. To assess variable importance in this unsupervised learning setting, we adopt a leave-one-variable-out procedure by refitting the model without each variable and calculating the resulting change in the Bayesian information criterion. Our analysis of 21 continuous indicators across 1964 zones in Glasgow and the surrounding areas in Scotland shows that socioeconomic measures, particularly income and employment rates, are major drivers of deprivation, while certain health- and crime-related indicators appear less influential. These findings are consistent across the approach of variable importance and the analysis of the fitted vine structures of the identified clusters.

stat.AP