SearcharxivSearch

arXiv subjects

Tristan Miller

Publications and source records attributed to Tristan Miller.

7 recordsLinked to original sources

Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation

With the advent of large multimodal language models, science is now at a threshold of an AI-based technological transformation. An emerging ecosystem of models and tools aims to support researchers throughout the scientific lifecycle, including (1) searching for relevant literature, (2) generating research ideas and conducting experiments, (3) producing text-based content, (4) creating multimodal artifacts such as figures and diagrams, and (5) evaluating scientific work, as in peer review. In this survey, we provide a curated overview of literature representative of the core techniques, evaluation practices, and emerging trends in AI-assisted scientific discovery. Across the five tasks outlined above, we discuss datasets, methods, results, evaluation strategies, limitations, and ethical concerns, including risks to research integrity through the misuse of generative models. We aim for this survey to serve both as an accessible, structured orientation for newcomers to the field, as well as a catalyst for new AI-based initiatives and their integration into future ``AI4Science'' systems.

cs.CL

Remembering Netizens: An interview with Ronda Hauben, co-author of Netizens: On the history and impact of Usenet and the Internet (1997)

Netizens, Michael and Ronda Hauben's foundational treatise on Usenet and the Internet, was first published in print 25 years ago. In this piece, we trace the history and impact of the book and of Usenet itself, contextualising them within the contemporary and modern-day scholarship on virtual communities, online culture, and Internet history. We discuss the Net as a tool of empowerment, and touch on the social, technical, and economic issues related to the maintenance of shared network infrastructures and to the preservation and commodification of Usenet archives. Our interview with Ronda Hauben offers a retrospective look at the development of online communities, their impact, and how they are studied. She recounts her own introduction to the online world, as well as the impetus and writing process for Netizens. She presents Michael Hauben's conception of "netizens" as contributory citizens of the Net (rather than mere users of it) and the "electronic commons" they built up, and argues that this collaborative and collectivist model has been overwhelmed and endangered by the privatisation and commercialisation of the Internet and its communities.

cs.CY

Predicting the Humorousness of Tweets Using Gaussian Process Preference Learning

Most humour processing systems to date make at best discrete, coarse-grained distinctions between the comical and the conventional, yet such notions are better conceptualized as a broad spectrum. In this paper, we present a probabilistic approach, a variant of Gaussian process preference learning (GPPL), that learns to rank and rate the humorousness of short texts by exploiting human preference judgments and automatically sourced linguistic annotations. We apply our system, which is similar to one that had previously shown good performance on English-language one-liners annotated with pairwise humorousness annotations, to the Spanish-language data set of the HAHA@IberLEF2019 evaluation campaign. We report system performance for the campaign's two subtasks, humour detection and funniness score prediction, and discuss some issues arising from the conversion between the numeric scores used in the HAHA@IberLEF2019 data and the pairwise judgment annotations required for our method.

cs.CL

GPP, the Generic Preprocessor

In computer science, a preprocessor (or macro processor) is a tool that programatically alters its input, typically on the basis of inline annotations, to produce data that serves as input for another program. Preprocessors are used in software development and document processing workflows to translate or extend programming or markup languages, as well as for conditional or pattern-based generation of source code and text. Early preprocessors were relatively simple string replacement tools that were tied to specific programming languages and application domains, and while these have since given rise to more powerful, general-purpose tools, these often require the user to learn and use complex macro languages with their own syntactic conventions. In this paper, we present GPP, an extensible, general-purpose preprocessor whose principal advantage is that its syntax and behaviour can be customized to suit any given preprocessing task. This makes GPP of particular benefit to research applications, where it can be easily adapted for use with novel markup, programming, and control languages.

cs.PL

Cross-topic Argument Mining from Heterogeneous Sources Using Attention-based Neural Networks

Argument mining is a core technology for automating argument search in large document collections. Despite its usefulness for this task, most current approaches to argument mining are designed for use only with specific text types and fall short when applied to heterogeneous texts. In this paper, we propose a new sentential annotation scheme that is reliably applicable by crowd workers to arbitrary Web texts. We source annotations for over 25,000 instances covering eight controversial topics. The results of cross-topic experiments show that our attention-based neural network generalizes best to unseen topics and outperforms vanilla BiLSTM models by 6% in accuracy and 11% in F-score.

cs.CL

Stimulated emission of Cooper pairs in a high-temperature cuprate superconductor

The concept of stimulated emission of bosons has played an important role in modern science and technology, and constitutes the working principle for lasers. In a stimulated emission process, an incoming photon enhances the probability that an excited atomic state will transition to a lower energy state and generate a second photon of the same energy. It is expected, but not experimentally shown, that stimulated emission contributes significantly to the zero resistance current in a superconductor by enhancing the probability that scattered Cooper pairs will return to the macroscopically occupied condensate instead of entering any other state. Here, we use time- and angle-resolved photoemission spectroscopy to study the initial rise of the non-equilibrium quasiparticle population in a Bi$_2$Sr$_2$CaCu$_2$O$_{8+δ}$ cuprate superconductor induced by an ultrashort laser pulse. Our finding reveals significantly slower buildup of quasiparticles in the superconducting state than in the normal state. The slower buildup only occurs when the pump pulse is too weak to deplete the superconducting condensate, and for cuts inside the Fermi arc region. We propose this is a manifestation of stimulated recombination of broken Cooper pairs, and signals an important momentum space dichotomy in the formation of Cooper pairs inside and outside the Fermi arc region.

cond-mat.supr-con

Signatures of superconductivity and pseudogap formation in non-equilibrium nodal quasiparticles revealed by ultrafast angle-resolved photoemission

We use time- and angle-resolved photoemission to measure the nodal non-equilibrium electronic states in various dopings of Bi$_2$Sr$_2$CaCu$_2$O$_{8+δ}$. We find that the initial pump-induced transient signal of these ungapped states is strongly affected by the onset of the superconducting gap at $T_c$, superconducting pairing fluctuations at $T_p$, and the pseudogap at $T^*$. Moreover, $T_p$ marks a suggestive threshold in the fluence-dependent transient signal, with the appearance of a critical fluence below $T_p$ that corresponds to the energy required to break apart all Cooper pairs. These results challenge the notion of a nodal-antinodal dichotomy in cuprate superconductors by establishing a new link between nodal quasiparticles and the cuprate phase diagram.

cond-mat.supr-con