Searcharxiv⌕ Search

arXiv subjects

Cyril Labbé

Publications and source records attributed to Cyril Labbé.

At least 19 recordsLinked to original sources

From skew to reflected SPDEs

In this article, we prove that the solution of the skew stochastic heat equation converges to the solution of the stochastic heat equation with reflection in the limit of infinite skewness. The result holds at equilibrium. This provides an SPDE analog of the link between skew and reflected Brownian motion. One of the intermediate results is the convergence in law of a penalised Brownian bridge to a 3-Bessel bridge.

math.PR↗

Dynamical interface above a hard wall and reflected SPDE on the half-line

We consider a dynamical random interface on the infinite lattice $\mathbb{N}$ evolving according to a "corner flip" dynamic above a hard wall, with an additional pinning at the origin. We study the stationary fluctuations under a diffusive scaling and prove convergence in law towards the solution of an SPDE of Nualart-Pardoux's type, namely the Reflected Stochastic Heat Equation on the half-line. We also obtain that the law of the 3-dimensional Bessel process is an invariant measure for this SPDE.

math.PR↗

Construction and spectrum of the Anderson Hamiltonian with white noise potential on $\mathbf{R}^2$ and $\mathbf{R}^3$

We propose a simple construction of the Anderson Hamiltonian with white noise potential on $\mathbf{R}^2$ and $\mathbf{R}^3$ based on the solution theory of the parabolic Anderson model. It relies on a theorem of Klein and Landau [KL81] that associates a unique self-adjoint generator to a symmetric semigroup satisfying some mild assumptions. Then, we show that almost surely the spectrum of this random Schrödinger operator is $\mathbf{R}$. To prove this result, we extend the method of Kotani [Kot85] to our setting of singular random operators.

math.PR↗

Top of the spectrum of discrete Anderson Hamiltonians with correlated Gaussian potentials

We investigate the top of the spectrum of discrete Anderson Hamiltonians with correlated Gaussian noise in the large volume limit. The class of Gaussian noises under consideration allows for long-range correlations. We show that the largest eigenvalues converge to a Poisson point process and we obtain a very precise description of the associated eigenfunctions near their localisation centres. We also relate these localisation centres with the locations of the maxima of the noise. Actually, our analysis reveals that this relationship depends in a subtle way on the behaviour near $0$ of the covariance function of the noise: in some situations, the largest eigenfunctions are not associated with the largest values of the noise.

math.PR↗

Convergence of dynamical stationary fluctuations

We present a general black box theorem that ensures convergence of a sequence of stationary Markov processes, provided a few assumptions are satisfied. This theorem relies on a control of the resolvents of the sequence of Markov processes, and on a suitable characterization of the resolvents of the limit. One major advantage of this approach is that it circumvents the use of the Boltzmann-Gibbs principle: for instance, we deduce in a rather simple way that the stationary fluctuations of the one-dimensional zero-range process converge to the stochastic heat equation. More importantly, it allows to establish results that were probably out of reach of existing methods: using the black box result, we are able to prove that the stationary fluctuations of a discrete model of ordered interfaces, that was considered previously in the statistical physics literature, converge to a system of reflected stochastic PDEs.

math.PR↗

Detection of metadata manipulations: Finding sneaked references in the scholarly literature

We report evidence of a new set of sneaked references discovered in the scientific literature. Sneaked references are references registered in the metadata of publications without being listed in reference section or in the full text of the actual publications where they ought to be found. We document here 80,205 references sneaked in metadata of the International Journal of Innovative Science and Research Technology (IJISRT). These sneaked references are registered with Crossref and all cite -- thus benefit -- this same journal. Using this dataset, we evaluate three different methods to automatically identify sneaked references. These methods compare reference lists registered with Crossref against the full text or the reference lists extracted from PDF files. In addition, we report attempts to scale the search for sneaked references to the scholarly literature.

cs.DL↗

ChatGPT as speechwriter for the French presidents

Generative AI proposes several large language models (LLMs) to automatically generate a message in response to users' requests. Such scientific breakthroughs promote new writing assistants but with some fears. The main focus of this study is to analyze the written style of one LLM called ChatGPT by comparing its generated messages with those of the recent French presidents. To achieve this, we compare end-of-the-year addresses written by Chirac, Sarkozy, Hollande, and Macron with those automatically produced by ChatGPT. We found that ChatGPT tends to overuse nouns, possessive determiners, and numbers. On the other hand, the generated speeches employ less verbs, pronouns, and adverbs and include, in mean, too standardized sentences. Considering some words, one can observe that ChatGPT tends to overuse "to must" (devoir), "to continue" or the lemma "we" (nous). Moreover, GPT underuses the auxiliary verb "to be" (^etre), or the modal verbs "to will" (vouloir) or "to have to" (falloir). In addition, when a short text is provided as example to ChatGPT, the machine can generate a short message with a style closed to the original wording. Finally, we reveal that ChatGPT style exposes distinct features compared to real presidential speeches.

cs.CL↗

Cutoff phenomenon in nonlinear recombinations

We investigate a quadratic dynamical system known as nonlinear recombinations. This system models the evolution of a probability measure over the Boolean cube, converging to the stationary state obtained as the product of the initial marginals. Our main result reveals a cutoff phenomenon for the total variation distance in both discrete and continuous time. Additionally, we derive the explicit cutoff profiles in the case of monochromatic initial distributions. These profiles are different in the discrete and continuous time settings. The proof leverages a pathwise representation of the solution in terms of a fragmentation process associated to a binary tree. In continuous time, the underlying binary tree is given by a branching random process, thus requiring a more elaborate probabilistic analysis.

math.PR↗

Detection of tortured phrases in scientific literature

This paper presents various automatic detection methods to extract so called tortured phrases from scientific papers. These tortured phrases, e.g. flag to clamor instead of signal to noise, are the results of paraphrasing tools used to escape plagiarism detection. We built a dataset and evaluated several strategies to flag previously undocumented tortured phrases. The proposed and tested methods are based on language models and either on embeddings similarities or on predictions of masked token. We found that an approach using token prediction and that propagates the scores to the chunk level gives the best results. With a recall value of .87 and a precision value of .61, it could retrieve new tortured phrases to be submitted to domain experts for validation.

cs.IR↗

NanoNER: Named Entity Recognition for nanobiology using experts' knowledge and distant supervision

Here we present the training and evaluation of NanoNER, a Named Entity Recognition (NER) model for Nanobiology. NER consists in the identification of specific entities in spans of unstructured texts and is often a primary task in Natural Language Processing (NLP) and Information Extraction. The aim of our model is to recognise entities previously identified by domain experts as constituting the essential knowledge of the domain. Relying on ontologies, which provide us with a domain vocabulary and taxonomy, we implemented an iterative process enabling experts to determine the entities relevant to the domain at hand. We then delve into the potential of distant supervision learning in NER, supporting how this method can increase the quantity of annotated data with minimal additional manpower. On our full corpus of 728 full-text nanobiology articles, containing more than 120k entity occurrences, NanoNER obtained a F1-score of 0.98 on the recognition of previously known entities. Our model also demonstrated its ability to discover new entities in the text, with precision scores ranging from 0.77 to 0.81. Ablation experiments further confirmed this and allowed us to assess the dependency of our approach on the external resources. It highlighted the dependency of the approach to the resource, while also confirming its ability to rediscover up to 30% of the ablated terms. This paper details the methodology employed, experimental design, and key findings, providing valuable insights and directions for future related researches on NER in specialized domain. Furthermore, since our approach require minimal manpower , we believe that it can be generalized to other specialized fields.

cs.IR↗

Sneaked references: Cooked reference metadata inflate citation counts

We report evidence of an undocumented method to manipulate citation counts involving 'sneaked' references. Sneaked references are registered as metadata for scientific articles in which they do not appear. This manipulation exploits trusted relationships between various actors: publishers, the Crossref metadata registration agency, digital libraries, and bibliometric platforms. By collecting metadata from various sources, we show that extra undue references are actually sneaked in at Digital Object Identifier (DOI) registration time, resulting in artificially inflated citation counts. As a case study, focusing on three journals from a given publisher, we identified at least 9% sneaked references (5,978/65,836) mainly benefiting two authors. Despite not existing in the articles, these sneaked references exist in metadata registries and inappropriately propagate to bibliometric dashboards. Furthermore, we discovered 'lost' references: the studied bibliometric platform failed to index at least 56% (36,939/65,836) of the references listed in the HTML version of the publications. The extent of the sneaked and lost references in the global literature remains unknown and requires further investigations. Bibliometric platforms producing citation counts should identify, quantify, and correct these flaws to provide accurate data to their patrons and prevent further citation gaming.

cs.DL↗

Anderson localization for the $1$-d Schrödinger operator with white noise potential

We consider the random Schrödinger operator on $\mathbb{R}$ obtained by perturbing the Laplacian with a white noise. We prove that Anderson localization holds for this operator: almost surely the spectral measure is pure point and the eigenfunctions are exponentially localized. We give two separate proofs of this result. We also present a detailed construction of the operator and relate it to the parabolic Anderson model. Finally, we discuss the case where the noise is smoothed out.

math.PR↗

Investigating the detection of Tortured Phrases in Scientific Literature

With the help of online tools, unscrupulous authors can today generate a pseudo-scientific article and attempt to publish it. Some of these tools work by replacing or paraphrasing existing texts to produce new content, but they have a tendency to generate nonsensical expressions. A recent study introduced the concept of 'tortured phrase', an unexpected odd phrase that appears instead of the fixed expression. E.g. counterfeit consciousness instead of artificial intelligence. The present study aims at investigating how tortured phrases, that are not yet listed, can be detected automatically. We conducted several experiments, including non-neural binary classification, neural binary classification and cosine similarity comparison of the phrase tokens, yielding noticeable results.

cs.CL↗

The 'Problematic Paper Screener' automatically selects suspect publications for post-publication (re)assessment

Post publication assessment remains necessary to check erroneous or fraudulent scientific publications. We present an online platform, the 'Problematic Paper Screener' (https://www.irit.fr/~Guillaume.Cabanac/problematic-paper-screener) that leverages both automatic machine detection and human assessment to identify and flag already published problematic articles. We provide a new effective tool to curate the scientific literature.

cs.DL↗

Improper legitimization of hijacked journals through citations

The goal is to study the prevalence of citajacked papers: papers in authentic scientific journals citing hijacked journals, in academic literature. A Citejacked detector was designed as a part of the Problematic Paper Screener (https://www.irit.fr/~Guillaume.Cabanac/problematic-paper-screener/citejacked) to trace if the references to articles originating from hijacked journals infiltrate scientific communication. A full-text search was performed between November 2021 and January 2022 in the Dimensions database using the name of 1 of the 12 hijacked journals. The analysis of the bibliography in these articles revealed that 828 of them cite unreliable articles from hijacked journals. During 01.Jan.2021-31.Jan.2022, an average of 2 citejacked articles has been published daily in established journals. Given the limited number of titles included in this study, the phenomenon might be wider and is not yet systematically studied.

cs.DL↗

Universal cutoff for Dyson Ornstein Uhlenbeck process

We study the Dyson-Ornstein-Uhlenbeck diffusion process, an evolving gas of interacting particles. Its invariant law is the beta Hermite ensemble of random matrix theory, a non-product log-concave distribution. We explore the convergence to equilibrium of this process for various distances or divergences, including total variation, relative entropy, and transportation cost. When the number of particles is sent to infinity, we show that a cutoff phenomenon occurs: the distance to equilibrium vanishes abruptly at a critical time. A remarkable feature is that this critical time is independent of the parameter beta that controls the strength of the interaction, in particular the result is identical in the non-interacting case, which is nothing but the Ornstein-Uhlenbeck process. We also provide a complete analysis of the non-interacting case that reveals some new phenomena. Our work relies among other ingredients on convexity and functional inequalities, exact solvability, exact Gaussian formulas, coupling arguments, stochastic calculus, variational formulas and contraction properties. This work leads, beyond the specific process that we study, to questions on the high-dimensional analysis of heat kernels of curved diffusions.

math.PR↗

Asymptotic of the smallest eigenvalues of the continuous Anderson Hamiltonian in $d \leq 3$

We consider the continuous Anderson Hamiltonian with white noise potential on $(-L/2,L/2)^d$ in dimension $d\le 3$, and derive the asymptotic of the smallest eigenvalues when $L$ goes to infinity. We show that these eigenvalues go to $-\infty$ at speed $(\log L)^{1/(2-d/2)}$ and identify the prefactor in terms of the optimal constant of the Gagliardo-Nirenberg inequality. This result was already known in dimensions $1$ and $2$, but appears to be new in dimension $3$. We present some conjectures on the fluctuations of the eigenvalues and on the asymptotic shape of the corresponding eigenfunctions near their localisation centers.

math.PR↗

Hydrodynamic limit and cutoff for the biased adjacent walk on the simplex

We investigate the asymptotic in $N$ of the mixing times of a Markov dynamics on $N-1$ ordered particles in an interval. This dynamics consists in resampling at independent Poisson times each particle according to a probability measure on the segment formed by its nearest neighbours. In the setting where the resampling probability measures are symmetric, the asymptotic of the mixing times were obtained and a cutoff phenomenon holds. In the present work, we focus on an asymmetric version of the model and we establish a cutoff phenomenon. An important part of our analysis consists in the derivation of a hydrodynamic limit, which is given by a non-linear Hamilton-Jacobi equation with degenerate boundary conditions.

math.PR↗