SearcharxivSearch

arXiv subjects

Marie-Anne Guerry

Publications and source records attributed to Marie-Anne Guerry.

10 recordsLinked to original sources

A Data-Driven Approach to State Construction in Markov Models

A Markov chain is a widely used stochastic process modelling random events over time. These models are built on subsets of the entire dataset, referred to as states, which are considered to be homogeneous regarding transition probabilities. However, the creation of these states is often disregarded or based on prior assumption, potentially violating the homogeneity requirement and thus decreasing the validity and predictive power of the model. In order to fill this gap, this paper combines supervised feature selection with unsupervised learning techniques for data-driven state construction. Density-based clustering, spectral clustering, and Kohonen self-organizing maps are examined for their ability to identify latent groups without prior assumptions. The contribution of this study is twofold. First, the paper presents a methodological framework for state construction incorporating suitable unsupervised learning techniques, with appropriate measures both for classification performance and Markov model accuracy. Secondly, the framework is tested on an application, resulting in a comparative analysis showing that spectral clustering and Kohonen self-organizing maps are best at capturing inherent structure. These results serve as a cornerstone in providing theoretical and methodological guidance for improving state definition in applied Markov modelling.

stat.ML

Nonnegativity of the second largest eigenvalue of $4 \times 4$ tridiagonal stochastic matrices

The spectral study of nonnegative and more specifically stochastic matrices is an important topic in matrix theory. In this paper, we prove a conjecture, formulated by Ran and Teng, which states that the second largest eigenvalue of an irreducible $4\times4$ tridiagonal stochastic matrix is nonnegative. We establish this conjecture and extend the result to arbitrary $4\times4$ tridiagonal stochastic matrices, including both irreducible and reducible cases.

math.PR

On the Spectral Region of 4-Cycle Stochastic Matrices

We study the spectrum of 4-cycle row-stochastic matrices. For real eigenvalues the spectral region is [-1,1]. For nonreal eigenvalues a+ib we derive necessary conditions in terms of the real and imaginary parts, including the inequality a+|b| <= 1 and the condition (b^2+a^2+a)^2+2a^2-b^2 >= 0. We also prove conversely that every point in the corresponding interior region occurs as an eigenvalue of a 4-cycle matrix. The proof is organized through a reformulation of the characteristic equation, an argument parametrization, a convex-analytic criterion, and explicit boundary constructions. Hence, the spectral region for the 4-cycle row-stochastic matrices is exactly and explicitly determined.

math.SP

Human-in-the-Loop LLM Grading for Handwritten Mathematics Assessments

Providing timely and individualised feedback on handwritten student work is highly beneficial for learning but difficult to achieve at scale. This challenge has become more pressing as generative AI undermines the reliability of take-home assessments, shifting emphasis toward supervised, in-class evaluation. We present a scalable, end-to-end workflow for LLM-assisted grading of short, pen-and-paper assessments. The workflow spans (1) constructing solution keys, (2) developing detailed rubric-style grading keys used to guide the LLM, and (3) a grading procedure that combines automated scanning and anonymisation, multi-pass LLM scoring, automated consistency checks, and mandatory human verification. We deploy the system in two undergraduate mathematics courses using six low-stakes in-class tests. Empirically, LLM assistance reduces grading time by approximately 23% while achieving agreement comparable to, and in several cases tighter than, fully manual grading. Occasional model errors occur but are effectively contained by the hybrid design. Overall, our results show that carefully embedded human-in-the-loop LLM grading can substantially reduce workload while maintaining fairness and accuracy.

cs.CY

Early Evidence of Vibe-Proving with Consumer LLMs: A Case Study on Spectral Region Characterization with ChatGPT-5.2 (Thinking)

Large Language Models (LLMs) are increasingly used as scientific copilots, but evidence on their role in research-level mathematics remains limited, especially for workflows accessible to individual researchers. We present early evidence for vibe-proving with a consumer subscription LLM through an auditable case study that resolves Conjecture 20 of Ran and Teng (2024) on the exact nonreal spectral region of a 4-cycle row-stochastic nonnegative matrix family. We analyze seven shareable ChatGPT-5.2 (Thinking) threads and four versioned proof drafts, documenting an iterative pipeline of generate, referee, and repair. The model is most useful for high-level proof search, while human experts remain essential for correctness-critical closure. The final theorem provides necessary and sufficient region conditions and explicit boundary attainment constructions. Beyond the mathematical result, we contribute a process-level characterization of where LLM assistance materially helps and where verification bottlenecks persist, with implications for evaluation of AI-assisted research workflows and for designing human-in-the-loop theorem proving systems.

cs.AI

Eigenvalue regions and realising monotone stochastic matrices

Eigenvalues of stochastic matrices have been studied from two complementary perspectives. The individual eigenvalues are characterised through the well-established Karpelevich regions. The spectrum as a whole has also been analysed, yielding powerful results such as the Johnson-Loewy-London (JLL) inequalities. Current research now turns toward particular subsets of stochastic matrices, among others the doubly stochastic matrices. This paper studies spectral properties of monotone stochastic matrices which are characterised by the fact that each row stochastically dominates the preceding one, and which arise in contexts such as intergenerational mobility, equal-input models, and credit-rating systems. This paper analyses the dominance matrix associated with a monotone matrix, which is a non-negative matrix that preserves the non-trivial eigenvalues. Properties are established and the conditions are given under which a non-negative matrix can be regarded as a dominance matrix. In analogy with the stochastic matrices, this study examines for the monotone stochastic matrices both the individual eigenvalues as the spectrum as a whole. Individually, the eigenvalue region for all monotone matrices up till order 3 is completely determined, and realising matrices are provided. Collectively, the set of possible pairs of non-trivial eigenvalues arising from 3x3 monotone matrices is characterised, accompanied by realising matrices. In both perspectives, the resulting regions are substantially smaller than those for general stochastic matrices. Finally, this paper proves a reduction theorem stating that, for all n from 4 on, the eigenvalue region of n x n monotone matrices is contained within that of (n-1) x (n-1) stochastic matrices.

math.SP

State Re-union Maintainability for Semi-Markov Models in Manpower Planning

In previous research the importance of both Markov and semi-Markov models in manpower planning is highlighted. Maintainability of population structures for different types of personnel strategies (i.e. under control by promotion and control by recruitment) were extensively investigated for various types of Markov models (homogeneous as well as non-homogeneous) (Bartholomew, 1967; Vassiliou and Tsantas, 1984a). Semi-Markov models are extensions of Markov models that account for duration of stay in the states. Less attention is paid to the study of maintainability for semi-Markov models. Although, some interesting maintainability results were obtained for non-homogeneous semi-Markov models (Vassiliou and Papadopoulou, 1992). The current paper focuses on discrete-time homogeneous semi-Markov models, and explores the concept of maintainable population structures in this setting for a system with constant total size or one with a growth factor. In particular, a new concept of maintainability is introduced, the so called State Re-union maintainability (SR-maintainability). Moreover, we show that, under certain conditions, the seniority-based paths associated with the SR-maintainable structures converge. This allows to characterize the convex set of SR-maintainable structures.

math.PR

The Markov chain embedding problem in a low jump frequency context

We consider the problem of finding the transition rates of a continuous-time homogeneous Markov chain under the empirical condition that the state changes at most once during a time interval of unit length. It is proven that this conditional embedding approach results in a unique intensity matrix for a transition matrix with non-zero diagonal entries. Hence, the presented conditional embedding approach has the merit to avoid the identification phase as well as regularization for the embedding problem. The resulting intensity matrix is compared to the approximation for the Markov generator found by Jarrow.

math.PR

Cost-Sensitive Stacking: an Empirical Evaluation

Many real-world classification problems are cost-sensitive in nature, such that the misclassification costs vary between data instances. Cost-sensitive learning adapts classification algorithms to account for differences in misclassification costs. Stacking is an ensemble method that uses predictions from several classifiers as the training data for another classifier, which in turn makes the final classification decision. While a large body of empirical work exists where stacking is applied in various domains, very few of these works take the misclassification costs into account. In fact, there is no consensus in the literature as to what cost-sensitive stacking is. In this paper we perform extensive experiments with the aim of establishing what the appropriate setup for a cost-sensitive stacking ensemble is. Our experiments, conducted on twelve datasets from a number of application domains, using real, instance-dependent misclassification costs, show that for best performance, both levels of stacking require cost-sensitive classification decision.

cs.LG

To do or not to do: cost-sensitive causal decision-making

Causal classification models are adopted across a variety of operational business processes to predict the effect of a treatment on a categorical business outcome of interest depending on the process instance characteristics. This allows optimizing operational decision-making and selecting the optimal treatment to apply in each specific instance, with the aim of maximizing the positive outcome rate. While various powerful approaches have been presented in the literature for learning causal classification models, no formal framework has been elaborated for optimal decision-making based on the estimated individual treatment effects, given the cost of the various treatments and the benefit of the potential outcomes. In this article, we therefore extend upon the expected value framework and formally introduce a cost-sensitive decision boundary for double binary causal classification, which is a linear function of the estimated individual treatment effect, the positive outcome probability and the cost and benefit parameters of the problem setting. The boundary allows causally classifying instances in the positive and negative treatment class to maximize the expected causal profit, which is introduced as the objective at hand in cost-sensitive causal classification. We introduce the expected causal profit ranker which ranks instances for maximizing the expected causal profit at each possible threshold for causally classifying instances and differs from the conventional ranking approach based on the individual treatment effect. The proposed ranking approach is experimentally evaluated on synthetic and marketing campaign data sets. The results indicate that the presented ranking method effectively outperforms the cost-insensitive ranking approach and allows boosting profitability.

cs.LG