SearcharxivSearch

arXiv subjects

Amelia Taylor

Publications and source records attributed to Amelia Taylor.

At least 19 recordsLinked to original sources

mwBTFreddy: A Dataset for Flash Flood Damage Assessment in Urban Malawi

This paper describes the mwBTFreddy dataset, a resource developed to support flash flood damage assessment in urban Malawi, specifically focusing on the impacts of Cyclone Freddy in 2023. The dataset comprises paired pre- and post-disaster satellite images sourced from Google Earth Pro, accompanied by JSON files containing labelled building annotations with geographic coordinates and damage levels (no damage, minor, major, or destroyed). Developed by the Kuyesera AI Lab at the Malawi University of Business and Applied Sciences, this dataset is intended to facilitate the development of machine learning models tailored to building detection and damage classification in African urban contexts. It also supports flood damage visualisation and spatial analysis to inform decisions on relocation, infrastructure planning, and emergency response in climate-vulnerable regions.

cs.LG

Using Machine Learning to Detect Fraudulent SMSs in Chichewa

SMS enabled fraud is of great concern globally. Building classifiers based on machine learning for SMS fraud requires the use of suitable datasets for model training and validation. Most research has centred on the use of datasets of SMSs in English. This paper introduces a first dataset for SMS fraud detection in Chichewa, a major language in Africa, and reports on experiments with machine learning algorithms for classifying SMSs in Chichewa as fraud or non-fraud. We answer the broader research question of how feasible it is to develop machine learning classification models for Chichewa SMSs. To do that, we created three datasets. A small dataset of SMS in Chichewa was collected through primary research from a segment of the young population. We applied a label-preserving text transformations to increase its size. The enlarged dataset was translated into English using two approaches: human translation and machine translation. The Chichewa and the translated datasets were subjected to machine classification using random forest and logistic regression. Our findings indicate that both models achieved a promising accuracy of over 96% on the Chichewa dataset. There was a drop in performance when moving from the Chichewa to the translated dataset. This highlights the importance of data preprocessing, especially in multilingual or cross-lingual NLP tasks, and shows the challenges of relying on machine-translated text for training machine learning models. Our results underscore the importance of developing language specific models for SMS fraud detection to optimise accuracy and performance. Since most machine learning models require data preprocessing, it is essential to investigate the impact of the reliance on English-specific tools for data preprocessing.

cs.LG

MasakhaPOS: Part-of-Speech Tagging for Typologically Diverse African Languages

In this paper, we present MasakhaPOS, the largest part-of-speech (POS) dataset for 20 typologically diverse African languages. We discuss the challenges in annotating POS for these languages using the UD (universal dependencies) guidelines. We conducted extensive POS baseline experiments using conditional random field and several multilingual pre-trained language models. We applied various cross-lingual transfer models trained with data available in UD. Evaluating on the MasakhaPOS dataset, we show that choosing the best transfer language(s) in both single-source and multi-source setups greatly improves the POS tagging performance of the target languages, in particular when combined with cross-lingual parameter-efficient fine-tuning methods. Crucially, transferring knowledge from a language that matches the language family and morphosyntactic properties seems more effective for POS tagging in unseen languages.

cs.CL

MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity Recognition

African languages are spoken by over a billion people, but are underrepresented in NLP research and development. The challenges impeding progress include the limited availability of annotated datasets, as well as a lack of understanding of the settings where current methods are effective. In this paper, we make progress towards solutions for these challenges, focusing on the task of named entity recognition (NER). We create the largest human-annotated NER dataset for 20 African languages, and we study the behavior of state-of-the-art cross-lingual transfer methods in an Africa-centric setting, demonstrating that the choice of source language significantly affects performance. We show that choosing the best transfer language improves zero-shot F1 scores by an average of 14 points across 20 languages compared to using English. Our results highlight the need for benchmark datasets and models that cover typologically-diverse African languages.

cs.CL

AI4D -- African Language Program

Advances in speech and language technologies enable tools such as voice-search, text-to-speech, speech recognition and machine translation. These are however only available for high resource languages like English, French or Chinese. Without foundational digital resources for African languages, which are considered low-resource in the digital context, these advanced tools remain out of reach. This work details the AI4D - African Language Program, a 3-part project that 1) incentivised the crowd-sourcing, collection and curation of language datasets through an online quantitative and qualitative challenge, 2) supported research fellows for a period of 3-4 months to create datasets annotated for NLP tasks, and 3) hosted competitive Machine Learning challenges on the basis of these datasets. Key outcomes of the work so far include 1) the creation of 9+ open source, African language datasets annotated for a variety of ML tasks, and 2) the creation of baseline models for these datasets through hosting of competitive ML challenges.

cs.CL

Clique decompositions of multipartite graphs and completion of Latin squares

Our main result essentially reduces the problem of finding an edge-decomposition of a balanced r-partite graph of large minimum degree into r-cliques to the problem of finding a fractional r-clique decomposition or an approximate one. Together with very recent results of Bowditch and Dukes as well as Montgomery on fractional decompositions into triangles and cliques respectively, this gives the best known bounds on the minimum degree which ensures an edge-decomposition of an r-partite graph into r-cliques (subject to trivially necessary divisibility conditions). The case of triangles translates into the setting of partially completed Latin squares and more generally the case of r-cliques translates into the setting of partially completed mutually orthogonal Latin squares.

math.CO

Developing a statistically powerful measure for quartet tree inference using phylogenetic identities and Markov invariants

Recently there has been renewed interest in phylogenetic inference methods based on phylogenetic invariants, alongside the related Markov invariants. Broadly speaking, both these approaches give rise to polynomial functions of sequence site patterns that, in expectation value, either vanish for particular evolutionary trees (in the case of phylogenetic invariants) or have well understood transformation properties (in the case of Markov invariants). While both approaches have been valued for their intrinsic mathematical interest, it is not clear how they relate to each other, and to what extent they can be used as practical tools for inference of phylogenetic trees. In this paper, by focusing on the special case of binary sequence data and quartets of taxa, we are able to view these two different polynomial-based approaches within a common framework. To motivate the discussion, we present three desirable statistical properties that we argue any phylogenetic method should satisfy: (1) sensible behaviour under reordering of input sequences; (2) stability as the taxa evolve independently according to a Markov process; and (3) ability to detect if the conditions of a continuous-time process are violated. Motivated by these statistical properties, we develop and explore several new phylogenetic inference methods. In particular, we develop a statistical bias-corrected version of the Markov invariants approach which satisfies all three properties. We also extend previous work by showing that the phylogenetic invariants can be implemented in such a way as to satisfy property (3). A simulation study shows that, in comparison to other methods, our new proposed approach based on bias-corrected Markov invariants is extremely powerful for phylogenetic inference.

q-bio.QM

On the exact decomposition threshold for even cycles

A graph $G$ has a $C_k$-decomposition if its edge set can be partitioned into cycles of length $k$. We show that if $δ(G)\geq 2|G|/3-1$, then $G$ has a $C_4$-decomposition, and if $δ(G)\geq |G|/2$, then $G$ has a $C_{2k}$-decomposition, where $k\in \mathbb{N}$ and $k\geq 4$ (we assume $G$ is large and satisfies necessary divisibility conditions). These minimum degree bounds are best possible and provide exact versions of asymptotic results obtained by Barber, Kühn, Lo and Osthus. In the process, we obtain asymptotic versions of these results when $G$ is bipartite or satisfies certain expansion properties.

math.CO

Arbitrary Orientations of Hamilton Cycles in Digraphs

Let $n$ be sufficiently large and suppose that $G$ is a digraph on $n$ vertices where every vertex has in- and outdegree at least $n/2$. We show that $G$ contains every orientation of a Hamilton cycle except, possibly, the antidirected one. The antidirected case was settled by DeBiasio and Molla, where the threshold is $n/2+1$. Our result is best possible and improves on an approximate result by Häggkvist and Thomason.

math.CO

On the random greedy F-free hypergraph process

Let $F$ be a strictly $k$-balanced $k$-uniform hypergraph with $e(F)\geq |F|-k+1$ and maximum co-degree at least two. The random greedy $F$-free process constructs a maximal $F$-free hypergraph as follows. Consider a random ordering of the hyperedges of the complete $k$-uniform hypergraph $K_n^k$ on $n$ vertices. Start with the empty hypergraph on $n$ vertices. Successively consider the hyperedges $e$ of $K_n^k$ in the given ordering, and add $e$ to the existing hypergraph provided that $e$ does not create a copy of $F$. We show that asymptotically almost surely this process terminates at a hypergraph with $\tilde{O}(n^{k-(|F|-k)/(e(F)-1)})$ hyperedges. This is best possible up to logarithmic factors.

math.CO

The regularity method for graphs and digraphs

This MSci thesis surveys results in extremal graph theory, in particular relating to Hamilton cycles. Szeméredi's Regularity Lemma plays a central role. We also investigate the robust outexpansion property for digraphs. Kelly showed that every sufficiently large oriented graph on $n$ vertices with minimum in- and outdegree at least $3n/8 +o(n)$ contains any orientation of a Hamilton cycle. We use Kelly's arguments to extend his result to any robustly expanding digraph of linear degree.

math.CO

The Structure of N-Player Games when Influence and Independence Collide

We study the mathematical properties of probabilistic processes in which the independent actions of $n$ players (`causes') can influence the outcome of each player (`effects'). In such a setting, each pair of outcomes will generally be statistically correlated, even if the actions of all the players provide a complete causal description of the players' outcomes, and even if we condition on the outcome of any one player's action. This correlation always holds when $n=2$, but when $n=3$ there exists a highly symmetric process, recently studied, in which each cause can influence each effect, and yet each pair of effects is probabilistically independent (even upon conditioning on any one cause). We study such symmetric processes in more detail, obtaining a complete classification for all $n \geq 3$. Using a variety of mathematical techniques, we describe the geometry and topology of the underlying probability space that allows independence and influence to coexist.

math.PR

A semialgebraic description of the general Markov model on phylogenetic trees

Many of the stochastic models used in inference of phylogenetic trees from biological sequence data have polynomial parameterization maps. The image of such a map --- the collection of joint distributions for a model --- forms the model space. Since the parameterization is polynomial, the Zariski closure of the model space is an algebraic variety which is typically much larger than the model space, but has been usefully studied with algebraic methods. Of ultimate interest, however, is not the full variety, but only the model space. Here we develop complete semialgebraic descriptions of the model space arising from the k-state general Markov model on a tree, with slightly restricted parameters. Our approach depends upon both recently-formulated analogs of Cayley's hyperdeterminant, and the construction of certain quadratic forms from the joint distribution whose positive (semi-)definiteness encodes information about parameter values. We additionally investigate the use of Sturm sequences for obtaining similar results.

q-bio.PE

Minimal primes of ideals arising from conditional independence statements

We consider ideals arising in the context of conditional independence models that generalize the class of ideals considered by Fink [7] in a way distinct from the generalizations of Herzog-Hibi-Hreinsdottir-Kahle-Rauh [13] and Ay-Rauh [1]. We introduce switchable sets to give a combinatorial description of the minimal prime ideals, and for some classes we describe the minimal components. We discuss many possible interpretations of the ideals we study, including as 2 \times 2 minors of generic hypermatrices. We also introduce a definition of diagonal monomial orders on generic hypermatrices and we compute some Groebner bases.

math.AC

Second symmetric powers of chain complexes

We investigate Buchbaum and Eisenbud's construction of the second symmetric power S^2_R(X) of a chain complex X of modules over a commutative ring R. We state and prove a number of results from the folklore of the subject for which we know of no good direct references. We also provide several explicit computations and examples. We use this construction to prove the following version of a result of Avramov, Buchweitz, and Sega: Let R \to S be a module-finite ring homomorphism such that R is noetherian and local, and such that 2 is a unit in R. Let X be a complex of finite rank free S-modules such that X_n = 0 for each n < 0. If \cup_n Ass_R(H_n(X \otimes_S X)) \subseteq Ass(R) and if X_P \simeq S_P for each P \in Ass(R), then X \simeq S.

math.AC

Relations between semidualizing complexes

We study the following question: Given two semidualizing complexes B and C over a commutative noetherian ring R, does the vanishing of Ext^n_R(B,C) for n>>0 imply that B is C-reflexive? This question is a natural generalization of one studied by Avramov, Buchweitz, and Sega. We begin by providing conditions equivalent to B being C-reflexive, each of which is slightly stronger than the condition Ext^n_R(B,C)=0 for all n>>0. We introduce and investigate an equivalence relation \approx on the set of isomorphism classes of semidualizing complexes. This relation is defined in terms of a natural action of the derived Picard group and is well-suited for the study of semidualizing complexes over nonlocal rings. We identify numerous alternate characterizations of this relation, each of which includes the condition Ext^n_R(B,C)=0 for all n>>0. Finally, we answer our original question in some special cases.

math.AC

Borel Fixed Initial Ideals of Prime Ideals in Dimension Two

We prove that if the initial ideal of a prime ideal is Borel-fixed and the dimension of the quotient ring is less than or equal to two, then given any non-minimal associated prime ideal of the initial ideal it contains another associated prime ideal of dimension one larger.

math.AC

Methods for Computing Normalisations of Affine Rings

Our main purpose is to give multiple examples for using the available implementations for computing the normalization of an affine ring, computing the minimial generators of the normalization as an algebra over the original ring and integral closures of ideals. Some such examples have been published for Singular, but not for Macaulay 2 and we present both in this paper. We also briefly describe the implementations.

math.AC