Searcharxiv⌕ Search

arXiv subjects

Brian Davis

Publications and source records attributed to Brian Davis.

29 records · Page 2Linked to original sources

Antichain Simplices

To each lattice simplex $Δ$ we associate a poset encoding the additive structure of lattice points in the fundamental parallelepiped for $Δ$. When this poset is an antichain, we say $Δ$ is antichain. To each partition $λ$ of $n$, we associate a lattice simplex $Δ_λ$ having one unimodular facet, and we investigate their associated posets. We give a number-theoretic characterization of the relations in these posets, as well as a simplified characterization in the case where each part of $λ$ is relatively prime to $n-1$. We use these characterizations to experimentally study $Δ_λ$ for all partitions of $n$ with $n\leq 73$. We also investigate the structure of these posets when $λ$ has only one or two distinct parts. Finally, we explain how this work relates to Poincaré series for the semigroup algebra associated to $Δ$, and we prove that this series is rational when $Δ$ is antichain.

math.CO↗

Language Model Supervision for Handwriting Recognition Model Adaptation

Training state-of-the-art offline handwriting recognition (HWR) models requires large labeled datasets, but unfortunately such datasets are not available in all languages and domains due to the high cost of manual labeling.We address this problem by showing how high resource languages can be leveraged to help train models for low resource languages.We propose a transfer learning methodology where we adapt HWR models trained on a source language to a target language that uses the same writing script.This methodology only requires labeled data in the source language, unlabeled data in the target language, and a language model of the target language. The language model is used in a bootstrapping fashion to refine predictions in the target language for use as ground truth in training the model.Using this approach we demonstrate improved transferability among French, English, and Spanish languages using both historical and modern handwriting datasets. In the best case, transferring with the proposed methodology results in character error rates nearly as good as full supervised training.

cs.CV↗

Predicting the Integer Decomposition Property via Machine Learning

In this paper we investigate the ability of a neural network to approximate algebraic properties associated to lattice simplices. In particular we attempt to predict the distribution of Hilbert basis elements in the fundamental parallelepiped, from which we detect the integer decomposition property (IDP). We give a gentle introduction to neural networks and discuss the results of this prediction method when scanning very large test sets for examples of IDP simplices.

math.CO↗

Semantic Relation Classification: Task Formalisation and Refinement

The identification of semantic relations between terms within texts is a fundamental task in Natural Language Processing which can support applications requiring a lightweight semantic interpretation model. Currently, semantic relation classification concentrates on relations which are evaluated over open-domain data. This work provides a critique on the set of abstract relations used for semantic relation classification with regard to their ability to express relationships between terms which are found in a domain-specific corpora. Based on this analysis, this work proposes an alternative semantic relation model based on reusing and extending the set of abstract relations present in the DOLCE ontology. The resulting set of relations is well grounded, allows to capture a wide range of relations and could thus be used as a foundation for automatic classification of semantic relations.

cs.CL↗

Semantic Relatedness for All (Languages): A Comparative Analysis of Multilingual Semantic Relatedness Using Machine Translation

This paper provides a comparative analysis of the performance of four state-of-the-art distributional semantic models (DSMs) over 11 languages, contrasting the native language-specific models with the use of machine translation over English-based DSMs. The experimental results show that there is a significant improvement (average of 16.7% for the Spearman correlation) by using state-of-the-art machine translation approaches. The results also show that the benefit of using the most informative corpus outweighs the possible errors introduced by the machine translation. For all languages, the combination of machine translation over the Word2Vec English distributional model provided the best results consistently (average Spearman correlation of 0.68).

cs.CL↗

Composite Semantic Relation Classification

Different semantic interpretation tasks such as text entailment and question answering require the classification of semantic relations between terms or entities within text. However, in most cases it is not possible to assign a direct semantic relation between entities/terms. This paper proposes an approach for composite semantic relation classification, extending the traditional semantic relation classification task. Different from existing approaches, which use machine learning models built over lexical and distributional word vector features, the proposed model uses the combination of a large commonsense knowledge base of binary relations, a distributional navigational algorithm and sequence classification to provide a solution for the composite semantic relation classification problem.

cs.CL↗

DINFRA: A One Stop Shop for Computing Multilingual Semantic Relatedness

This demonstration presents an infrastructure for computing multilingual semantic relatedness and correlation for twelve natural languages by using three distributional semantic models (DSMs). Our demonsrator - DInfra (Distributional Infrastructure) provides researchers and developers with a highly useful platform for processing large-scale corpora and conducting experiments with distributional semantics. We integrate several multilingual DSMs in our webservice so the end user can obtain a result without worrying about the complexities involved in building DSMs. Our webservice allows the users to have easy access to a wide range of comparisons of DSMs with different parameters. In addition, users can configure and access DSM parameters using an easy to use API.

cs.IR↗

Unlabeled Signed Graph Coloring

We extend the work of Hanlon on the chromatic polynomial of an unlabeled graph to define the unlabeled chromatic polynomial of an unlabeled signed graph. Explicit formulas are presented for labeled and unlabeled signed chromatic polynomials as summations over distinguished order-ideals of the signed partition lattice. We also define the quotient of a signed graph by a signed permutation, and show that its signed graphic arrangement is closely related to an induced arrangement on a distinguished subspace. Lastly, a formula for the number of unlabeled acyclic orientations of a signed graph is presented which recalls classical reciprocity theorems of Stanley and Zaslavsky.

math.CO↗

Complete intersection P-partition rings

We present an alternate proof of a result of Féray and Reiner characterizing posets whose $P$-partition rings are complete intersections. This shortened proof relates the complete intersection property to a simple structural property of a graph associated to $P$.

math.CO↗

PageNet: Page Boundary Extraction in Historical Handwritten Documents

When digitizing a document into an image, it is common to include a surrounding border region to visually indicate that the entire document is present in the image. However, this border should be removed prior to automated processing. In this work, we present a deep learning based system, PageNet, which identifies the main page region in an image in order to segment content from both textual and non-textual border noise. In PageNet, a Fully Convolutional Network obtains a pixel-wise segmentation which is post-processed into the output quadrilateral region. We evaluate PageNet on 4 collections of historical handwritten documents and obtain over 94% mean intersection over union on all datasets and approach human performance on 2 of these collections. Additionally, we show that PageNet can segment documents that are overlayed on top of other documents.

cs.CV↗

A Brief State of the Art for Ontology Authoring

One of the main challenges for building the Semantic web is Ontology Authoring. Controlled Natural Languages CNLs offer a user friendly means for non-experts to author ontologies. This paper provides a snapshot of the state-of-the-art for the core CNLs for ontology authoring and reviews their respective evaluations.

cs.CL↗