SearcharxivSearch

arXiv subjects

Brian Rabern

Publications and source records attributed to Brian Rabern.

4 recordsLinked to original sources

LogicSkills: A Structured Benchmark for Formal Reasoning in Large Language Models

Large language models perform well on many logical reasoning benchmarks, but it remains unclear which core logical skills they truly master. To address this, we introduce LogicSkills, a benchmark that isolates three fundamental logical skills: (i) $\textit{formal symbolization}\unicode{x2014}{}$translating premises into first-order logic; (ii) $\textit{countermodel construction}\unicode{x2014}$showing that an argument is logically invalid by constructing a finite countermodel; and (iii) $\textit{validity assessment}\unicode{x2014}$determining whether a conclusion follows from a set of premises. Items are drawn from the two-variable fragment of first-order logic without identity and are presented in both English and a Carrollian nonce-word language. All instances are solver-verified with Z3 for correctness and non-triviality. Across conventional instruction-tuned LLMs, performance is high on $\textit{validity assessment}$ but substantially lower on $\textit{formal symbolization}$ and $\textit{countermodel construction}$, highlighting that high task-level accuracy can mask weaknesses in core logical skills. In contrast, recent reasoning-tuned models perform strongly across all three tasks, suggesting a more systematic logical skill profile.

cs.AI

Playing cards with Vizing's demon

We analyze a solitaire game in which a demon rearranges some cards after each move. The graph edge coloring theorems of K\H{o}nig (1931) and Vizing (1964) follow from the winning strategies developed.

math.CO

Structural fixed-point theorems

The semantic paradoxes are associated with self-reference or referential circularity. However, there are infinitary versions of the paradoxes, such as Yablo's paradox, that do not involve this form of circularity. It remains an open question what relations of reference between collections of sentences afford the structure necessary for paradoxicality -- these are the so-called "dangerous" directed graphs. Building on Rabern, et. al (2013) we reformulate this problem in terms of fixed points of certain functions, thereby boiling it down to get a purely mathematical problem.

math.CO

A Novel Proof of the Heine-Borel Theorem

Every beginning real analysis student learns the classic Heine-Borel theorem, that the interval [0,1] is compact. In this article, we present a proof of this result that doesn't involve the standard techniques such as constructing a sequence and appealing to the completeness of the reals. We put a metric on the space of infinite binary sequences and prove that compactness of this space follows from a simple combinatorial lemma. The Heine-Borel theorem is an immediate corollary.

math.HO