Searcharxiv⌕ Search

arXiv subjects

Paul B. Kantor

Publications and source records attributed to Paul B. Kantor.

8 recordsLinked to original sources

Interpretable GOHR Agents via Sparse Autoencoders

A central challenge in interpreting learned decision-making systems is to determine whether their internal representations contain concepts that help explain their behavior. We report interpretability experiments for a tokenized autoregressive Transformer agent in the Game of Hidden Rules (GOHR). We focus on a compact two-rule task in which both hidden rules map object shapes to target buckets, but with different permutations. The policy is trained on episodes sampled from these two hidden rules and then evaluated with fixed weights. It is never given a rule label and does not use an explicit rule classifier; any rule information must be inferred implicitly from interaction history. In this setting, the correct rule is not identifiable before the agent tries an informative move and observes accept/reject feedback. Sparse autoencoders (SAEs) trained on the agent's decision-token embeddings recover this structure. When held-out decisions are labeled by simple concepts such as the chosen shape or bucket, SAE dimensions that are highly selective for a concept cover most decisions where that concept is present. Individual SAE dimensions also correspond to interpretable strategies such as probing one rule hypothesis and switching after negative feedback.

cs.LG↗

AI Learning and Conceptual Transfer in the Game of Hidden Rules

This report summarizes the work conducted on the Game of Hidden Rules (GOHR), focusing on reinforcement learning agents trained to infer hidden rules from trial-and-error feedback, representation design, rule difficulty analysis, transfer learning, generalization, and pseudo-bot-assisted human learning analysis. The report focuses on the Transformer-based A2C framework, Feature-Centric and Object-Centric representations, experimental findings, and classification of human learning data.

cs.AI↗

Toward a Metrology for Artificial Intelligence: Hidden-Rule Environments and Reinforcement Learning

We investigate reinforcement learning in the Game Of Hidden Rules (GOHR) environment, a complex puzzle in which an agent must infer and execute hidden rules to clear a 6$\times$6 board by placing game pieces into buckets. We explore two state representation strategies, namely Feature-Centric (FC) and Object-Centric (OC), and employ a Transformer-based Advantage Actor-Critic (A2C) algorithm for training. The agent has access only to partial observations and must simultaneously infer the governing rule and learn the optimal policy through experience. We evaluate our models across multiple rule-based and trial-list-based experimental setups, analyzing transfer effects and the impact of representation on learning efficiency.

cs.LG↗

Randomization to Reduce Terror Threats at Large Venues

Can randomness be better than scheduled practices, for securing an event at a large venue such as a stadium or entertainment arena? Perhaps surprisingly, from several perspectives the answer is "yes." This note examines findings from an extensive study of the problem, including interviews and a survey of selected venue security directors. That research indicates that: randomness has several goals; many security directors recognize its potential; but very few have used it much, if at all. Some fear they will not be able to defend using random methods if an adversary does slip through security. Others are concerned that staff may not be able to perform effectively. We discuss ways in which it appears that randomness can improve effectiveness, ways it can be effectively justified to those who must approve security processes, and some potential research or regulatory advances.

cs.CY↗

Confidence Assertions in Cyber-Security for an Information-Sharing Environment

Information sharing is vital in resisting cyberattacks, and the volume and severity of these attacks is increasing very rapidly. Therefore responders must triage incoming warnings in deciding how to act. This study asked a very specific question: "how can the addition of confidence information to alerts and warnings improve overall resistance to cyberattacks." We sought, in particular, to identify current practices, and if possible, to identify some "best practices." The research involved literature review and interviews with subject matter experts at every level from system administrators to persons who develop broad principles of policy. An innovative Modified Online Delphi Panel technique was used to elicit judgments and recommendations from experts who were able to speak with each other and vote anonymously to rank proposed practices.

cs.CR↗

Soft Triangles for Expert Aggregation

We consider the problem of eliciting expert assessments of an uncertain parameter. The context is risk control, where there are, in fact, three uncertain parameters to be estimates. Two of these are probabilities, requiring the that the experts be guided in the concept of "uncertainty about uncertainty." We propose a novel formulation for expert estimates, which relies on the range and the median, rather than the variance and the mean. We discuss the process of elicitation, and provide precise formulas for these new distributions.

cs.AI↗

Can We Distinguish Machine Learning from Human Learning?

What makes a task relatively more or less difficult for a machine compared to a human? Much AI/ML research has focused on expanding the range of tasks that machines can do, with a focus on whether machines can beat humans. Allowing for differences in scale, we can seek interesting (anomalous) pairs of tasks T, T'. We define interesting in this way: The "harder to learn" relation is reversed when comparing human intelligence (HI) to AI. While humans seems to be able to understand problems by formulating rules, ML using neural networks does not rely on constructing rules. We discuss a novel approach where the challenge is to "perform well under rules that have been created by human beings." We suggest that this provides a rigorous and precise pathway for understanding the difference between the two kinds of learning. Specifically, we suggest a large and extensible class of learning tasks, formulated as learning under rules. With these tasks, both the AI and HI will be studied with rigor and precision. The immediate goal is to find interesting groundtruth rule pairs. In the long term, the goal will be to understand, in a generalizable way, what distinguishes interesting pairs from ordinary pairs, and to define saliency behind interesting pairs. This may open new ways of thinking about AI, and provide unexpected insights into human learning.

cs.LG↗

Quantum Message Disruption: A Two-State Model

A game in which one player makes unitary transformations of a simple system, and another seeks to confound the resulting state by a randomly chosen action is analyzed carefully. It is shown that the second player can reduce any system to a completely random one by rotation through an angle of 120 degrees, about an axis chosen at random. If, on the other hand, the second player is forced to behave ``classically'' by reducing the wave function, then the first play retains an advantage, which the second player may eliminate by repeated measurement using randomly selected bases.

quant-ph↗