SearcharxivSearch

arXiv subjects

Marcin Abram

Publications and source records attributed to Marcin Abram.

15 recordsLinked to original sources

Emergent Inference-Time Semantic Contamination via In-Context Priming

Recent work has shown that fine-tuning large language models (LLMs) on insecure code or culturally loaded numeric codes can induce emergent misalignment, causing models to produce harmful content in unrelated downstream tasks. The authors of that work concluded that $k$-shot prompting alone does not induce this effect. We revisit this conclusion and show that inference-time semantic drift is real and measurable; however, it requires models of large-enough capability. Using a controlled experiment in which five culturally loaded numbers are injected as few-shot demonstrations before a semantically unrelated prompt, we find that models with richer cultural-associative representations exhibit significant distributional shifts toward darker, authoritarian, and stigmatized themes, while a simpler/smaller model does not. We additionally find that structurally inert demonstrations (nonsense strings) perturb output distributions, suggesting two separable mechanisms: structural format contamination and semantic content contamination. Our results map the boundary conditions under which inference-time contamination occurs, and carry direct implications for the security of LLM-based applications that use few-shot prompting.

cs.CL

Toward Evaluation Frameworks for Multi-Agent Scientific AI Systems

We analyze the challenges of benchmarking scientific (multi)-agentic systems, including the difficulty of distinguishing reasoning from retrieval, the risks of data/model contamination, the lack of reliable ground truth for novel research problems, the complications introduced by tool use, and the replication challenges due to the continuously changing/updating knowledge base. We discuss strategies for constructing contamination-resistant problems, generating scalable families of tasks, and the need for evaluating systems through multi-turn interactions that better reflect real scientific practice. As an early feasibility test, we demonstrate how to construct a dataset of novel research ideas to test the out-of-sample performance of our system. We also discuss the results of interviews with several researchers and engineers working in quantum science. Through those interviews, we examine how scientists expect to interact with AI systems and how these expectations should shape evaluation methods.

cs.CY

Hofstadter Butterflies in Topological Insulators

In this chapter, we investigate the energy spectra as well as the bulk and surface states in a two-dimensional system composed of a coupled stack of one-dimensional dimerized chains in the presence of an external magnetic field. Specifically, we analyze the Hofstadter butterfly patterns that emerge in a 2D stack of coupled 1D Su-Schrieffer-Heeger (SSH) chains subject to an external transverse magnetic field. Depending on the parameter regime, we find that the energy spectra of this hybrid topological system can exhibit topologically non-trivial bulk bands separated by energy gaps. Upon introducing boundaries into the system, we observe topologically protected in-gap surface states, which are protected either by a non-trivial Chern number or by inversion symmetry. We examine the resilience of these surface states against perturbations, confirming their expected stability against local symmetry-preserving perturbations.

cond-mat.mes-hall

Why does in-context learning fail sometimes? Evaluating in-context learning on open and closed questions

We measure the performance of in-context learning as a function of task novelty and difficulty for open and closed questions. For that purpose, we created a novel benchmark consisting of hard scientific questions, each paired with a context of various relevancy. We show that counter-intuitively, a context that is more aligned with the topic does not always help more than a less relevant context. This effect is especially visible for open questions and questions of high difficulty or novelty. This result reveals a fundamental difference between the treatment of close-form and open-form questions by large-language models and shows a need for a more robust evaluation of in-context learning on the variety of different types of questions. It also poses a new question of how to optimally select a context for large language models, especially in the context of Retrieval Augmented Generation (RAG) systems. Our results suggest that the answer to this question can be highly application-dependent and might be contingent on factors including the format of the question, the perceived difficulty level of the questions, and the novelty or popularity of the information we seek.

cs.CL

Context Matters: Data-Efficient Augmentation of Large Language Models for Scientific Applications

In this paper, we explore the challenges inherent to Large Language Models (LLMs) like GPT-4, particularly their propensity for hallucinations, logic mistakes, and incorrect conclusions when tasked with answering complex questions. The capacity of LLMs to present erroneous answers in a coherent and semantically rigorous manner further complicates the detection of factual inaccuracies. This issue is especially pronounced in fields that require specialized expertise. Our work delves into these challenges, aiming to enhance the understanding and mitigation of such errors, thereby contributing to the improvement of LLM accuracy and reliability in scientific and other specialized domains. Our findings reveal a non-linear relationship between the context's relevancy and the answers' measured quality. In addition, we demonstrate that with the correct calibration, it is possible to automate the grading procedure -- a finding suggesting that, at least to some degree, the LLMs can be used to self-examine the quality of their own performance. Finally, we describe an experimental platform that can be seen as a proof-of-concept of the techniques described in this work.

cs.CL

Inferring topological transitions in pattern-forming processes with self-supervised learning

The identification and classification of transitions in topological and microstructural regimes in pattern-forming processes are critical for understanding and fabricating microstructurally precise novel materials in many application domains. Unfortunately, relevant microstructure transitions may depend on process parameters in subtle and complex ways that are not captured by the classic theory of phase transition. While supervised machine learning methods may be useful for identifying transition regimes, they need labels which require prior knowledge of order parameters or relevant structures describing these transitions. Motivated by the universality principle for dynamical systems, we instead use a self-supervised approach to solve the inverse problem of predicting process parameters from observed microstructures using neural networks. This approach does not require predefined, labeled data about the different classes of microstructural patterns or about the target task of predicting microstructure transitions. We show that the difficulty of performing the inverse-problem prediction task is related to the goal of discovering microstructure regimes, because qualitative changes in microstructural patterns correspond to changes in uncertainty predictions for our self-supervised problem. We demonstrate the value of our approach by automatically discovering transitions in microstructural regimes in two distinct pattern-forming processes: the spinodal decomposition of a two-phase mixture and the formation of concentration modulations of binary alloys during physical vapor deposition of thin films. This approach opens a promising path forward for discovering and understanding unseen or hard-to-discern transition regimes, and ultimately for controlling complex pattern-forming processes.

cond-mat.mtrl-sci

Performance Weighting for Robust Federated Learning Against Corrupted Sources

Federated Learning has emerged as a dominant computational paradigm for distributed machine learning. Its unique data privacy properties allow us to collaboratively train models while offering participating clients certain privacy-preserving guarantees. However, in real-world applications, a federated environment may consist of a mixture of benevolent and malicious clients, with the latter aiming to corrupt and degrade federated model's performance. Different corruption schemes may be applied such as model poisoning and data corruption. Here, we focus on the latter, the susceptibility of federated learning to various data corruption attacks. We show that the standard global aggregation scheme of local weights is inefficient in the presence of corrupted clients. To mitigate this problem, we propose a class of task-oriented performance-based methods computed over a distributed validation dataset with the goal to detect and mitigate corrupted clients. Specifically, we construct a robust weight aggregation scheme based on geometric mean and demonstrate its effectiveness under random label shuffling and targeted label flipping attacks.

cs.LG

Emulating Quantum Dynamics with Neural Networks via Knowledge Distillation

High-fidelity quantum dynamics emulators can be used to predict the time evolution of complex physical systems. Here, we introduce an efficient training framework for constructing machine learning-based emulators. Our approach is based on the idea of knowledge distillation and uses elements of curriculum learning. It works by constructing a set of simple, but rich-in-physics training examples (a curriculum). These examples are used by the emulator to learn the general rules describing the time evolution of a quantum system (knowledge distillation). The goal is not only to obtain high-quality predictions, but also to examine the process of how the emulator learns the physics of the underlying problem. This allows us to discover new facts about the physical system, detect symmetries, and measure relative importance of the contributing physical processes. We illustrate this approach by training an artificial neural network to predict the time evolution of quantum wave packages propagating through a potential landscape. We focus on the question of how the emulator learns the rules of quantum dynamics from the curriculum of simple training examples and to which extent it can generalize the acquired knowledge to solve more challenging cases.

quant-ph

Democratising blockchain: A minimal agency consensus model

We propose a novel consensus protocol based on a hybrid approach, that combines a directed acyclic graph (DAG) and a classical chain of blocks. This architecture allows us to enforce collective block construction, minimising the monopolistic power of the round-leader. In this way, we decrease the possibility for collusion among senders and miners, as well as miners themselves, allowing the use of more incentive compatible and fair pricing strategies. We investigate these possibilities alongside the ability to use the DAG structure to minimise the risk of transaction censoring. We conclude by providing preliminary benchmarks of our protocol and by exploring further research directions.

cs.CR

Antiferromagnetism, charge density wave, and d-wave superconductivity in the extended $t$--$J$--$U$ model: role of intersite Coulomb interaction and a critical overview of renormalized mean field theory

In the first part of the paper, we study the stability of antiferromagnetic (AF), charge density wave (CDW), and superconducting (SC) states within the $t$-$J$-$U$-$V$ model of strongly correlated electrons by using the statistically consistent Gutzwiller approximation (SGA). We concentrate on the role of the intersite Coulomb interaction term $V$ in stabilizing the CDW phase. In particular, we show that the charge ordering appears only above a critical value of $V$ in a limited hole-doping range $δ$. The effect of the $V$ term on SC and AF phases is that a strong interaction suppresses SC, whereas the AF order is not significantly influenced by its presence. In the second part, separate calculations for the case of pure SC phase have been carried out within an extended approach (the diagrammatic expansion for the Gutzwiller wave function, DE-GWF) in order to analyze the influence of the intersite Coulomb repulsion on the SC phase with the higher-order corrections included beyond the SGA method. In the Appendices we discuss the ambiguity connected with the choice of the Gutzwiller renormalization factors within the renormalized mean filed theory when either AF or CDW orders are considered. At the end we overview briefly the possible extensions of the current models to make description of the SC, AF, and CDW states on equal footing.

cond-mat.str-el

Tricritical wings in UGe$_2$: A microscopic interpretation

In the present work we analyze the second order transition line that connect the tricritical point and the quantum critical ending point on the temperature--magnetic-field plane in UGe$_2$. For the microscopic modeling we employ the Anderson lattice model recently shown to provide a fairly complete description of the full magnetic phase diagram of UGe$_2$ including all the criticalities. The shape of the so-called tricritical wings, i.e. surfaces of the first-order transitions, previously reported by us to quantitatively agree with the experimental data, is investigated here with respect to the change of the total filling and the Landé factor for $f$ electrons which can differ from the free electron value. The analysis of the total filling dependence demonstrates sensitivity of our prediction when the respective positions of the critical ending point at the metamagnetic transition and tricritical point are mismatched as compared to the experiment.

cond-mat.str-el

Criticalities in the itinerant ferromagnet UGe$_{2}$

We provide a microscopic description of the magnetic properties of UGe$_2$ and in particular, of its both classical and quantum critical behavior. Namely, we account for all the critical points: the critical ending point (CEP) at the metamagnetic phase transition, the tricritical point, and the quantum critical end point at the ferromagnetic to paramagnetic phase transition. Their position agrees quantitatively with experiment. Additionally, we predict that the metamagnetic CEP can be traced down to zero temperature and becomes quantum critical point by a small decrease of both the total electron concentration and the external pressure. The system properties are then determined by the quantum critical fluctuations appearing near the instability point of the Fermi surface topology.

cond-mat.str-el

t-t'-J-U model in mean-field approximation: Coexistence of superconductivity and antiferromagnetism

We discuss the $t$-$J$-$U$ model in the mean-field approximation. The role of spin-exchange coupling $J$ and the second nearest hopping $t'$ are examined in the context of the coexistence of superconductivity (SC) and antiferromagnetism (AF). Stability of the phases is studied with respect to temperature. The coexistence region exists for the sufficiently large Coulomb repulsion ($U>U_{cr}$), and in the vicinity of the half-filled band (hole doping $δ< δ_{cr}$). The critical hole doping is relatively small ($δ_{cr} \approx 0.006$ for $J/|t| = 1/3$) and linear with respect to $J$. The decrease of $U_{cr}$ is proportional to $J$, except the limit of small $J$ ($J/|t|< 0.03)$, where $U_{cr}$ grows rapidly with decreasing $J$. The effect of the second nearest hopping is limited -- the phase diagram does not change in a qualitative manner when the $t'$ value is changed. In the limit of $T \rightarrow 0$, SC phase is stable even for large hole-doping (such as $δ= 0.5$). Additional paramagnetic (PM) phase appears for large $δ$ or small $U$ at non-zero temperature. When temperature increases, both SC and AF+SC phase regions are reduced.

cond-mat.supr-con

Ferromagnetism in UGe2 : A microscopic model

Anderson lattice model is used to rationalize the principal features of the heavy fermion compound UGe2 by means of the generalized Gutzwiller approach (the SGA method). This microscopic approach successfully reproduces magnetic and electronic properties of this material, in a qualitative agreement with experimental findings from the magnetization measurements, the neutron scattering, and the de Haas-van Alphen oscillations. Most importantly, it explains the appearance, sequence, character, and evolution in an applied magnetic field of the observed in UGe2 ferro- and, para-magnetic phases as an effect of a competition between the f-f electrons Coulomb interaction energy and f-conduction electrons kinetic energy (hybridization)

cond-mat.str-el

d-wave superconductivity and its coexistence with antiferromagnetism in t-J-U model: Statistically consistent Gutzwiller approach

We discuss the coexistence of antiferromagnetism and d-wave superconductivity within the so-called statistically-consistent Gutzwiller approximation (SGA) applied to the t-J-U model. In this approach, the averages calculated in a self-consistent manner coincide with those determined variationally. Such consistency is not guaranteed within the standard renormalized mean field theory. With the help of SGA, we show that for the typical value J/|t| = 1/3, coexistence of antiferromagnetism (AF) and superconductivity (SC) appears only for U/|t| > 10.6 and in a very narrow range of doping (δ< 0.006) in the vicinity of the Mott insulating state, in contrast to some previous reports. In the coexistent AF+SC phase, a staggered spin-triplet component of the superconducting gap appears also naturally; its value is very small.

cond-mat.supr-con