SearcharxivSearch

arXiv subjects

Matthew Nguyen

Publications and source records attributed to Matthew Nguyen.

13 recordsLinked to original sources

On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness

Model capabilities have improved in large part due to scaling chain of thought. This has been a promising development for AI safety--where models verbalize their reasoning, it is possible to monitor it. However, in some cases, models do not verbalize important steps in their reasoning process. For example, models prompted with a cue suggesting the incorrect answer may fail to acknowledge that cue, even when it appears instrumental to their conclusion. When chain of thought (CoT) fails to disclose instrumental reasoning steps, we describe it as unfaithful. Prior work has shown that activation steering can be a useful method to improve faithfulness in CoT. We extend this line of work by studying how well steering for faithfulness generalizes across cue types, datasets, and methods of constructing the steering vector for three models (Gemma-3 4B, Qwen-3.5 9B, Gemma-3 12B) in a cued question-answering setting. While steering reliably increases cue acknowledgment for only the largest model (Gemma-3 12B), we find that when steering is effective, its effect generalizes broadly across cue types and datasets--in cross-cue and cross-dataset analyses, effect size is determined primarily by the evaluation setting, rather than the vector's train setting. How the vector is built also matters little--four construction methods, including one whose optimization target mentions no specific cue, yield similar effect sizes. Finally, we consider the possibility that steering promotes the salience of the cue and causes greater cue use, rather than targeting verbalization behaviors. However, we find no evidence for this--steering leaves the rate of cue use roughly unchanged while reducing hidden cue use, i.e., cue use that is not acknowledged.

cs.AI

Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations

Recent research has shown that large language models (LLMs) favor their own outputs when acting as judges, undermining the integrity of automated post-training and evaluation workflows. However, it is difficult to disentangle which behaviors are explained by narcissism versus experimental confounds. Specifically, LLM evaluators may deliver self-preferring verdicts when comparing responses to questions they fail on; these verdicts may not depend on the identity of the author, but on evaluator quality. We correct this by directly comparing the judge's voting distribution in cases where it evaluates itself versus another model. This evaluator quality baseline reveals that only 51% of examples in previous findings retain statistical significance against this null hypothesis, covering 89.6% of total self-preference probability mass. Finally, we compare the entropy of voting distributions, suggesting uncertainty-driven overlap, and show that our procedure enables more careful documentation against the backdrop of judge-bias research.

cs.CL

Minimal and Mechanistic Conditions for Behavioral Self-Awareness in LLMs

Recent studies have revealed that LLMs can exhibit behavioral self-awareness: the ability to accurately describe or predict their own learned behaviors without explicit supervision. This capability raises safety concerns as it may, for example, allow models to better conceal their true abilities during evaluation. We attempt to characterize the minimal conditions under which such self-awareness emerges, and the mechanistic processes through which it manifests. Through controlled finetuning experiments on instruction-tuned LLMs with low-rank adapters (LoRA), we find: (1) that self-awareness can be reliably induced using a single rank-1 LoRA adapter; (2) that the learned self-aware behavior can be largely captured by a single steering vector in activation space, recovering nearly all of the fine-tune's behavioral effect; and (3) that self-awareness is non-universal and domain-localized, with independent representations across tasks. Together, these findings suggest that behavioral self-awareness emerges as a domain-specific, linear feature that can be easily induced and modulated.

cs.CL

Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators

Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer from "self-preference bias": a tendency to favor their own outputs over those of other models. This bias undermines fairness and reliability in evaluation pipelines, particularly for tasks like preference tuning and model routing. We investigate whether lightweight steering vectors can mitigate this problem at inference time without retraining. We introduce a curated dataset that distinguishes self-preference bias into justified examples of self-preference and unjustified examples of self-preference, and we construct steering vectors using two methods: Contrastive Activation Addition (CAA) and an optimization-based approach. Our results show that steering vectors can reduce unjustified self-preference bias by up to 97\%, substantially outperforming prompting and direct preference optimization baselines. Yet steering vectors are unstable on legitimate self-preference and unbiased agreement, implying self-preference spans multiple or nonlinear directions. This underscores both their promise and limits as safeguards for LLM-as-judges and motivates more robust interventions.

cs.CL

Novel tools and observables for jet physics in heavy-ion collisions

Studies of fully-reconstructed jets in heavy-ion collisions aim at extracting thermodynamical and transport properties of hot and dense QCD matter. Recently, a plethora of new jet substructure observables have been theoretically and experimentally developed that provide novel precise insights on the modifications of the parton radiation pattern induced by a QCD medium. This report, summarizing the main lines of discussion at the 5th Heavy Ion Jet Workshop and CERN TH institute "Novel tools and observables for jet physics in heavy-ion collisions" in 2017, presents a first attempt at outlining a strategy for isolating and identifying the relevant physical processes that are responsible for the observed medium-induced jet modifications. These studies combine theory insights, based on the Lund parton splitting map, with sophisticated jet reconstruction techniques, including grooming and background subtraction algorithms.

hep-ph

b-jet Identification in PbPb Collisions with CMS

The flavor dependence of jet quenching is a powerful handle to discriminate between models of parton energy loss in heavy-ion collisions. We demonstrate the capacity of CMS to identify jets initiated by bottom quarks using displaced vertices reconstructed in the silicon tracking system. The b-jet to inclusive jet ratio is measured in PbPb collisions and compared to pp collisions and simulations at the same center-of-mass energy.

nucl-ex

Jet Reconstruction with Particle Flow in Heavy-Ion Collisions with CMS

In the particle-flow approach information from all available sub-detector systems is combined to reconstruct all stable particles. The global event reconstruction has been shown to improve, in particular, the resolution of jet energy and missing transverse energy in pp collisions compared to purely calorimetric measurements. This improvement is achieved primarily by combining the precise momentum determination of charged hadrons in the silicon tracker with the associated energy depositions in the calorimeters. By resolving individual particles inside jets, particle flow reduces the sensitivity of the jet energy scale to the jet fragmentation pattern, which is known to be one of the largest sources of systematic uncertainty in jet reconstruction. Particle flow reconstruction is thus potentially well-suited for the study of potential modifications to jet fragmentation in heavy-ion collisions. The particle flow algorithm has been adapted to the heavy-ion environment. The performance of jet reconstruction from particle flow objects in PbPb collisions using the anti-kT jet reconstruction algorithm is presented.

nucl-ex

Studies of Jet Quenching in PbPb collisions at CMS

Jets are an important tool to probe the hot, dense medium produced in ultra-relativistic heavy-ion collisions. At the collision energies available at the Large Hadron Collider (LHC), there is copious production of hard processes, such that high p_T jets may be differentiated from the heavy-ion underlying event. The multipurpose Compact Muon Solenoid (CMS) detector is well designed to measure hard scattering processes with its high quality calorimeters and high precision silicon tracker. Jet quenching has been studied in CMS in PbPb collisions at sqrt(s_NN)= 2.76 TeV. As a function of centrality, dijet events with a high p_T leading jet were found to have an increasing momentum imbalance that was significantly larger than predicted by simulations. The angular distribution of jet fragmentation products has been explored by associating charged tracks with the jets measured in the calorimeters. By projecting the momenta of charged tracks onto the leading jet axis it is shown that the apparent momentum imbalance of the leading dijet pair can be recovered if low p_T tracks are considered. A large fraction of the balancing momentum carried by these soft particles is radiated at large angle relative to the jets.

nucl-ex

Jet Fragmentation in Vacuum and Medium with gamma-hadron Correlations in PHENIX

Jet fragmentation in p+p and Au+Au collisions is studied via back-to-back correlations of direct photons and charged hadrons. The direct photon correlations are obtained by statical subtraction of the background from decay photons. Results on the nuclear modification to the associated charged hadron yields are reviewed. Further studies of jet fragmentation in p+p using isolated direct photons are also presented. A kT-smeared LO pQCD calculation is used to interpret the data. The sensitivity of the data to the underlying fragmentation function is tested and the results are found to be compatible with expectations of a sample dominated by quark jet fragmentation.

nucl-ex

Jet Fragmentation in Medium and Vacuum with the PHENIX Detector

One of the most active areas of investigation in relativistic heavy-ion collisions is the study of the jet quenching phenomenon whereby hard partons lose their energy as they traverse the hot, dense matter created in such collisions. Strong parton energy loss has been observed in central nucleus-nucleus collisions as evidenced by the a large suppression of the yield of high pT hadrons as compared to the expected yield based on measurements in p+p collisions. Moreover, measurements of back-to-back correlations of charged hadrons suggest that jet shapes are strongly modified modified by the medium. The quantitative interpretation of single and di-hadron measurements is, however, complicated by the fact that the initial parton energy is unknown. A more informative measurement would be one in which the initial parton energy is known, allowing the determination of the fragmentation function, which may be effectively modified from its vacuum form by the presence of the medium. Two measurements in which the initial parton energy may be estimated are discussed in these proceedings: jet reconstruction and two- particle correlations using direct photons. Jet reconstruction in nuclear collisions is challenging due to the large background of soft particles, fluctuations of which give rise to fake jets. Direct photons can be used to estimate the initial parton energy of the recoil jet without recourse to jet reconstruction algorithms. However, such studies suffer from a smaller rate and the direct photon signal must be disentangled from a large background of decay photons. We present jet reconstruction results which use an algorithm suitable for a high multiplicity environment. We also present results of two-particle correlations using direct photons. These results are discussed in the context of medium modification to the fragmentation function.

nucl-ex

Aspects of Jet Production with PHENIX

Measurement of the in-medium energy loss of fast partons is one of the most active topics in heavy-ion physics. Such studies provide an opportunity to gain insight into the fundamental behavior of QCD processes by studying them away from vacuum conditions. A promising channel for relating theoretical models to data are two particle correlations using direct photon triggers. Recent results on this observeable using the PHENIX detector are presented.

nucl-ex

High p_T Direct Photon-Hadron Correlations Using the PHENIX Detector

Jet tomography, the study of differential energy loss of hard scattered partons to infer the density profile of the medium, is greatly improved by precise knowledge of the initial energy of the hard probe. As photons are not strongly interacting, the momentum of the recoil jet from a direct photon trigger is balanced, to a good approximation, by the momentum of the photon. The energy loss of the away-side jet may be viewed as an effective modification of the fragmentation function. Direct photon-hadron correlations in A+A collisions should be sensitive to modified jet fragmentation as well as to medium response effects. Complementary measurements from p+p collisions are necessary to benchmark jet fragmentation expectations at \sqrt{s_NN} = 200 GeV as well as to constrain perturbative calculations in the γ+jet channel. Here we present new results from p+p and Au+Au collisions which use a statistical method to subtract the background from decay photons.

nucl-ex

Resonances and Quantum Scattering for the Morse Potential as a Barrier

Quantum scattering in the presence of a potential valley followed by a barrier is examined for the case of a Morse potential, for which exact analytic solutions to the Schr\UNICODE{0xf6}dinger equation are known in terms of confluent hypergeometric functions. For our application the potential is characterized by three parameters: the height of the barrier, the distance of the barrier from the origin of the radial variable $r,$ and a diffuseness parameter. The wave function, defined in the interval $0\leq r<\infty ,$ is required to vanish at $r=0,$ and hence represents a radial partial wave for zero angular momentum. The vanishing at $r=0$ requires a special combination of hypergeometric functions, and can lead to resonances for incident energies which occur below the top of the barrier. Numerical values for the analytical phase shifts are presented in and outside the resonant regions, and the corresponding properties of the scattering S-matrix are examined in the complex momentum plane, mainly for pedagogical reasons. The validity of the Breit-Wigner approximation to the resonant phase shifts is tested, and the motion of a ''resonant'' wave packet slowly leaking out of the valley region is also displayed.

nucl-th