SearcharxivSearch

arXiv subjects

Kevin Liu

Publications and source records attributed to Kevin Liu.

At least 19 recordsLinked to original sources

Evaluation of the Exradin A30 Parallel Plate Ion Chamber as a Reference Dosimeter in Ultra-High Dose Rate (UHDR) Electron Beams

Reliable reference dosimetry for ultra-high dose-rate (UHDR) beams (>40 Gy/s) is challenging because conventional ionization chambers (ICs) exhibit saturation from ion recombination. The Exradin A30 IC uses an ultra-thin 0.3-mm electrode spacing to improve charge-collection efficiency (CCE). This study evaluated the commercial A30 as a reference dosimeter for UHDR electron beams by characterizing leakage current, CCE, polarity correction (Ppol), and beam-quality correction factors (kQ). Measurements were performed with a 9-MeV IntraOp Mobetron from the accelerator head, achieving up to 9 Gy per pulse (DPP) and an instantaneous dose rate of 2.25 MGy/s. Data were acquired in grounded water-equivalent plastic, distilled water, and saline water. DPP was varied by changing SSD at a fixed 4-{\mu}s pulse width, while pulse repetition frequency (PRF) ranged from 5 to 90 Hz. CCE was determined using EBT-XD film under matched UHDR and conventional dose and energy conditions. CCE and Ppol were also evaluated as functions of DPP and PRF in distilled and saline water. Values of kQ were calculated using Monte Carlo simulations and measured in TrueBeam electron beams. Leakage current was <2 fA. Both CCE and Ppol decreased with increasing DPP; however, CCE remained 90-99% across all three phantoms, while Ppol decreased from 0.990 to 0.981 in liquid and solid water. Neither CCE nor Ppol depended on PRF over 5-90 Hz. Measured and calculated kQ values agreed within 0.8% at all energies except 9 MeV, where they differed by 2%. The A30 exhibited 5% recombination at DPP up to 5 Gy in distilled and saline water. Its response in solid phantoms was affected by charge buildup, which was mitigated by grounding. With appropriate CCE corrections and grounded solid phantoms, the commercial A30 is suitable for reference dosimetry in UHDR electron beams.

physics.med-ph

A ChatGPT-assisted Triangle Characterization of Affine Permutation Inversion Graphs

Inversion sets of permutations in the affine symmetric group $\widetilde{S}_n$ were studied extensively by Bj\"orner and Brenti. One of their methods for encoding an inversion set is through an affine inversion graph, which is a certain weighted graph on vertex set $[n]=\{1,2,\ldots,n\}$. Subsequent work by Papi characterized which graphs arise as affine inversion graphs. In this paper, we provide an alternative characterization in terms of a simple local condition on each triangle in a weighted tournament graph. This new characterization was produced with the assistance of ChatGPT, which suggested several key insights that simplified portions of Papi's original characterization. Consequences of our characterization include efficient algorithms for recognizing inversion graphs and inversion sets. Furthermore, we give bounds on the weights along directed paths, and we show that standardizing the labels on an induced subgraph results in another inversion graph. We conclude with a new order $O(|R|+n^{3})$ algorithm for testing if a given set $R$ is the inversion set of an affine permutation.

math.CO

Commutation classes of reduced words and higher Bruhat orders for affine permutations

The higher Bruhat orders are partial orders that generalize the weak order on the symmetric group $S_n$, and the second higher Bruhat order is a poset on commutation classes of reduced words for the longest element in $S_n$, where covering relations correspond to braid relations. Constructing analogs in other settings is an area of recent interest, and we present an analog that generalizes any interval $[id,w]$ in the weak order of both the symmetric group and the affine symmetric group. Paralleling the classical case, we show that the second higher Bruhat order is a poset on commutation classes of reduced words for any affine permutation. For the symmetric group, we also establish results for all higher Bruhat orders that are direct analogs of those in the classical case.

math.CO

Descent sets of cyclic permutations in types B and D

Elizalde constructed a bijection $\phi$ from the cyclic permutations $\pi\in S_{n+1}$ to the symmetric group $S_n$ satisfying $\operatorname{Des}(\pi)\cap \{1,2,\ldots,n-1\}=\operatorname{Des}(\phi(\pi))$. We give a corresponding result on the signed symmetric group $B_n$ by constructing a function $\Phi$ from the cyclic signed permutations $\pi\in B_{n+1}$ to $B_n$ satisfying $\operatorname{Des}(\pi)\cap \{0,1,\ldots,n-1\}=\operatorname{Des}(\Phi(\pi))$. Moreover, letting $D_{n+1}\subseteq B_{n+1}$ be the subgroup consisting of signed permutations with an even number of sign changes, we show that the restriction of $\Phi$ to the cyclic signed permutations in $D_{n+1}$ or its complement is a bijection. Our function $\Phi$ reduces to Elizalde's original bijection $\phi$ under the natural identification of the symmetric groups as subgroups of the signed symmetric groups. One application of our results is asymptotic normality of the descent and flag major index statistics on the cyclic signed permutations in $B_{n}$ and $D_n$.

math.CO

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation

Internal world models (WMs) enable agents to understand the world's state and predict transitions, serving as the basis for advanced deliberative reasoning. Recent large Vision-Language Models (VLMs), such as OpenAI o3, GPT-4o and Gemini, exhibit potential as general-purpose WMs. While the latest studies have evaluated and shown limitations in specific capabilities such as visual understanding, a systematic evaluation of VLMs' fundamental WM abilities remains absent. Drawing on comparative psychology and cognitive science, we propose a two-stage framework that assesses Perception (visual, spatial, temporal, quantitative, and motion) and Prediction (mechanistic simulation, transitive inference, compositional inference) to provide an atomic evaluation of VLMs as WMs. Guided by this framework, we introduce WM-ABench, a large-scale benchmark comprising 23 fine-grained evaluation dimensions across 6 diverse simulated environments with controlled counterfactual simulations. Through 660 experiments on 15 latest commercial and open-source VLMs, we find that these models exhibit striking limitations in basic world modeling abilities. For instance, almost all models perform at near-random accuracy when distinguishing motion trajectories. Additionally, they lack disentangled understanding -- e.g., some models tend to believe blue objects move faster than green ones. More rich results and analyses reveal significant gaps between VLMs and human-level world modeling.

cs.CL

Descents and flag major index on conjugacy classes of colored permutation groups without short cycles

We consider the descent and flag major index statistics on the colored permutation groups, which are wreath products of the form $\mathfrak{S}_{n,r}=\mathbb{Z}_r\wr \mathfrak{S}_n$. We show that the $k$-th moments of these statistics on $\mathfrak{S}_{n,r}$ will coincide with the corresponding moments on all conjugacy classes without cycles of lengths $1,2,\ldots,2k$. Using this, we establish the asymptotic normality of the descent and flag major index statistics on conjugacy classes of $\mathfrak{S}_{n,r}$ with sufficiently long cycles. Our results generalize prior work of Fulman involving the descent and major index statistics on the symmetric group $\mathfrak{S}_n$. Our methods involve an intricate extension of Fulman's work on $\mathfrak{S}_n$ combined with the theory of the degree for a colored permutation statistic, as introduced by Campion Loth, Levet, Liu, Sundaram, and Yin.

math.CO

Reconstruction of caterpillar tanglegrams

A tanglegram consists of two rooted binary trees with the same number of leaves and a perfect matching between the leaves of the trees. Given a size-$n$ tanglegram, i.e., a tanglegram for two trees with $n$ leaves, a multiset of induced size-$(n-1)$ tanglegrams is obtained by deleting a pair of matched leaves in every possible way. Here, we analyze whether a size-$n$ tanglegram is uniquely encoded by this multiset of size-$(n-1)$ tanglegrams. We answer this question affirmatively in the case that at least one of the two trees of the tanglegram is a caterpillar tree.

math.CO

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering. To this end, we curate 75 ML engineering-related competitions from Kaggle, creating a diverse set of challenging tasks that test real-world ML engineering skills such as training models, preparing datasets, and running experiments. We establish human baselines for each competition using Kaggle's publicly available leaderboards. We use open-source agent scaffolds to evaluate several frontier language models on our benchmark, finding that the best-performing setup--OpenAI's o1-preview with AIDE scaffolding--achieves at least the level of a Kaggle bronze medal in 16.9% of competitions. In addition to our main results, we investigate various forms of resource scaling for AI agents and the impact of contamination from pre-training. We open-source our benchmark code (github.com/openai/mle-bench/) to facilitate future research in understanding the ML engineering capabilities of AI agents.

cs.CL

Figuring out Figures: Using Textual References to Caption Scientific Figures

Figures are essential channels for densely communicating complex ideas in scientific papers. Previous work in automatically generating figure captions has been largely unsuccessful and has defaulted to using single-layer LSTMs, which no longer achieve state-of-the-art performance. In our work, we use the SciCap datasets curated by Hsu et al. and use a variant of a CLIP+GPT-2 encoder-decoder model with cross-attention to generate captions conditioned on the image. Furthermore, we augment our training pipeline by creating a new dataset MetaSciCap that incorporates textual metadata from the original paper relevant to the figure, such as the title, abstract, and in-text references. We use SciBERT to encode the textual metadata and use this encoding alongside the figure embedding. In our experimentation with different models, we found that the CLIP+GPT-2 model performs better when it receives all textual metadata from the SciBERT encoder in addition to the figure, but employing a SciBERT+GPT2 model that uses only the textual metadata achieved optimal performance.

cs.CL

On the acceptance, commissioning, and quality assurance of electron FLASH units

Background & Purpose: FLASH or ultra-high dose rate (UHDR) radiation therapy (RT) has gained attention in recent years for its ability to spare normal tissues relative to conventional dose rate (CDR) RT in various preclinical trials. However, clinical implementation of this promising treatment option has been limited because of the lack of availability of accelerators capable of delivering UHDR RT. We established a framework for the acceptance, commissioning, and periodic quality assurance (QA) of electron FLASH units and present an example of commissioning. Methods: A protocol for acceptance, commissioning, and QA of UHDR linear accelerators was established by combining and adapting standards and professional recommendations for standard linear accelerators based on the experience with UHDR at four clinical centers that use different UHDR devices. Non-standard dosimetric beam parameters considered included pulse width, pulse repetition frequency, dose per pulse, and instantaneous dose rate, together with recommendations on how to acquire these measurements. Results: The 6 and 9 MeV beams of an UHDR electron device were commissioned by using this developed protocol. Measurements were acquired with a combination of ion chambers, beam current transformers (BCTs), and dose rate independent passive dosimeters. The unit was calibrated according to the concept of redundant dosimetry using a reference setup. Conclusions: This study provides detailed recommendations for the acceptance testing, commissioning, and routine QA of low-energy electron UHDR linear accelerators. The proposed framework is not limited to any specific unit, making it applicable to all existing eFLASH units in the market. Through practical insights and theoretical discourse, this document establishes a benchmark for the commissioning of UHDR devices for clinical use.

physics.med-ph

Development of novel ionization chambers for reference dosimetry in electron FLASH radiotherapy

The aim of this study was to optimize the design and performance of parallel plate ion chambers for use in ultra-high dose rate (UHDR) dosimetry applications, and evaluate their potential as reference class chambers for calibration purposes. Three chambers were designed and produced: the A11-VAR (0.2-1.0 mm electrode gap, 20 mm diameter collector), the A11-TPP (0.3 mm electrode gap, 20 mm diameter collector), and the A30 (0.3 mm electrode gap, 5.4 mm diameter collector).The chambers underwent full characterization using an UHDR 9 MeV electron beam with individually varied beam parameters of pulse repetition frequency (PRF, 10-120Hz), pulse width (PW, 0.5-4us), and pulse amplitude (0.01-9 Gy/pulse). The response of the ion chambers was evaluated as a function of the dose per pulse (DPP), PRF, PW, dose rate, electric field strength, and electrode gap. The chamber response was found to be dependent on DPP and PW, whose dependencies were mitigated with larger electric field strengths and smaller electrode spacing. At a constant electric field strength, we measured a larger charge collection efficiency (CCE) as a function of DPP for ion chambers with a smaller electrode gap in the A11-VAR. For ion chambers with identical electrode gap (A11-TPP and A30), higher electric field strengths were found to yield better CCE at higher DPP. A PW dependence was observed at low electric field strengths (500 V/mm) for DPP values ranging from 1-5 Gy at PWs ranging from 0.5-4 {\mu}s, but at electric field strengths of 1000 V/mm and higher, these effects become negligible. This study confirmed that the charge collection efficiency of ion chambers depends strongly on the electrode spacing and the electric field strength, and also on the DPP and the PW of the UHDR beam. The new finding of this study is that the PW dependence becomes negligible with reduced electrode spacing and increased electric field.

physics.med-ph

Characterization of a novel time-resolved, real-time scintillation dosimetry system for ultra-high dose rate radiation therapy applications

Background: Scintillation dosimetry has promising qualities for ultra-high dose rate (UHDR) radiotherapy (RT), but no system has shown compatibility with mean dose rates ($\bar{DR}$) above 100 Gy/s and doses per pulse ($D_p$) exceeding 1.5 Gy typical of UHDR (FLASH)-RT. The aim of this study was to characterize a novel scintillator dosimetry system with the potential of accommodating UHDRs. Methods: A thorough dosimetric characterization of the system was performed on an UHDR electron beamline. The system's response as a function of dose, $\bar{DR}$, $D_p$, and the pulse dose rate ${DR}_p$ was investigated, together with the system's dose sensitivity (signal per unit dose) as a function of dose history. The capabilities of the system for time-resolved dosimetric readout were also evaluated. Results: Within a tolerance of $\pm$3% the system exhibited dose linearity and was independent of $\bar{DR}$ and $D_p$ within the tested ranges of 1.8-1341 Gy/s and 0.005-7.68 Gy, respectively. A 6% reduction in the signal per unit dose was observed as ${DR}_p$ was increased from 8.9e4-1.8e6 Gy/s. Additionally, the dose delivered per integration window of the continuously sampling photodetector had to remain between 0.028 and 11.64 Gy to preserve a stable signal response per unit dose. The system accurately measured $D_p$ of individual pulses delivered at up to 120 Hz. The day-to-day variation of the signal per unit dose at a reference setup varied by up to $\pm$13% but remained consistent (<$\pm$2%) within each day of measurements and showed no signal loss as a function of dose history. Conclusions: With daily calibrations and ${DR}_p$ specific correction factors, the system reliably provides real-time, millisecond-resolved dosimetric measurements of pulsed conventional and UHDR beams from typical electron linacs, marking an important advancement in UHDR dosimetry.

physics.med-ph

Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?

Neural language models (LMs) can be used to evaluate the truth of factual statements in two ways: they can be either queried for statement probabilities, or probed for internal representations of truthfulness. Past work has found that these two procedures sometimes disagree, and that probes tend to be more accurate than LM outputs. This has led some researchers to conclude that LMs "lie" or otherwise encode non-cooperative communicative intents. Is this an accurate description of today's LMs, or can query-probe disagreement arise in other ways? We identify three different classes of disagreement, which we term confabulation, deception, and heterogeneity. In many cases, the superiority of probes is simply attributable to better calibration on uncertain answers rather than a greater fraction of correct, high-confidence answers. In some cases, queries and probes perform better on different subsets of inputs, and accuracy can further be improved by ensembling the two. Code is available at github.com/lingo-mit/lm-truthfulness.

cs.CL

Nova: Generative Language Models for Assembly Code with Hierarchical Attention and Contrastive Learning

Binary code analysis is the foundation of crucial tasks in the security domain; thus building effective binary analysis techniques is more important than ever. Large language models (LLMs) although have brought impressive improvement to source code tasks, do not directly generalize to assembly code due to the unique challenges of assembly: (1) the low information density of assembly and (2) the diverse optimizations in assembly code. To overcome these challenges, this work proposes a hierarchical attention mechanism that builds attention summaries to capture the semantics more effectively and designs contrastive learning objectives to train LLMs to learn assembly optimization. Equipped with these techniques, this work develops Nova, a generative LLM for assembly code. Nova outperforms existing techniques on binary code decompilation by up to 14.84 -- 21.58% (absolute percentage point improvement) higher Pass@1 and Pass@10, and outperforms the latest binary code similarity detection techniques by up to 6.17% Recall@1, showing promising abilities on both assembly generation and understanding tasks.

cs.SE

Decks of rooted binary trees

We consider extremal problems related to decks and multidecks of rooted binary trees (a.k.a. rooted phylogenetic tree shapes). Here, the deck (resp. multideck) of a tree $T$ refers to the set (resp. multiset) of leaf induced binary subtrees of $T$. On the one hand, we consider the reconstruction of trees from their (multi)decks. We give lower and upper bounds on the minimum (multi)deck size required to uniquely encode a rooted binary tree on $n$ leaves. On the other hand, we consider problems related to deck cardinalities. In particular, we characterize trees with minimum-size as well as maximum-size decks. Finally, we present some exhaustive computations for $k$-universal trees, i.e., rooted binary trees that contain all $k$-leaf rooted binary trees as induced subtrees.

math.CO

Multi-Institutional Audit of FLASH and Conventional Dosimetry with a 3D-Printed Anatomically Realistic Mouse Phantom

We conducted a multi-institutional audit of dosimetric variability between FLASH and conventional dose rate (CONV) electron irradiations by using an anatomically realistic 3D-printed mouse phantom. A CT scan of a live mouse was used to create a 3D model of bony anatomy, lungs, and soft tissue. A dual-nozzle 3D printer was used to print the mouse phantom using acrylonitrile butadiene styrene ($~1.02 g/cm^3$) and polylactic acid ($~1.24 g/cm^3$) simultaneously to simulate soft tissue and bone densities, respectively. The lungs were printed separately using lightweight polylactic acid ($~0.64 g/cm^3$). Hounsfield units (HU) and densities were compared with the reference CT scan of the live mouse. Print-to-print reproducibility of the phantom was assessed. Three institutions were each provided a phantom, and each institution performed two replicates of irradiations at selected mouse anatomic regions. The average dose difference between FLASH and CONV dose distributions and deviation from the prescribed dose were measured with radiochromic film. Compared to the reference CT scan, CT scans of the phantom demonstrated mass density differences of $0.10 g/cm^3$ for bone, $0.12 g/cm^3$ for lung, and $0.03 g/cm^3$ for soft tissue regions. Between phantoms, the difference in HU for soft tissue and bone was <10 HU from print to print. Lung exhibited the most variation (54 HU) but minimally affected dose distribution (<0.5% dose differences between phantoms). The mean difference between FLASH and CONV from the first replicate to the second decreased from 4.3% to 1.2%, and the mean difference from the prescribed dose decreased from 3.6% to 2.5% for CONV and 6.4% to 2.7% for FLASH. The framework presented here is promising for credentialing of multi-institutional studies of FLASH preclinical research to maximize the reproducibility of biological findings.

physics.med-ph

Universal rooted phylogenetic tree shapes and universal tanglegrams

We provide an $\Omega(n\log n) $ lower bound and an $O(n^2)$ upper bound for the smallest size of rooted binary trees (a.k.a. phylogenetic tree shapes), which are universal for rooted binary trees with $n$ leaves, i.e., contain all of them as induced binary subtrees. We explicitly compute the smallest universal trees for $n\leq 11$. We also provide an $\Omega(n^2) $ lower bound and an $O(n^4)$ upper bound for the smallest size of tanglegrams, which are universal for size $n$ tanglegrams, i.e., which contain all of them as induced subtanglegrams. Some of our results generalize to rooted $d$-ary trees and to $d$-ary tanglegrams.

math.CO

New Structures and their Applications to Variants of Zero Forcing and Propagation Time

We introduce a generalization of the concept of a chronological list of forces, called a relaxed chronology. This concept is used to introduce a new way of formulating the standard zero forcing process, which we refer to as parallel increasing path covers, or PIPs. The combinatorial properties of PIPs are utilized to identify bounds comparing standard zero forcing propagation time to positive semidefinite propagation time. A collection of paths within a set of PSD forcing trees, called a path bundle, is used to identify the PSD forcing analog of the reversal of a standard zero forcing process, as well as to draw a connection between PSD forcing and rigid-linkage forcing.

math.CO