SearcharxivSearch

arXiv subjects

Andrew Smith

Publications and source records attributed to Andrew Smith.

At least 19 recordsLinked to original sources

Readout and PID using AIML for SoLID High Background Cherenkov Detectors

We present the development of readout electronics and artificial-intelligence-based particle-identification methods for the SoLID Cherenkov detectors at Jefferson Lab. To operate in the high-rate, high-background SoLID environment, we designed a MAROC sum readout system for multianode photomultiplier tubes that provides simultaneous pixel, quadrant-sum, and total-sum signals. Bench studies show that the system can sustain rates at or above those expected for SoLID while maintaining acceptable pedestal behavior and signal linearity. Using realistic Geant4 simulations for the heavy-gas Cherenkov detector, we then investigate $\pi/K$ separation with beam-related background. A simple photoelectron-counting cut is insufficient under these conditions, whereas multilayer perceptron models trained on PMT, quad, and pixel readout data perform substantially better. The quad and pixel readout schemes achieve pion and kaon efficiencies above 90\% and clearly outperform PMT-only readout. These results demonstrate that the combination of high-rate MAROC sum electronics and AIML-based pattern recognition provides a practical path toward robust SoLID Cherenkov PID.

physics.ins-det

Scaling Latent Reasoning via Looped Language Models

Modern LLMs are trained to "think" primarily via explicit text generation, such as chain-of-thought (CoT), which defers reasoning to post-training and under-leverages pre-training data. We present and open-source Ouro, named after the recursive Ouroboros, a family of pre-trained Looped Language Models (LoopLM) that instead build reasoning into the pre-training phase through (i) iterative computation in latent space, (ii) an entropy-regularized objective for learned depth allocation, and (iii) scaling to 7.7T tokens. Ouro 1.4B and 2.6B models enjoy superior performance that match the results of up to 12B SOTA LLMs across a wide range of benchmarks. Through controlled experiments, we show this advantage stems not from increased knowledge capacity, but from superior knowledge manipulation capabilities. We also show that LoopLM yields reasoning traces more aligned with final outputs than explicit CoT. We hope our results show the potential of LoopLM as a novel scaling direction in the reasoning era. Our model is available here: http://ouro-llm.github.io.

cs.CL

A Survey on Latent Reasoning

Large Language Models (LLMs) have demonstrated impressive reasoning capabilities, especially when guided by explicit chain-of-thought (CoT) reasoning that verbalizes intermediate steps. While CoT improves both interpretability and accuracy, its dependence on natural language reasoning limits the model's expressive bandwidth. Latent reasoning tackles this bottleneck by performing multi-step inference entirely in the model's continuous hidden state, eliminating token-level supervision. To advance latent reasoning research, this survey provides a comprehensive overview of the emerging field of latent reasoning. We begin by examining the foundational role of neural network layers as the computational substrate for reasoning, highlighting how hierarchical representations support complex transformations. Next, we explore diverse latent reasoning methodologies, including activation-based recurrence, hidden state propagation, and fine-tuning strategies that compress or internalize explicit reasoning traces. Finally, we discuss advanced paradigms such as infinite-depth latent reasoning via masked diffusion models, which enable globally consistent and reversible reasoning processes. By unifying these perspectives, we aim to clarify the conceptual landscape of latent reasoning and chart future directions for research at the frontier of LLM cognition. An associated GitHub repository collecting the latest papers and repos is available at: https://github.com/multimodal-art-projection/LatentCoT-Horizon/.

cs.CL

Large Language Models as 'Hidden Persuaders': Fake Product Reviews are Indistinguishable to Humans and Machines

Reading and evaluating product reviews is central to how most people decide what to buy and consume online. However, the recent emergence of Large Language Models and Generative Artificial Intelligence now means writing fraudulent or fake reviews is potentially easier than ever. Through three studies we demonstrate that (1) humans are no longer able to distinguish between real and fake product reviews generated by machines, averaging only 50.8% accuracy overall - essentially the same that would be expected by chance alone; (2) that LLMs are likewise unable to distinguish between fake and real reviews and perform equivalently bad or even worse than humans; and (3) that humans and LLMs pursue different strategies for evaluating authenticity which lead to equivalently bad accuracy, but different precision, recall and F1 scores - indicating they perform worse at different aspects of judgment. The results reveal that review systems everywhere are now susceptible to mechanised fraud if they do not depend on trustworthy purchase verification to guarantee the authenticity of reviewers. Furthermore, the results provide insight into the consumer psychology of how humans judge authenticity, demonstrating there is an inherent 'scepticism bias' towards positive reviews and a special vulnerability to misjudge the authenticity of fake negative reviews. Additionally, results provide a first insight into the 'machine psychology' of judging fake reviews, revealing that the strategies LLMs take to evaluate authenticity radically differ from humans, in ways that are equally wrong in terms of accuracy, but different in their misjudgments.

cs.CL

Bridging Quantum and Classical Computing in Drug Design: Architecture Principles for Improved Molecule Generation

Hybrid quantum-classical machine learning offers a path to leverage noisy intermediate-scale quantum (NISQ) devices for drug discovery, but optimal model architectures remain unclear. We systematically optimize the quantum-classical bridge architecture of generative adversarial networks (GANs) for molecule discovery using multi-objective Bayesian optimization. Our optimized model (BO-QGAN) significantly improves performance, achieving a 2.27-fold higher Drug Candidate Score (DCS) than prior quantum-hybrid benchmarks and 2.21-fold higher than the classical baseline, while reducing parameter count by more than 60%. Key findings favor layering multiple (3-4) shallow (4-8 qubit) quantum circuits sequentially, while classical architecture shows less sensitivity above a minimum capacity. This work provides the first empirically-grounded architectural guidelines for hybrid models, enabling more effective integration of current quantum computers into pharmaceutical research pipelines.

cs.LG

Monitoring morphometric drift in lifelong learning segmentation of the spinal cord

Morphometric measures derived from spinal cord segmentations can serve as diagnostic and prognostic biomarkers in neurological diseases and injuries affecting the spinal cord. While robust, automatic segmentation methods to a wide variety of contrasts and pathologies have been developed over the past few years, whether their predictions are stable as the model is updated using new datasets has not been assessed. This is particularly important for deriving normative values from healthy participants. In this study, we present a spinal cord segmentation model trained on a multisite $(n=75)$ dataset, including 9 different MRI contrasts and several spinal cord pathologies. We also introduce a lifelong learning framework to automatically monitor the morphometric drift as the model is updated using additional datasets. The framework is triggered by an automatic GitHub Actions workflow every time a new model is created, recording the morphometric values derived from the model's predictions over time. As a real-world application of the proposed framework, we employed the spinal cord segmentation model to update a recently-introduced normative database of healthy participants containing commonly used measures of spinal cord morphometry. Results showed that: (i) our model outperforms previous versions and pathology-specific models on challenging lumbar spinal cord cases, achieving an average Dice score of $0.95 \pm 0.03$; (ii) the automatic workflow for monitoring morphometric drift provides a quick feedback loop for developing future segmentation models; and (iii) the scaling factor required to update the database of morphometric measures is nearly constant among slices across the given vertebral levels, showing minimum drift between the current and previous versions of the model monitored by the framework. The code and model are open-source and accessible via Spinal Cord Toolbox v7.0.

cs.CV

The Amazon Nova Family of Models: Technical Report and Model Card

We present Amazon Nova, a new generation of state-of-the-art foundation models that deliver frontier intelligence and industry-leading price performance. Amazon Nova Pro is a highly-capable multimodal model with the best combination of accuracy, speed, and cost for a wide range of tasks. Amazon Nova Lite is a low-cost multimodal model that is lightning fast for processing images, video, documents and text. Amazon Nova Micro is a text-only model that delivers our lowest-latency responses at very low cost. Amazon Nova Canvas is an image generation model that creates professional grade images with rich customization controls. Amazon Nova Reel is a video generation model offering high-quality outputs, customization, and motion control. Our models were built responsibly and with a commitment to customer trust, security, and reliability. We report benchmarking results for core capabilities, agentic performance, long context, functional adaptation, runtime performance, and human evaluation.

cs.AI

Enhancing Exoplanet Ephemerides by Leveraging Professional and Citizen Science Data: A Test Case with WASP-77A b

We present an updated ephemeris and physical parameters for the exoplanet WASP-77 A b. In this effort, we combine 64 ground- and space-based transit observations, 6 space-based eclipse observations, and 32 radial velocity observations to produce the most precise orbital solution to date for this target, aiding in the planning of James Webb Space Telescope (JWST) and Ariel observations and atmospheric studies. We report a new orbital period of 1.360029395 +- 5.7e-8 days, a new mid-transit time of 2459957.337860 +- 4.3e-5 BJDTDB (Barycentric Julian Date in the Barycentric Dynamical Time scale; arXiv:1005.4415) and a new mid-eclipse time of 2459956.658192 +- 6.7e-5 BJDTDB. Furthermore, the methods presented in this study reduce the uncertainties in the planet mass to 1.6654 +- 4.5e-3 Mjup and orbital period to 1.360029395 +- 5.7e-8 days by factors of 15.1 and 10.9, respectively. Through a joint fit analysis comparison of transit data taken by space-based and citizen science-led initiatives, our study demonstrates the power of including data collected by citizen scientists compared to a fit of the space-based data alone. Additionally, by including a vast array of citizen science data from ExoClock, Exoplanet Transit Database (ETD), and Exoplanet Watch, we can increase our observational baseline and thus acquire better constraints on the forward propagation of our ephemeris than what is achievable with TESS data alone.

astro-ph.EP

Ultra-Resolution Cascaded Diffusion Model for Gigapixel Image Synthesis in Histopathology

Diagnoses from histopathology images rely on information from both high and low resolutions of Whole Slide Images. Ultra-Resolution Cascaded Diffusion Models (URCDMs) allow for the synthesis of high-resolution images that are realistic at all magnification levels, focusing not only on fidelity but also on long-distance spatial coherency. Our model beats existing methods, improving the pFID-50k [2] score by 110.63 to 39.52 pFID-50k. Additionally, a human expert evaluation study was performed, reaching a weighted Mean Absolute Error (MAE) of 0.11 for the Lower Resolution Diffusion Models and a weighted MAE of 0.22 for the URCDM.

eess.IV

Towards contrast-agnostic soft segmentation of the spinal cord

Spinal cord segmentation is clinically relevant and is notably used to compute spinal cord cross-sectional area (CSA) for the diagnosis and monitoring of cord compression or neurodegenerative diseases such as multiple sclerosis. While several semi and automatic methods exist, one key limitation remains: the segmentation depends on the MRI contrast, resulting in different CSA across contrasts. This is partly due to the varying appearance of the boundary between the spinal cord and the cerebrospinal fluid that depends on the sequence and acquisition parameters. This contrast-sensitive CSA adds variability in multi-center studies where protocols can vary, reducing the sensitivity to detect subtle atrophies. Moreover, existing methods enhance the CSA variability by training one model per contrast, while also producing binary masks that do not account for partial volume effects. In this work, we present a deep learning-based method that produces soft segmentations of the spinal cord. Using the Spine Generic Public Database of healthy participants ($\text{n}=267$; $\text{contrasts}=6$), we first generated participant-wise soft ground truth (GT) by averaging the binary segmentations across all 6 contrasts. These soft GT, along with aggressive data augmentation and a regression-based loss function, were used to train a U-Net model for spinal cord segmentation. We evaluated our model against state-of-the-art methods and performed ablation studies involving different loss functions and domain generalization methods. Our results show that using the soft segmentations along with a regression loss function reduces CSA variability ($p < 0.05$, Wilcoxon signed-rank test). The proposed spinal cord segmentation model generalizes better than the state-of-the-art methods amongst unseen datasets, vendors, contrasts, and pathologies (compression, lesions), while accounting for partial volume effects.

eess.IV

Easy-plane spin Hall oscillator

Spin Hall oscillators (SHOs) based on bilayers of a ferromagnet (FM) and a non-magnetic heavy metal (HM) are electrically tunable nanoscale microwave signal generators. Achieving high output power in SHOs requires driving large-amplitude magnetization dynamics by a direct spin Hall current. The maximum possible amplitude of such oscillations with the precession cone angle nearing $90^\circ$ is predicted for FM layers with easy-plane magnetic anisotropy and spin Hall current polarization perpendicular to the easy plane. While many FMs exhibit natural easy-plane anisotropy in the FM film plane, the spin Hall current in a HM|FM bilayer is polarized in this plane and thus cannot drive large-amplitude magneto-dynamics. Here we present a new type of SHO engineered to have the easy-plane anisotropy oriented normal to the film plane, enabling large-amplitude easy-plane dynamics driven by spin Hall current. Our experiments and micromagnetic simulations demonstrate that the desired easy-plane anisotropy can be achieved by tuning the magnetic shape anisotropy and perpendicular magnetic anisotropy in a nanowire SHO, leading to a significant enhancement of the generated microwave power. The easy-plane SHO experimentally demonstrated here is an ideal candidate for realization of a spintronic spiking neuron. Our results provide a new approach to design of high-power SHOs for wireless communications, neuromorphic computing, and microwave assisted magnetic recording.

cond-mat.mes-hall

A Double Layered Water Cherenkov Detector Array for Gamma-Ray Astronomy

Ground-level particle detection is now a well-established approach to TeV gamma-ray astronomy. Detection of Cherenkov light produced in water-filled detection units is a proven and cost-effective method. Here we discuss the optimization of the units towards the future Southern Wide-field Gamma-ray Observatory (SWGO). In this context, we investigate a new type of configuration in which each water Cherenkov detector (WCD) unit in the array comprises two chambers with black or reflective walls and a single photomultiplier tube (PMT) in each chamber. We find that this is a cost-effective approach that improves the performance of the WCD array with respect to current approaches. A shallow lower chamber with a PMT facing downwards enables muon tagging and the identification of hadron-induced air showers, which are the primary source of background in gamma-ray astronomy. We investigate how gamma/hadron separation power and achievable angular resolution depend on the geometry and wall reflectivity of the detector units in this configuration. We find that excellent angular resolution, background rejection power and low-energy response are achievable in this double-layer configuration, with the aid of reflective surfaces in both chambers.

astro-ph.IM

A stochastic household model for vector-borne diseases

We introduce a stochastic household model for vector-borne diseases, in particular as relevant to prominent vectors belonging to the Aedes genus and hence the Zika, chikungunya, and dengue viruses. In this model, vectors remain local to each household, while hosts mix for a proportion of their time in their household and the remaining proportion in the population at random. This is approximated with a two-type branching process, allowing us to efficiently calculate a number of useful epidemiological characteristics, such as reproductive numbers, early growth rates and household-type proportions, offspring distributions, probabilities of a major outbreak, and within-household final size distributions. We compare control interventions of spraying -- reducing the number of vectors in each household -- and social-distancing -- having individuals spend more time at home -- in terms of these characteristics.

q-bio.PE

Characterization of Multianode Photomultiplier Tubes for use in the CLAS12 RICH Detector

We present results of the detailed study of several hundred Hamamatsu H12700 Multianode Photomultiplier Tubes (MaPMTs), characterizing their response to the Cherenkov light photons in the second Ring Imaging Cherenkov detector, a part of the CLAS12 upgrade at Jefferson Lab. The total number of pixels studied was 25536. The single photoelectron spectra were measured for each pixel at different high voltages and light intensities of the laser test setup. Using the same dedicated front-end electronics as in the first RICH detector, the setup allowed us to characterize each pixel's properties such as gain, quantum efficiency, signal crosstalk between neighboring pixels, and determine the signal threshold values to optimize their efficiency to detect Cherenkov photons. A recently published state-of-the-art mathematical model, describing photon detector response functions measured in low light conditions, was extended to include the description of the crosstalk contributions to the spectra. The database of extracted parameters will be used for the final selection of the MaPMTs, their arrangement in the new RICH detector, and the optimization of the operational settings of the front-end electronics. The results show that the characteristics of the H12700 MaPMTs satisfy our requirements for the position-sensitive single photoelectron detectors.

physics.ins-det

Parametric resonance of spin waves in ferromagnetic nanowires tuned by spin Hall torque

We present a joint experimental and theoretical study of parametric resonance of spin wave eigenmodes in Ni$_{80}$Fe$_{20}$/Pt bilayer nanowires. Using electrically detected magnetic resonance, we measure the spectrum of spin wave eigenmodes in transversely magnetized nanowires and study parametric excitation of these eigenmodes by a microwave magnetic field. We also develop an analytical theory of spin wave eigenmodes and their parametric excitation in the nanowire geometry that takes into account magnetic dilution at the nanowire edges. We measure tuning of the parametric resonance threshold by antidamping spin Hall torque from a direct current for the edge and bulk eigenmodes, which allows us to independently evaluate frequency, damping and ellipticity of the modes. We find good agreement between theory and experiment for parametric resonance of the bulk eigenmodes but significant discrepancies arise for the edge modes. The data reveals that ellipticity of the edge modes is significantly lower than expected, which can be attributed to strong modification of magnetism at the nanowire edges. Our work demonstrates that parametric resonance of spin wave eigenmodes is a sensitive probe of magnetic properties at edges of thin-film nanomagnets.

cond-mat.mes-hall

Application of Machine Learning to Sleep Stage Classification

Sleep studies are imperative to recapitulate phenotypes associated with sleep loss and uncover mechanisms contributing to psychopathology. Most often, investigators manually classify the polysomnography into vigilance states, which is time-consuming, requires extensive training, and is prone to inter-scorer variability. While many works have successfully developed automated vigilance state classifiers based on multiple EEG channels, we aim to produce an automated and open-access classifier that can reliably predict vigilance state based on a single cortical electroencephalogram (EEG) from rodents to minimize the disadvantages that accompany tethering small animals via wires to computer programs. Approximately 427 hours of continuously monitored EEG, electromyogram (EMG), and activity were labeled by a domain expert out of 571 hours of total data. Here we evaluate the performance of various machine learning techniques on classifying 10-second epochs into one of three discrete classes: paradoxical, slow-wave, or wake. Our investigations include Decision Trees, Random Forests, Naive Bayes Classifiers, Logistic Regression Classifiers, and Artificial Neural Networks. These methodologies have achieved accuracies ranging from approximately 74% to approximately 96%. Most notably, the Random Forest and the ANN achieved remarkable accuracies of 95.78% and 93.31%, respectively. Here we have shown the potential of various machine learning classifiers to automatically, accurately, and reliably classify vigilance states based on a single EEG reading and a single EMG reading.

cs.LG

Low-Force Elastocaloric Refrigeration via Bending

Elastocaloric cooling has been identified as a promising alternative to high global warming potential vapor compression cooling. Two key bottlenecks to adoption are the need for bulky/expensive actuators to provide sufficient uniaxial stress and inadequate elastocaloric material fatigue life. This paper defines the physics that govern performance of axisymmetric flexural bending for use as an emerging low-force and low-fatigue elastocaloric heating and cooling mechanism and further demonstrates a continuous rotary-driven cooling prototype using polycrcrystalline Ni50.7Ti48.9. Elastocaloric material performance is determined using infrared thermography during uniaxial-tension and four-point bending thermomechanical testing. A systematic study reveals the effects of strain rate (from 0.001 to 0.025 s-1), maximum strain (from 2 to 8%), and strain mode on the temperature evolution, mechanical response, and coefficient of performance. Four-point bending experiments demonstrate a temperature reduction up to 11.3{\deg}C, material coefficients of performance between 2.31 and 21.71, and a 6.09- to 7.75-fold reduction in required actuation force compared to uniaxial tension. The absence of L\"uders bands and reduced mechanical dissipation during flexure represent reduced microstructure degradation and improved fatigue life. The rotary-based elastocaloric cooling prototype is shown to provide similar thermomechanical performance with the added benefit of discrete hot and cold zones, continuous cooling, inexpensive rotary actuation, and scalability, which represents a significant advancement for compact, long lifetime, and inexpensive elastocaloric cooling.

cond-mat.mtrl-sci

High-Capacity High-Power Thermal Energy Storage Using Solid-Solid Martensitic Transformations

Adding thermal conductivity enhancements to increase thermal power in solid-liquid phase-change thermal energy storage modules compromises volumetric energy density and often times reduces the mass and volume of active phase change material (PCM) by well over half. In this study, a new concept of building thermal energy storage modules using high-conductivity, solid-solid, shape memory alloys is demonstrated to eliminate this trade-off and enable devices that have both high heat transfer rate and high thermal capacity. Nickel titanium, Ni50.28Ti49.36, was solution heat treated and characterized using differential scanning calorimetry and Xenon Flash to determine transformation temperature (78deg-C), latent heat (183 kJm-3), and thermal conductivity in the Austenite and Martensite phases (12.92/12.64 Wm-1K-1). Four parallel-plate thermal energy storage demonstrators were designed, fabricated, and tested in a thermofluidic test setup. These included a baseline sensible heating module (aluminum), a conventional solid-liquid PCM module (aluminum/1-octadecanol), an all-solid-solid PCM module (Ni50.28Ti49.36), and a composite solid-solid/solid-liquid PCM module (Ni50.28Ti49.36/1-octadecanol). By using high-conductivity solid-solid PCMs, and eliminating the need for encapsulants and conductivity enhancements, we are able to demonstrate a 1.73-3.38 times improvement in volumetric thermal capacity and a 2.03-3.21 times improvement in power density as compared to the conventional approaches. These experimental results are bolstered by analytical models to explain the observed heat transfer physics and reveal a 5.86 times improvement in thermal time constant. This work demonstrates the ability to build high-capacity and high-power thermal energy storage modules using multifunctional shape memory alloys and opens the door for leap ahead improvement in thermal energy storage performance.

cond-mat.mtrl-sci