SearcharxivSearch

arXiv subjects

Amy Rouillard

Publications and source records attributed to Amy Rouillard.

8 recordsLinked to original sources

Evaluating Multimodal LLMs for Inpatient Diagnosis: Real-World Performance, Safety, and Cost Across Ten Frontier Models

Background: Large language models (LLMs) are increasingly proposed for diagnostic support, but few evaluations use real-world multimodal inpatient data, particularly in low and middle-income country (LMIC) public hospitals. Methods: We conducted VALID, a retrospective evaluation of 539 multimodal inpatient cases from a tertiary public hospital in South Africa. Inputs included radiology imaging (CT, MRI, CXR) and reports, laboratory results, clinical notes, and vital signs. Expert panels adjudicated 300 cases (balanced and discordant subsets) to establish ground truth diagnoses, differentials, and reasoning. Ten multimodal LLMs generated zero-shot outputs. A calibrated three-model LLM Jury scored all outputs and routine ward diagnoses across diagnostic accuracy, differential quality, reasoning, and patient safety (>10,000 evaluations). Primary outcomes were composite scores ($S_3$, $S_4$) and win rates. Results: (i) LLM performance was tightly clustered (<15% variation) despite large cost differences; low-cost models performed comparably to top models. (ii) All LLMs significantly outperformed routine ward diagnoses on average diagnostic and safety scores. (iii) Top performance was achieved by GPT-5.1, followed by Gemini models. (vi) Adding radiology reports improved performance by 6%. (v) Diagnostic and reasoning scores were highly correlated ($\rho = 0.85$). (vi) Output rates varied (65-100%) due to input constraints. Results were robust across subsets and evaluation design. Conclusions: Across a real-world LMIC dataset, multimodal LLMs showed similar diagnostic performance despite large cost differences and outperformed routine care on average safety metrics. Affordability, robustness, and deployment constraints may outweigh marginal performance differences in LMIC settings.

cs.LG

Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?

Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alternative adjudicators. Here, we evaluate an LLM Jury, composed of three frontier AI models, for scoring 3334 diagnoses on 300 real-world low- and middle-income country (LMIC) hospital cases. Both LLM- and clinician-generated diagnoses are scored against expert panel diagnoses across four dimensions: diagnosis, differential diagnosis, clinical reasoning, and negative treatment risk. The LLM Jury scores are compared with expert and independent re-scoring panel scores to assess error metrics, inter-rater agreement, severe-risk errors, and the effect of post hoc calibration using isotonic regression. In our data, we find that: (i) the uncalibrated LLM Jury scores preserve ordinal agreement with the expert clinician panel scores, but are systematically lower; (ii) the probability of severe-risk errors is lower for the LLM Jury than the human expert re-score panels; (iii) the LLM Jury combined with LLM diagnoses can be used to identify diagnoses at high risk of error, enabling targeted expert review and improved panel efficiency; (iv) the calibrated LLM Jury scores and rankings of diagnosing agents show excellent agreement with those of the primary expert panels; (v) LLM Jury models show no self-preference bias, they did not score diagnoses generated by their own underlying model or models from the same vendor more (or less) favourably than those generated by other models. Together, these results provide evidence that a calibrated LLM Jury is a trustworthy and reliable proxy for expert clinician evaluation in medical AI benchmarking. Confirming these findings in other clinical settings is an important direction for future work.

cs.LG

What You Prompt is What You Get: Increasing Transparency of Prompting Using Prompt Cards

The rapid advancement and impressive capabilities of large language models (LLMs) have given rise to the field of prompt engineering, the practice of crafting inputs to guide LLMs toward high-quality, task-relevant outputs. A critical challenge facing the field is the lack of standardised prompt documentation and evaluation practices. Prompts can be long, complex and difficult to evaluate on subjective tasks. To address this challenge, we propose the use of prompt cards, structured summaries of prompt engineering practices inspired by the concept of model cards. Through prompt cards, the specific goals, considerations and steps taken during prompt engineering can be systematically documented and assessed. We present the prompt card approach and illustrate it on a specific task called wordalisation, in which structured numerical data is transformed into text. We argue that a well-structured prompt card can enable better reproducibility, transparency, improve prompt methodology and give an effective alternative to benchmarking for judging the quality of generated texts. By systemically capturing underlying model details, prompt intent, contextualisation strategies, evaluation practices and ethical considerations, prompt cards make explicit the often implicit design decisions that shape system behaviour. Documenting these choices is important as prompting increasingly involves complex pipelines with multiple moving parts.

cs.CY

Automated Quantum Algorithm Design using a Domain-Specific Language

We present a computational method to automatically design the n-qubit realisations of quantum algorithms. Our approach leverages a domain-specific language (DSL) that enables the construction of quantum circuits via modular building blocks, making it well-suited for evolutionary search. In this DSL quantum circuits are abstracted beyond the usual gate-sequence description and scale automatically to any problem size. This enables us to learn the algorithm structure rather than a specific unitary implementation. We demonstrate our method by automatically designing three known quantum algorithms-the Quantum Fourier Transform, the Deutsch-Jozsa algorithm, and Grover's search. Remarkably, we were able to learn the general implementation of each algorithm by considering examples of circuits containing at most 5-qubits. Our method proves robust, as it maintains performance across increasingly large search spaces. Convergence to the relevant algorithm is achieved with high probability and with moderate computational resources.

quant-ph

Representing data in words: A context engineering approach

Large language models (LLMs) have demonstrated remarkable potential across a broad range of applications. However, producing reliable text that faithfully represents data remains a challenge. While prior work has shown that task-specific conditioning through in-context learning and knowledge augmentation can improve performance, LLMs continue to struggle with interpreting and reasoning about numerical data. To address this, we introduce wordalisations, a methodology for generating stylistically natural narratives from data. Much like how visualisations display numerical data in a way that is easy to digest, wordalisations abstract data insights into descriptive texts. To illustrate the method's versatility, we apply it to three application areas: scouting football players, personality tests, and international survey data. Due to the absence of standardized benchmarks for this specific task, we conduct LLM-as-a-judge and human-as-a-judge evaluations to assess accuracy across the three applications. We found that wordalisation produces engaging texts that accurately represent the data. We further describe best practice methods for open and transparent development of communication about data.

cs.HC

Measurement-based Feedback Control of a Quantum System in a Harmonic Potential

We present a formulation of measurement-based feedback control of a single quantum particle in one spatial dimension. An arbitrary linear combination of the position and momentum of the particle is continuously monitored, and feedback proportional to the measured signal is used to control the system. We derive a feedback master equation and discuss a general approach to computing the steady-state solutions for arbitrary potentials. For a quantum harmonic oscillator or a free particle, we show that it is possible to cool and confine the system using feedback that simultaneously damps the measured observable and its conjugate momentum, as well as compensates for noise introduced by the measurement. In addition, we demonstrate that appropriate feedback adds a quadratic term in the measured observable to the Hamiltonian of the system. For the particular case of the harmonic potential, we describe the exact dynamics of the system that can be cooled to the ground state. Moreover, we provide an argument for the possibility to cool systems with arbitrary potentials provided that the measurement is strong enough to localise the particle on an interval smaller than the characteristic length scale of the potential.

quant-ph

Deterministic preparation of optical qubits with coherent feedback control

We propose a class of preparation schemes for orbital angular momentum and polarisation qubits carried by single photons or classical states of light based on coherent feedback control by an ancillary degree of freedom of light. The preparation methods use linear optics and include the transcription of an arbitrary polarisation state onto a two-level OAM system (swap) for arbitrary OAM values plus/minus "l" within a light beam, i.e. without spatial interferometer. The preparations can be carried out with unit efficiency independent from the potentially unknown initial state of the system. Moreover, we show how to translate measurement-based qubit control channels into coherent feedback schemes for optical implementation.

quant-ph

Robust control of quantum systems by quantum systems

Quantum systems can be controlled by other quantum systems in a reversible way, without any information leaking to the outside of the system-controller compound. Such coherent quantum control is deterministic, is less noisy than measurement-based feedback control, and has potential applications in a variety of quantum technologies, including quantum computation, quantum communication and quantum metrology. Here we introduce a coherent feedback protocol, consisting of a sequence of identical interactions with controlling quantum systems, that steers a quantum system from an arbitrary initial state towards a target state. We determine the broad class of such coherent feedback channels that achieve convergence to the target state, and then stabilise as well as protect it against noise. Our results imply that also weak system-controller interactions can counter noise if they occur with suitably high frequency. We provide an example of a control scheme that does not require knowledge of the target state encoded in the controllers, which could be the result of a quantum computation. It thus provides a mechanism for autonomous, purely quantum closed-loop control.

quant-ph