SearcharxivSearch

arXiv subjects

Ashutosh Joshi

Publications and source records attributed to Ashutosh Joshi.

7 recordsLinked to original sources

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Primary care guards against diagnostic errors and unsafe care; agents assisting in this domain warrant evaluation against the same risks. Current benchmarks focus on medical knowledge, assessed through isolated question-answering or clinician-facing tasks. PatientAgentBench benchmarks patient-facing agentic healthcare; it evaluates a foundation model, wrapped in an agent with a sandbox of healthcare tools, conversing with a simulated patient. Each conversation is scored by an LLM-as-a-Jury across six dimensions via over a hundred conversation-agnostic, clinician-grounded criteria. To validate alignment, licensed clinicians annotated shared conversations, yielding 79-93% adjacent agreement between jury and expert raters, on par with or exceeding clinician inter-rater agreement. We benchmarked 10 models across four families on the same 1,200 scenarios and found clinical gaps. Triage quality is the most discriminating dimension: pass rates rise from 32% for the weakest models to 88% for the strongest, with agents often acting on administrative requests without clinical screening. Clinical safety and workflow accuracy follow the same pattern: the weakest models fail often, fabricating unexecuted actions, while frontier models fail on only 1-3% of cases, from unverified tool outputs and omitted crisis resources in an emergency. More capable models narrow these gaps but do not close them; the strongest scores only 4.25 of 5 overall. These failures surface only in sustained, tool-using conversations against realistic patient records, confirming that static benchmarks are insufficient as healthcare agentic systems gain autonomy. We release the framework as a reproducible, clinician-validated evaluation standard to help the field close this gap.

cs.AI

IMCBench: A benchmark for multimodal LLMs in Image-grounded Medical Conversations

Recent advances in large language models and vision-language models have enabled reasoning over multimodal data, offering opportunities for clinical applications such as decision support and triaging. However, existing medical AI benchmarks are fragmented: some support multi-turn dialogues but lack images, while others provide multimodal inputs but focus on single-turn QA tasks. To address this gap, we introduce IMCBench, an image-grounded, multi-turn medical conversation benchmark that pairs real, publicly available clinical images with synthetic patient profiles to simulate realistic patient-clinician interactions. Each conversation is evaluated across three clinical dimensions: safety, accuracy, and appropriate use of uncertainty in diagnosis. We benchmark eight multimodal frontier models across four model families (Claude, GPT, Nova, and Llama), scoring each on a 1-5 scale using LLM-as-Jury scoring calibrated against expert clinician annotations. Our results show that Claude Opus 4.6 achieves the highest overall score (3.61), followed by Claude Sonnet 4.6 (3.30) and GPT-5.2 (3.29), though no model dominates all dimensions and safety degrades for both malignant and rare conditions ($\Delta$ = -0.27 each). Ablation studies further reveal that both visual input and EHR context contribute to safe guidance (safety drops of 0.18 and 0.23 on average when each is removed), with stronger models leveraging visual features more effectively. Together, these findings demonstrate that accurate clinical description does not guarantee safe patient guidance, motivating the need for multi-dimensional evaluation frameworks in medical AI.

cs.AI

Linking Knowledge to Care: Knowledge Graph-Augmented Medical Follow-Up Question Generation

Clinical diagnosis is time-consuming, requiring intensive interactions between patients and medical professionals. While large language models (LLMs) could ease the pre-diagnostic workload, their limited domain knowledge hinders effective medical question generation. We introduce a Knowledge Graph-augmented LLM with active in-context learning to generate relevant and important follow-up questions, KG-Followup, serving as a critical module for the pre-diagnostic assessment. The structured medical domain knowledge graph serves as a seamless patch-up to provide professional domain expertise upon which the LLM can reason. Experiments demonstrate that KG-Followup outperforms state-of-the-art methods by 5% - 8% on relevant benchmarks in recall.

cs.CL

Exploring Query Understanding for Amazon Product Search

Online shopping platforms, such as Amazon, offer services to billions of people worldwide. Unlike web search or other search engines, product search engines have their unique characteristics, primarily featuring short queries which are mostly a combination of product attributes and structured product search space. The uniqueness of product search underscores the crucial importance of the query understanding component. However, there are limited studies focusing on exploring this impact within real-world product search engines. In this work, we aim to bridge this gap by conducting a comprehensive study and sharing our year-long journey investigating how the query understanding service impacts Amazon Product Search. Firstly, we explore how query understanding-based ranking features influence the ranking process. Next, we delve into how the query understanding system contributes to understanding the performance of a ranking model. Building on the insights gained from our study on the evaluation of the query understanding-based ranking model, we propose a query understanding-based multi-task learning framework for ranking. We present our studies and investigations using the real-world system on Amazon Search.

cs.IR

REAPER: Reasoning based Retrieval Planning for Complex RAG Systems

Complex dialog systems often use retrieved evidence to facilitate factual responses. Such RAG (Retrieval Augmented Generation) systems retrieve from massive heterogeneous data stores that are usually architected as multiple indexes or APIs instead of a single monolithic source. For a given query, relevant evidence needs to be retrieved from one or a small subset of possible retrieval sources. Complex queries can even require multi-step retrieval. For example, a conversational agent on a retail site answering customer questions about past orders will need to retrieve the appropriate customer order first and then the evidence relevant to the customer's question in the context of the ordered product. Most RAG Agents handle such Chain-of-Thought (CoT) tasks by interleaving reasoning and retrieval steps. However, each reasoning step directly adds to the latency of the system. For large models this latency cost is significant -- in the order of multiple seconds. Multi-agent systems may classify the query to a single Agent associated with a retrieval source, though this means that a (small) classification model dictates the performance of a large language model. In this work we present REAPER (REAsoning-based PlannER) - an LLM based planner to generate retrieval plans in conversational systems. We show significant gains in latency over Agent-based systems and are able to scale easily to new and unseen use cases as compared to classification-based planning. Though our method can be applied to any RAG system, we show our results in the context of a conversational shopping assistant.

cs.IR

Photoinduced modulation of refractive index in Langmuir-Blodgett films of Azo-based H-shaped liquid crystal molecules

The development of optically active area consisting of organic molecules are essential for the devices like optical switches and waveguides, as it can be easily maneuvered by the application of suitable electromagnetic (EM) waves. In this article, we report the development of a photoactive surface by the deposition of a single layer of Langmuir-Blodgett (LB) film of a novel H-shaped liquid crystal (HLC) molecule. The synthesized HLC molecules possess azo-groups and nitro-groups. The azo-group can be isomerized (trans-cis transformation) by irradiating them with ultraviolet (UV) light. The nitro-group can provide sufficient amphiphilicity to the HLC molecules to form a stable Langmuir monolayer at air-water interface. The Langmuir monolayer of the HLC molecules exhibited gas and liquid-like phases. A single layer of LB film of HLC molecules was deposited on a gold chip of a home-built surface plasmon resonance (SPR) instrument. The azo-groups of the molecules in LB film was excited by UV irradiation leading to a change in morphology due to trans-cis transformation. Such a change in morphology can lead to a miniscule change in refractive index (RI) of the LB film. SPR is a label free and highly sensitive optical phenomenon for the measurement of such changes in RI. In our studies, we found systematic changes in the resonance angle of the LB film of HLC molecules as a function of intensity of the UV irradiation. We measured switch-on and switch-off intensity which may suggest that the LB film of HLC molecules can find applications in optical switches or waveguides.

physics.optics

Surface plasmon resonance for in-plane birefringence measurement of anisotropic thin organic film

The measurement of in-plane birefringence ($Δ{n}$) of ultrathin film is challenging due to a significant deviation of physical properties of materials in ultrathin regime as compared to that in bulk state. Surface plasmon resonance (SPR) phenomenon can be employed to measure change in refractive index of ultrathin film at a very high resolution. This article discusses simulation of SPR phenomenon in Kretschmann configuration for the measurement of $Δ{n}$ in organic thin film exhibiting nematic-like ordering on the two dimensional gold surface. The distribution of plasmonic field on the gold surface was found to be anisotropic. This suggested that the coupling plasmonic field with that of organic thin film exhibiting nematic-like ordering on the gold surface will be non-isotropic. Therefore, a non-zero difference in resonance angle (RA) was obtained from SPR measurement performed along the optic-axis (OA) and orthogonal to OA of the in-plane nematic ordering ($Δθ$). A calibration surface showing the variation of ($Δθ$) as a function of $Δ{n}$ and thickness of thin organic film consisting of shape anisotropic tilted molecules exhibiting nematic-like ordering on gold surface was obtained. This calibration surface was employed for the measurement of $Δ{n}$ of single layer of Langmuir-Blodgett films of cadmium stearate (CdSA) and 4'-octyl-4-biphenylcarbonitrile (8CB) deposited on SPR chips. The thickness of the LB films was estimated using X-ray reflectivity measurement and $Δθ$ was measured using a home built SPR instrument. The $Δ{n}$ values were found to be 0.012 and 0.022 for ultrathin films of CdSA and 8CB molecules, respectively.

cond-mat.soft