SearcharxivSearch

arXiv subjects

Minjae Cho

Publications and source records attributed to Minjae Cho.

At least 19 recordsLinked to original sources

EHRNote-ChatQA: A Benchmark for Evidence-Grounded Multi-Turn Clinical Question Answering over Longitudinal Discharge Summaries

Discharge summaries are crucial clinical documents containing the context of a patient's overall hospital stay, and are routinely reviewed by medical experts for patient readmission, ongoing care, and diagnostic decision-making. When reviewing them, medical experts often must iteratively synthesize information across multiple summaries while verifying the evidence supporting each answer. Although large language models (LLMs) are increasingly explored for clinical question answering, existing benchmarks do not sufficiently reflect this setting: they often evaluate exam-style medical knowledge or focus on single-turn question answering with limited evidence-grounding evaluation. We introduce EHRNote-ChatQA, the first benchmark for evidence-grounded multi-turn clinical question answering over patients' multiple discharge summaries. Built from de-identified MIMIC-IV discharge summaries, EHRNote-ChatQA contains 967 patient-level multi-turn samples spanning one to five notes and 16,072 medical-expert-verified QA pairs (8,036 content questions, each paired with an evidence-grounding question) across eight clinical categories. The benchmark is constructed through an expert-informed pipeline combining discharge-summary structuring schema, expert-curated multi-turn QA templates, and LLM-based generation, followed by review and revision of every single QA sample by 11 medical experts. Benchmarking 22 open- and closed-source LLMs reveals several challenges, including that LLMs struggle more with evidence grounding than content answering, multi-turn errors compound across turns, and single-turn clinical QA performance does not reliably transfer to this setting. These findings establish EHRNote-ChatQA as a rigorous and practical benchmark for evaluating clinical QA systems. The dataset will be made publicly available through PhysioNet credentialed access.

cs.CL

Bootstrapping Open Quantum Many-body Systems with Absorbing Phase Transitions

We demonstrate that combining the positivity of density matrices with steady-state conditions yields a systematic bootstrap method for studying open quantum many-body systems governed by Lindblad master equations on infinite lattices, which exhibit absorbing phase transitions. As a concrete example, we apply this method to the quantum contact process with an absorbing state. We obtain bootstrap bounds on steady-state expectation values, the critical coupling, certain ratios of expectation values in the nontrivial steady state in the supercritical phase, and the Liouvillian spectral gap in the subcritical phase.

quant-ph

Intrinsic Reward Policy Optimization for Sparse-Reward Environments

Exploration is essential in reinforcement learning as an agent relies on trial and error to learn an optimal policy. However, when rewards are sparse, naive exploration strategies, like noise injection, are often insufficient. Intrinsic rewards can also provide principled guidance for exploration by, for example, combining them with extrinsic rewards to optimize a policy or using them to train subpolicies for hierarchical learning. However, the former approach suffers from unstable credit assignment, while the latter exhibits sample inefficiency and sub-optimality. We propose a policy optimization framework that leverages multiple intrinsic rewards to directly optimize a policy for an extrinsic reward without pretraining subpolicies. Our algorithm -- intrinsic reward policy optimization (IRPO) -- achieves this by using a surrogate policy gradient that provides a more informative learning signal than the true gradient in sparse-reward environments. We demonstrate that IRPO improves performance and sample efficiency relative to baselines in discrete and continuous environments, and formally analyze the optimization problem solved by IRPO. Our code is available at https://github.com/Mgineer117/IRPO.

cs.LG

On the Flux Sectors of Matrix String Theory

We analyze the gapped flux vacua of 2D (8,8) SU(N) super-Yang-Mills theory. Based on the matrix string theory duality, we conjecture the spectrum of massive resonances in the gauge theory at large N and beyond the 't Hooft scaling regime.

hep-th

Bootstrapping Euclidean Two-point Correlators

We develop a bootstrap approach to Euclidean two-point correlators, in the thermal or ground state of quantum mechanical systems. We formulate the problem of bounding the two-point correlator as a semidefinite programming problem, subject to the constraints of reflection positivity, the Heisenberg equations of motion, and the Kubo-Martin-Schwinger condition or ground-state positivity. In the dual formulation, the Heisenberg equations of motion become "inequalities of motion" on the Lagrange multipliers that enforce the constraints. This enables us to derive rigorous bounds on continuous-time two-point correlators using a finite-dimensional semidefinite or polynomial matrix program. We illustrate this method by bootstrapping the two-point correlators of the ungauged one-matrix quantum mechanics, from which we extract the spectrum and matrix elements of the low-lying adjoint states. Along the way, we provide a new derivation of the energy-entropy balance inequality and establish a connection between the high-temperature two-point correlator bootstrap and the matrix integral bootstrap.

hep-th

Nonequilibrium Phase Transitions in Large $N$ Matrix Quantum Mechanics

It is believed that the theory of quantum gravity describing our universe is unitary. Nonetheless, if we only have access to a subsystem, its dynamics is described by nonequilibrium physics. Motivated by this, we investigate the planar limit of large $N$ ungauged one-matrix quantum mechanics obeying the Lindblad master equation with dissipative jump terms, focusing on the existence, uniqueness, and properties of steady states that signal nonequilibrium phase transitions. In the first class of examples, where potentials are unbounded from below, we study nonequilibrium critical points above which strong dissipation allows for the existence of normalizable steady states that would otherwise not exist. In the second class of examples, termed matrix quantum optics, we find evidence of nonequilibrium phase transitions analogous to those recently reported in the quantum optics literature for driven-dissipative Kerr resonators. Preliminary results on two-matrix quantum mechanics are also presented. We implement bootstrap methods to obtain concrete and rigorous results for the nonequilibrium steady states of matrix quantum mechanics in the planar limit.

hep-th

On the $AdS_5\times S^5$ Solution of Superstring Field Theory

We determine the $AdS_5\times S^5$ solution of type IIB superstring field theory (SFT) to the third order in the expansion with respect to Ramond-Ramond (RR) flux, demonstrate its supersymmetry from the SFT gauge transformations, and identify the massless RR axion in the spectrum of linearized fluctuations. We present an all-order solution in the pp-wave limit and comment on potential obstructions to, and the existence of, the all-order $AdS_5\times S^5$ solution.

hep-th

Contraction-Aware Reinforcement Learning for Nonlinear Control with Statistical Robustness

Control contraction metrics (CCMs)-defined by Riemannian metrics under which a closed-loop system is incrementally exponentially stable-offer a constructive framework for synthesizing contracting policies in nonlinear path-tracking problems. However, while the synthesized policies ensure pointwise satisfaction of the CCM conditions, they may not ensure long-term optimality (i.e., minimizing cumulative trajectory-level tracking error) over both transient and steady-state regimes. Furthermore, the myopic nature of these policies could also make them more susceptible to learning biases when approximate dynamics are used to formulate CCMs. To address these issues, we propose to integrate CCMs into reinforcement learning (RL). CCMs provide dynamics-informed feedback for learning a policy that has a stability guarantee-i.e., is contraction-aware-while RL provides a framework for minimizing cumulative tracking error under approximate dynamics. Given a pretrained dynamics model, our algorithm, contraction-aware RL (CARL), simultaneously learns to generate CCMs and optimize a policy for rewards defined by those CCMs. We demonstrate that CARL enhances path-tracking performance and is robust to errors in approximated dynamics compared to relevant baselines in both simulated and real-world robot experiments. We also provide theoretical rationale for integrating CCMs into RL. Our code is available at https://github.com/Mgineer117/CARL, and a video of our real-world robot experiments can be found at https://youtu.be/sOJ4hulbop0.

cs.LG

Bootstrapping Nonequilibrium Stochastic Processes

We show that bootstrap methods based on the positivity of probability measures provide a systematic framework for studying both synchronous and asynchronous nonequilibrium stochastic processes on infinite lattices. First, we formulate linear programming problems that use positivity and invariance property of invariant measures to derive rigorous bounds on their expectation values. Second, for time evolution in asynchronous processes, we exploit the master equation along with positivity and initial conditions to construct linear and semidefinite programming problems that yield bounds on expectation values at both short and late times. We illustrate both approaches using two canonical examples: the contact process in 1+1 and 2+1 dimensions, and the Domany-Kinzel model in both synchronous and asynchronous forms in 1+1 dimensions. Our bounds on invariant measures yield rigorous lower bounds on critical rates, while those on time evolutions provide two-sided bounds on the half-life of the infection density and the temporal correlation length in the subcritical phase.

cond-mat.stat-mech

Hierarchical Meta-Reinforcement Learning via Automated Macro-Action Discovery

Meta-Reinforcement Learning (Meta-RL) enables fast adaptation to new testing tasks. Despite recent advancements, it is still challenging to learn performant policies across multiple complex and high-dimensional tasks. To address this, we propose a novel architecture with three hierarchical levels for 1) learning task representations, 2) discovering task-agnostic macro-actions in an automated manner, and 3) learning primitive actions. The macro-action can guide the low-level primitive policy learning to more efficiently transition to goal states. This can address the issue that the policy may forget previously learned behavior while learning new, conflicting tasks. Moreover, the task-agnostic nature of the macro-actions is enabled by removing task-specific components from the state space. Hence, this makes them amenable to re-composition across different tasks and leads to promising fast adaptation to new tasks. Also, the prospective instability from the tri-level hierarchies is effectively mitigated by our innovative, independently tailored training schemes. Experiments in the MetaWorld framework demonstrate the improved sample efficiency and success rate of our approach compared to previous state-of-the-art methods.

cs.LG

Coarse-grained Bootstrap of Quantum Many-body Systems

We present a new computational framework combining coarse-graining techniques with bootstrap methods to study quantum many-body systems. The method efficiently computes rigorous upper and lower bounds on both zero- and finite-temperature expectation values of any local observables of infinite quantum spin chains. This is achieved by using tensor networks to coarse-grain bootstrap constraints, including positivity, translation invariance, equations of motion, and energy-entropy balance inequalities. Coarse-graining allows access to constraints from significantly larger subsystems than previously possible, yielding tighter bounds compared to those obtained without coarse-graining.

hep-th

Thermal Bootstrap of Matrix Quantum Mechanics

We implement a bootstrap method that combines stationary state conditions, thermal inequalities, and semidefinite relaxations of matrix logarithm in the ungauged one-matrix quantum mechanics, at finite rank N as well as in the large N limit, and determine finite temperature observables that interpolate between available analytic results in the low and high temperature limits respectively. We also obtain bootstrap bounds on thermal phase transition as well as preliminary results in the ungauged two-matrix quantum mechanics.

hep-th

Sparsity-based Safety Conservatism for Constrained Offline Reinforcement Learning

Reinforcement Learning (RL) has made notable success in decision-making fields like autonomous driving and robotic manipulation. Yet, its reliance on real-time feedback poses challenges in costly or hazardous settings. Furthermore, RL's training approach, centered on "on-policy" sampling, doesn't fully capitalize on data. Hence, Offline RL has emerged as a compelling alternative, particularly in conducting additional experiments is impractical, and abundant datasets are available. However, the challenge of distributional shift (extrapolation), indicating the disparity between data distributions and learning policies, also poses a risk in offline RL, potentially leading to significant safety breaches due to estimation errors (interpolation). This concern is particularly pronounced in safety-critical domains, where real-world problems are prevalent. To address both extrapolation and interpolation errors, numerous studies have introduced additional constraints to confine policy behavior, steering it towards more cautious decision-making. While many studies have addressed extrapolation errors, fewer have focused on providing effective solutions for tackling interpolation errors. For example, some works tackle this issue by incorporating potential cost-maximizing optimization by perturbing the original dataset. However, this, involving a bi-level optimization structure, may introduce significant instability or complicate problem-solving in high-dimensional tasks. This motivates us to pinpoint areas where hazards may be more prevalent than initially estimated based on the sparsity of available data by providing significant insight into constrained offline RL. In this paper, we present conservative metrics based on data sparsity that demonstrate the high generalizability to any methods and efficacy compared to using bi-level cost-ub-maximization.

cs.LG

Out-of-Distribution Adaptation in Offline RL: Counterfactual Reasoning via Causal Normalizing Flows

Despite notable successes of Reinforcement Learning (RL), the prevalent use of an online learning paradigm prevents its widespread adoption, especially in hazardous or costly scenarios. Offline RL has emerged as an alternative solution, learning from pre-collected static datasets. However, this offline learning introduces a new challenge known as distributional shift, degrading the performance when the policy is evaluated on scenarios that are Out-Of-Distribution (OOD) from the training dataset. Most existing offline RL resolves this issue by regularizing policy learning within the information supported by the given dataset. However, such regularization overlooks the potential for high-reward regions that may exist beyond the dataset. This motivates exploring novel offline learning techniques that can make improvements beyond the data support without compromising policy performance, potentially by learning causation (cause-and-effect) instead of correlation from the dataset. In this paper, we propose the MOOD-CRL (Model-based Offline OOD-Adapting Causal RL) algorithm, which aims to address the challenge of extrapolation for offline policy training through causal inference instead of policy-regularizing methods. Specifically, Causal Normalizing Flow (CNF) is developed to learn the transition and reward functions for data generation and augmentation in offline policy evaluation and training. Based on the data-invariant, physics-based qualitative causal graph and the observational data, we develop a novel learning scheme for CNF to learn the quantitative structural causal model. As a result, CNF gains predictive and counterfactual reasoning capabilities for sequential decision-making tasks, revealing a high potential for OOD adaptation. Our CNF-based offline RL approach is validated through empirical evaluations, outperforming model-free and model-based methods by a significant margin.

cs.LG

Constrained Meta-Reinforcement Learning for Adaptable Safety Guarantee with Differentiable Convex Programming

Despite remarkable achievements in artificial intelligence, the deployability of learning-enabled systems in high-stakes real-world environments still faces persistent challenges. For example, in safety-critical domains like autonomous driving, robotic manipulation, and healthcare, it is crucial not only to achieve high performance but also to comply with given constraints. Furthermore, adaptability becomes paramount in non-stationary domains, where environmental parameters are subject to change. While safety and adaptability are recognized as key qualities for the new generation of AI, current approaches have not demonstrated effective adaptable performance in constrained settings. Hence, this paper breaks new ground by studying the unique challenges of ensuring safety in non-stationary environments by solving constrained problems through the lens of the meta-learning approach (learning-to-learn). While unconstrained meta-learning al-ready encounters complexities in end-to-end differentiation of the loss due to the bi-level nature, its constrained counterpart introduces an additional layer of difficulty, since the constraints imposed on task-level updates complicate the differentiation process. To address the issue, we first employ successive convex-constrained policy updates across multiple tasks with differentiable convexprogramming, which allows meta-learning in constrained scenarios by enabling end-to-end differentiation. This approach empowers the agent to rapidly adapt to new tasks under non-stationarity while ensuring compliance with safety constraints.

cs.AI

A Worldsheet Description of Flux Compactifications

We demonstrate how recent developments in string field theory provide a framework to systematically study type II flux compactifications with non-trivial Ramond-Ramond profiles. We present an explicit example where physical observables can be computed order by order in a small parameter which can be effectively viewed as string coupling constant. We obtain the corresponding background solution of the string field equations of motions up to the second order in the expansion. Along the way, we show how the tadpole cancellations of the string field equations lead to the minimization of the F-term potential of the low energy supergravity description. String field action expanded around the obtained background solution furnishes a worldsheet description of the flux compactifications.

hep-th

Rolling tachyon and the Phase Space of Open String Field Theory

We construct the symplectic form on the covariant phase space of the open string field theory on a ZZ-brane in c=1 string theory, and determine the energy of the rolling tachyon solution, confirming Sen's earlier proposal based on boundary conformal field theory and closed string considerations.

hep-th

Bootstrapping the Stochastic Resonance

Stochastic resonance is a phenomenon where a noise of appropriate intensity enhances the input signal strength. In this work, by employing the recently developed convex optimization methods in the context of dynamical systems and stochastic processes, we derive rigorous two-sided bounds on the expected power at the input signal frequency for the prototypical example of stochastic resonance, the double-well potential with periodic forcing and Gaussian white noise.

math.DS