SearcharxivSearch

arXiv subjects

Hanwool Lee

Publications and source records attributed to Hanwool Lee.

At least 19 recordsLinked to original sources

Perfect Discrimination of Non-Orthogonal Quantum States via Adaptive Post-Measurement Queries

A set of pairwise non-orthogonal quantum states cannot be perfectly discriminated, and this remains true even when one bit of classical partial information is available prior to the measurement. Contrary to the usual intuition that earlier information is at least as valuable as later information, we show that the same bit can be more useful when it arrives after the measurement. We present a framework for state discrimination in which the sender provides classical information in response to a request from the receiver, which we refer to as a query. In some cases, pairwise non-orthogonal states can be perfectly discriminated when the query is allowed to depend on the measurement outcome. We give a general method for finding the optimal strategy in this setting, which turns out to be the standard minimum-error discrimination problem for an auxiliary ensemble.

quant-ph

NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms

Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness under sustained, adaptive adversarial pressure remains poorly characterized. We present NRT-Bench, a benchmark for multi-turn red-teaming of LLM agents acting as operators of a safety-critical system, instantiated in a simulated nuclear power plant control room. A five-role operator team, each backed by a configurable LLM, runs a plant governed by six critical safety functions (CSFs), while adversaries inject messages over four channels in bounded multi-turn sessions with per-turn feedback. Harm is an objective signal rather than LLM-judged text: a run terminates the moment any CSF is lost, attributed to the causing message. Evaluating four frontier operator models under a fixed-attack paired-replay protocol, we find that adaptive multi-turn attacks reliably push the operator team past a safety limit: across the four models, between 8.7% and 12.1% of attack sessions end with the plant losing a critical safety function. Although the four models look almost equally robust by this aggregate rate, their failures barely overlap: of $149$ sessions, none defeat all four models while a third defeat at least one, so vulnerabilities are nearly disjoint across models rather than nested. The effect of added defences is strongly model-dependent: the same guardrail stack or safety-advisor agent that lowers attack success for one model can raise it for another. We release the simulation venue, attack dataset, and replay tooling for reproducible safety evaluation of LLM agents.

cs.CR

Semi-Device-Independent Certification for Nonlocality without Entanglement

In this work, we present the framework for demonstrating and certifying the distinction between measurements in an entangled basis, called global measurements, and local operations and classical communication (LOCC), known as nonlocality without entanglement (NLWE). To be precise, we show NLWE via a maximum-confidence measurement in terms of a guess per detection event, called a confidence, a fine-grained guessing probability that encompasses both minimum-error and unambiguous state-discrimination strategies. We show that NLWE for unknown measurements can be certified, given the measurement-outcome rates, by bounding the confidence using LOCC; the certification is semi-device-independent in that state preparation is trusted. We illustrate the demonstration and certification of NLWE for antiparallel qubit states. Our results make it feasible to experimentally realize NLWE using measurement devices with imperfections, such as non-unit detection efficiency, since maximum-confidence measurements rely only on detected events.

quant-ph

Nonclassical traits in multi-copy state discrimination

Quantum state discrimination is a fundamental information processing task that serves as a key component in many applications while also carrying foundational significance. In this work, we consider minimum error discrimination of multi-copy states, where instead of preparing a single system we assume that multiple instances of the same state are prepared. Now the discrimination allows for measurements from multiple parties with different measurement strategies varying from global measurement strategy to ones restricted to different forms of local operations and classical communication strategies. By comparing the average success probabilities in quantum and classical cases, we find a qubit strategy that outperforms all the bit strategies. On the other hand, we show that the classical measurement strategy does not give benefit in qubit over bit. However, we find that there are other (qu)bit-like operational theories which can outperform the best qubit strategies even with a classical measurement strategy and we are able to identify instances of different theories where different measurement strategies are optimal. In this way, we are able to find instances of nonlocality without entanglement as well as provide general bounds for bit-like operational theories.

quant-ph

What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models

Current vision-language benchmarks predominantly feature well-structured questions with clear, explicit prompts. However, real user queries are often informal and underspecified. Users naturally leave much unsaid, relying on images to convey context. We introduce HAERAE-Vision, a benchmark of 653 real-world visual questions from Korean online communities (0.76% survival from 86K candidates), each paired with an explicit rewrite, yielding 1,306 query variants in total. Evaluating 39 VLMs, we find that even state-of-the-art models (GPT-5, Gemini 2.5 Pro) achieve under 50% on the original queries. Crucially, query explicitation alone yields 8 to 22 point improvements, with smaller models benefiting most. We further show that even with web search, under-specified queries underperform explicit queries without search, revealing that current retrieval cannot compensate for what users leave unsaid. Our findings demonstrate that a substantial portion of VLM difficulty stem from natural query under-specification instead of model capability, highlighting a critical gap between benchmark evaluation and real-world deployment.

cs.CV

Sharing quantum indistinguishability with multiple parties

Quantum indistinguishability of non-orthogonal quantum states is a valuable resource in quantum information applications such as cryptography and randomness generation. In this article, we present a sequential state-discrimination scheme that enables multiple parties to share quantum uncertainty, in terms of the max relative entropy, generated by a single party. Our scheme is based upon maximum-confidence measurements and takes advantages of weak measurements to allow a number of parties to perform state discrimination on a single quantum system. We review known sequential state discrimination and show how our scheme would work through a number of examples where ensembles may or may not contain symmetries. Our results will have a role to play in understanding the ultimate limits of sequential information extraction and guide the development of quantum resource sharing in sequential settings.

quant-ph

AI PB: A Grounded Generative Agent for Personalized Investment Insights

We present AI PB, a production-scale generative agent deployed in real retail finance. Unlike reactive chatbots that answer queries passively, AI PB proactively generates grounded, compliant, and user-specific investment insights. It integrates (i) a component-based orchestration layer that deterministically routes between internal and external LLMs based on data sensitivity, (ii) a hybrid retrieval pipeline using OpenSearch and the finance-domain embedding model, and (iii) a multi-stage recommendation mechanism combining rule heuristics, sequential behavioral modeling, and contextual bandits. Operating fully on-premises under Korean financial regulations, the system employs Docker Swarm and vLLM across 24 X NVIDIA H100 GPUs. Through human QA and system metrics, we demonstrate that grounded generation with explicit routing and layered safety can deliver trustworthy AI insights in high-stakes finance.

cs.AI

Sequential Semi-Device-Independent Quantum Randomness Certification

Quantum measurements under realistic conditions reveal only partial information about a system. Yet, by performing sequential measurements on the same system, additional information can be accessed. We investigate this problem in the context of semi-device-independent randomness certification using sequential maximum confidence measurements. We develop a general framework and versatile numerical methods to bound the amount of certifiable randomness in such scenarios. We further introduce a technique to compute min-tradeoff functions via semidefinite programming duality, thus making the framework suitable for bounding the certifiable randomness against adaptive attacking strategies through entropy accumulation. Our results establish sufficient criteria showing that maximum confidence measurements enable the distribution and certification of randomness across a sequential measurement chain.

quant-ph

NMIXX: Domain-Adapted Neural Embeddings for Cross-Lingual eXploration of Finance

General-purpose sentence embedding models often struggle to capture specialized financial semantics, especially in low-resource languages like Korean, due to domain-specific jargon, temporal meaning shifts, and misaligned bilingual vocabularies. To address these gaps, we introduce NMIXX (Neural eMbeddings for Cross-lingual eXploration of Finance), a suite of cross-lingual embedding models fine-tuned with 18.8K high-confidence triplets that pair in-domain paraphrases, hard negatives derived from a semantic-shift typology, and exact Korean-English translations. Concurrently, we release KorFinSTS, a 1,921-pair Korean financial STS benchmark spanning news, disclosures, research reports, and regulations, designed to expose nuances that general benchmarks miss. When evaluated against seven open-license baselines, NMIXX's multilingual bge-m3 variant achieves Spearman's rho gains of +0.10 on English FinSTS and +0.22 on KorFinSTS, outperforming its pre-adaptation checkpoint and surpassing other models by the largest margin, while revealing a modest trade-off in general STS performance. Our analysis further shows that models with richer Korean token coverage adapt more effectively, underscoring the importance of tokenizer design in low-resource, cross-lingual settings. By making both models and the benchmark publicly available, we provide the community with robust tools for domain-adapted, multilingual representation learning in finance.

cs.CL

Metainformation in Quantum Guessing Games

Quantum guessing games offer a structured approach to analyzing quantum information processing, where information is encoded in quantum states and extracted through measurement. An additional aspect of this framework is the influence of partial knowledge about the input on the optimal measurement strategies. This kind of side information can significantly influence the guessing strategy and earlier work has shown that the timing of such side information, whether revealed before or after the measurement, can affect the success probabilities. In this work, we go beyond this established distinction by introducing the concept of metainformation. Metainformation is information about information, and in our context it is knowledge that additional side information of certain type will become later available, even if it is not yet provided. We show that this seemingly subtle difference between having no expectation of further information versus knowing it will arrive can have operational consequences for the guessing task. Our results demonstrate that metainformation can, in certain scenarios, enhance the achievable success probability up to the point that post-measurement side information becomes as useful as prior-measurement side information, while in others it offers no benefit. By formally distinguishing metainformation from actual side information, we uncover a finer structure in the interplay between timing, information, and strategy, offering new insights into the capabilities of quantum systems in information processing tasks.

quant-ph

Redefining Evaluation Standards: A Unified Framework for Evaluating the Korean Capabilities of Language Models

Recent advancements in Korean large language models (LLMs) have driven numerous benchmarks and evaluation methods, yet inconsistent protocols cause up to 10 p.p performance gaps across institutions. Overcoming these reproducibility gaps does not mean enforcing a one-size-fits-all evaluation. Rather, effective benchmarking requires diverse experimental approaches and a framework robust enough to support them. To this end, we introduce HRET (Haerae Evaluation Toolkit), an open-source, registry-based framework that unifies Korean LLM assessment. HRET integrates major Korean benchmarks, multiple inference backends, and multi-method evaluation, with language consistency enforcement to ensure genuine Korean outputs. Its modular registry design also enables rapid incorporation of new datasets, methods, and backends, ensuring the toolkit adapts to evolving research needs. Beyond standard accuracy metrics, HRET incorporates Korean-focused output analyses-morphology-aware Type-Token Ratio (TTR) for evaluating lexical diversity and systematic keyword-omission detection for identifying missing concepts-to provide diagnostic insights into language-specific behaviors. These targeted analyses help researchers pinpoint morphological and semantic shortcomings in model outputs, guiding focused improvements in Korean LLM development.

cs.CE

(G)I-DLE: Generative Inference via Distribution-preserving Logit Exclusion with KL Divergence Minimization for Constrained Decoding

We propose (G)I-DLE, a new approach to constrained decoding that leverages KL divergence minimization to preserve the intrinsic conditional probability distribution of autoregressive language models while excluding undesirable tokens. Unlike conventional methods that naively set banned tokens' logits to $-\infty$, which can distort the conversion from raw logits to posterior probabilities and increase output variance, (G)I-DLE re-normalizes the allowed token probabilities to minimize such distortion. We validate our method on the K2-Eval dataset, specifically designed to assess Korean language fluency, logical reasoning, and cultural appropriateness. Experimental results on Qwen2.5 models (ranging from 1.5B to 14B) demonstrate that G-IDLE not only boosts mean evaluation scores but also substantially reduces the variance of output quality.

cs.CE

TWICE: What Advantages Can Low-Resource Domain-Specific Embedding Model Bring? -- A Case Study on Korea Financial Texts

Domain specificity of embedding models is critical for effective performance. However, existing benchmarks, such as FinMTEB, are primarily designed for high-resource languages, leaving low-resource settings, such as Korean, under-explored. Directly translating established English benchmarks often fails to capture the linguistic and cultural nuances present in low-resource domains. In this paper, titled TWICE: What Advantages Can Low-Resource Domain-Specific Embedding Models Bring? A Case Study on Korea Financial Texts, we introduce KorFinMTEB, a novel benchmark for the Korean financial domain, specifically tailored to reflect its unique cultural characteristics in low-resource languages. Our experimental results reveal that while the models perform robustly on a translated version of FinMTEB, their performance on KorFinMTEB uncovers subtle yet critical discrepancies, especially in tasks requiring deeper semantic understanding, that underscore the limitations of direct translation. This discrepancy highlights the necessity of benchmarks that incorporate language-specific idiosyncrasies and cultural nuances. The insights from our study advocate for the development of domain-specific evaluation frameworks that can more accurately assess and drive the progress of embedding models in low-resource settings.

cs.CL

Sequential Quantum Maximum Confidence Discrimination

Sequential quantum information processing may lie in the peaceful coexistence of no-go theorems on quantum operations, such as the no-cloning theorem, the monogamy of correlations, and the no-signalling principle. In this work, we investigate a sequential scenario of quantum state discrimination with maximum confidence, called maximum-confidence discrimination, which generalizes other strategies including minimum-error and unambiguous state discrimination. We show that sequential state discrimination with equally high confidence can be realized only when positive-operator-valued measure elements for a maximum-confidence measurement are linearly independent; otherwise, a party will have strictly less confidence in measurement outcomes than the previous one. We establish a tradeoff between the disturbance of states and information gain in sequential state discrimination, namely, that the less a party learn in state discrimination in terms of a guessing probability, the more parties can participate in the sequential scenario.

quant-ph

IVE: Enhanced Probabilistic Forecasting of Intraday Volume Ratio with Transformers

This paper presents a new approach to volume ratio prediction in financial markets, specifically targeting the execution of Volume-Weighted Average Price (VWAP) strategies. Recognizing the importance of accurate volume profile forecasting, our research leverages the Transformer architecture to predict intraday volume ratio at a one-minute scale. We diverge from prior models that use log-transformed volume or turnover rates, instead opting for a prediction model that accounts for the intraday volume ratio's high variability, stabilized via log-normal transformation. Our input data incorporates not only the statistical properties of volume but also external volume-related features, absolute time information, and stock-specific characteristics to enhance prediction accuracy. The model structure includes an encoder-decoder Transformer architecture with a distribution head for greedy sampling, optimizing performance on high-liquidity stocks across both Korean and American markets. We extend the capabilities of our model beyond point prediction by introducing probabilistic forecasting that captures the mean and standard deviation of volume ratios, enabling the anticipation of significant intraday volume spikes. Furthermore, an agent with a simple trading logic demonstrates the practical application of our model through live trading tests in the Korean market, outperforming VWAP benchmarks over a period of two and a half months. Our findings underscore the potential of Transformer-based probabilistic models for volume ratio prediction and pave the way for future research advancements in this domain.

q-fin.CP

ML-Promise: A Multilingual Dataset for Corporate Promise Verification

Promises made by politicians, corporate leaders, and public figures have a significant impact on public perception, trust, and institutional reputation. However, the complexity and volume of such commitments, coupled with difficulties in verifying their fulfillment, necessitate innovative methods for assessing their credibility. This paper introduces the concept of Promise Verification, a systematic approach involving steps such as promise identification, evidence assessment, and the evaluation of timing for verification. We propose the first multilingual dataset, ML-Promise, which includes English, French, Chinese, Japanese, and Korean, aimed at facilitating in-depth verification of promises, particularly in the context of Environmental, Social, and Governance (ESG) reports. Given the growing emphasis on corporate environmental contributions, this dataset addresses the challenge of evaluating corporate promises, especially in light of practices like greenwashing. Our findings also explore textual and image-based baselines, with promising results from retrieval-augmented generation (RAG) approaches. This work aims to foster further discourse on the accountability of public commitments across multiple languages and domains.

cs.CL

Strong Damping-Like Torques in Wafer-Scale MoTe${}_2$ Grown by MOCVD

The scalable synthesis of strong spin orbit coupling (SOC) materials such as 1T${}^\prime$ phase MoTe${}_2$ is crucial for spintronics development. Here, we demonstrate wafer-scale growth of 1T${}^\prime$ MoTe${}_2$ using metal-organic chemical vapor deposition (MOCVD) with sputtered Mo and (C${}_4$H${}_9$)${}_2$Te. The synthesized films show uniform coverage across the entire sample surface. By adjusting the growth parameters, a synthesis process capable of producing 1T${}^\prime$ and 2H MoTe${}_2$ mixed phase films was achieved. Notably, the developed process is compatible with back-end-of-line (BEOL) applications. The strong spin-orbit coupling of the grown 1T${}^\prime$ MoTe${}_2$ films was demonstrated through spin torque ferromagnetic resonance (ST-FMR) measurements conducted on a 1T${}^\prime$ MoTe${}_2$/permalloy bilayer RF waveguide. These measurements revealed a significant damping-like torque in the wafer-scale 1T${}^\prime$ MoTe${}_2$ film and indicated high spin-charge conversion efficiency. The BEOL compatible process and potent spin orbit torque demonstrate promise in advanced device applications.

cond-mat.mtrl-sci

KMMLU: Measuring Massive Multitask Language Understanding in Korean

We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM. While prior Korean benchmarks are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language. We test 27 public and proprietary LLMs and observe the best public model to score 50.5%, leaving significant room for improvement. This model was primarily trained for English and Chinese, not Korean. Current LLMs tailored to Korean, such as Polyglot-Ko, perform far worse. Surprisingly, even the most capable proprietary LLMs, e.g., GPT-4 and HyperCLOVA X do not exceed 60%. This suggests that further work is needed to improve LLMs for Korean, and we believe KMMLU offers the appropriate tool to track this progress. We make our dataset publicly available on the Hugging Face Hub and integrate the benchmark into EleutherAI's Language Model Evaluation Harness.

cs.CL