SearcharxivSearch

arXiv subjects

Kaiyuan Ji

Publications and source records attributed to Kaiyuan Ji.

16 recordsLinked to original sources

Research-Oriented Human-Centric Evaluation for Foundation Models

Most current evaluations of foundation models focus on objective benchmarks, such as knowledge coverage and reasoning accuracy, often overlooking users' subjective experiences in human-AI collaboration. To address this gap, we propose a research-oriented Human-Centric Evaluation framework. It captures user perceptions across three core dimensions: problem-solving ability, information quality, and interaction experience, providing a structured, fine-grained approach to understanding how users evaluate and respond to model behavior in multi-modal research contexts. We conduct 604 human evaluation sessions across various disciplines, involving recent advanced foundation models. Through open-ended, time-limited collaborative tasks, we gather rich subjective assessments that highlight model capabilities and user preferences. Additionally, we perform an LLM-as-a-judge experiment and find that even sophisticated models struggle to accurately replicate human subjective judgment, emphasizing the irreplaceable value of first-person human assessment. Our project link is https://github.com/yijinguo/Human-Centric-Evaluation.

cs.CL

Fundamental limits for thermodynamic control with quantum feedback

The study of feedback control inspired by Maxwell's demon is central to the understanding of the relationship between thermodynamics and information. In this paper, we establish fundamental lower limits on the work costs of system conversion with quantum feedback, where quantum side information acquired in advance can be fed back to the system coherently by a controller. From two basic operational principles that every physically admissible feedback-control scheme should satisfy, we derive the tightest possible bounds on the single-shot work of formation and extractable work of an arbitrary quantum system given arbitrary quantum side information held by the controller. These bounds are expressed in terms of information measures simultaneously generalizing conditional entropies, relative entropies, and mutual informations. In the asymptotic limit, we derive a generalized second law of thermodynamics with quantum feedback, featuring a conditional Helmholtz free energy, and we further show that it does not contradict the traditional second law. Our findings also provide precise thermodynamic meanings for the negativity of single-shot conditional entropies and resolve an open problem in the axiomatic reconstruction of such conditional entropies.

quant-ph

Retrocausal capacity of a quantum channel: Communicating through noisy closed timelike curves

We study the capacity of a quantum channel for retrocausal communication, where messages are transmitted backward in time, from a sender in the future to a receiver in the past, through a noisy postselected closed timelike curve mathematically represented by the channel. We completely characterize the one-shot retrocausal quantum and classical capacities, and we show that the corresponding asymptotic capacities are equal to the average and sum, respectively, of the channel's max-information and its regularized Doeblin information. This endows these information measures with a novel operational interpretation. Furthermore, our characterization can be generalized beyond quantum channels to all completely positive maps. This imposes information-theoretic limits on transmitting messages via postselected-teleportation-like mechanisms with arbitrary initial- and final-state boundary conditions, including those considered in various black-hole final-state models.

quant-ph

Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating

Prior work has shown that fine-tuning large language models on malicious or incorrect outputs in narrow domains can induce broad misalignment and harmful behavior, a phenomenon known as emergent misalignment. However, efficient methods for reversing such misalignment remain limited. In this work, we make two contributions. First, we identify sycophancy fine-tuning, i.e., training models to passively agree with users' incorrect opinions, as a previously underexplored driver of emergent misalignment, and show that it induces broad and severe misaligned behavior. Second, we propose Alignment Gating, an efficient method for reversing emergent misalignment that inserts learnable and controllable gates into the model during fine-tuning. Through fine-tuning, these gates learn to identify the internal representations responsible for unsafe responses. Thus, amplifying or suppressing these representations then exacerbates or mitigates EM, respectively. We further find that alignment gating module exhibits strong generalization: gating weights obtained from narrow-domain fine-tuning substantially suppress broad-domain misaligned behavior while preserving the model's general capabilities.

cs.CL

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond

The success of large language models (LLMs) in scientific domains has heightened safety concerns, prompting numerous benchmarks to evaluate their scientific safety. Existing benchmarks often suffer from limited risk coverage and a reliance on subjective evaluation. To address these problems, we introduce SafeSci, a comprehensive framework for safety evaluation and enhancement in scientific contexts. SafeSci comprises SafeSciBench, a multi-disciplinary benchmark with 0.25M samples, and SafeSciTrain, a large-scale dataset containing 1.5M samples for safety enhancement. SafeSciBench distinguishes between safety knowledge and risk to cover extensive scopes and employs objective metrics such as deterministically answerable questions to mitigate evaluation bias. We evaluate 24 advanced LLMs, revealing critical vulnerabilities in current models. We also observe that LLMs exhibit varying degrees of excessive refusal behaviors on safety-related issues. For safety enhancement, we demonstrate that fine-tuning on SafeSciTrain significantly enhances the safety alignment of models. Finally, we argue that knowledge is a double-edged sword, and determining the safety of a scientific question should depend on specific context, rather than universally categorizing it as safe or unsafe. Our work provides both a diagnostic tool and a practical resource for building safer scientific AI systems.

cs.LG

A$^3$: Towards Advertising Aesthetic Assessment

Advertising images significantly impact commercial conversion rates and brand equity, yet current evaluation methods rely on subjective judgments, lacking scalability, standardized criteria, and interpretability. To address these challenges, we present A^3 (Advertising Aesthetic Assessment), a comprehensive framework encompassing four components: a paradigm (A^3-Law), a dataset (A^3-Dataset), a multimodal large language model (A^3-Align), and a benchmark (A^3-Bench). Central to A^3 is a theory-driven paradigm, A^3-Law, comprising three hierarchical stages: (1) Perceptual Attention, evaluating perceptual image signals for their ability to attract attention; (2) Formal Interest, assessing formal composition of image color and spatial layout in evoking interest; and (3) Desire Impact, measuring desire evocation from images and their persuasive impact. Building on A^3-Law, we construct A^3-Dataset with 120K instruction-response pairs from 30K advertising images, each richly annotated with multi-dimensional labels and Chain-of-Thought (CoT) rationales. We further develop A^3-Align, trained under A^3-Law with CoT-guided learning on A^3-Dataset. Extensive experiments on A^3-Bench demonstrate that A^3-Align achieves superior alignment with A^3-Law compared to existing models, and this alignment generalizes well to quality advertisement selection and prescriptive advertisement critique, indicating its potential for broader deployment. Dataset, code, and models can be found at: https://github.com/euleryuan/A3-Align.

cs.CV

Barycentric bounds on the error exponents of quantum hypothesis exclusion

Quantum state exclusion is an operational task with application to ontological interpretations of quantum states. In such a task, one is given a system whose state is randomly selected from a finite set, and the goal is to identify a state from the set that is not the true state of the system. An error occurs if and only if the state identified is the true state. In this paper, we study the optimal error probability of quantum state exclusion and its error exponent from an information-theoretic perspective. Our main finding is a single-letter upper bound on the error exponent of state exclusion given by the multivariate log-Euclidean Chernoff divergence, and we prove that this improves upon the best previously known upper bound. We also extend our analysis to quantum channel exclusion, and we establish a single-letter and efficiently computable upper bound on its error exponent, admitting the use of adaptive strategies. We derive both upper bounds, for state and channel exclusion, based on one-shot analysis and formulate them as a type of multivariate divergence measure called a barycentric Chernoff divergence. Moreover, our result on channel exclusion has implications in two important special cases. First, when there are two hypotheses, our result provides the first known efficiently computable upper bound on the error exponent of symmetric binary channel discrimination. Second, when all channels are classical, we show that our upper bound is achievable by a parallel strategy, thus solving the exact error exponent of classical channel exclusion.

quant-ph

PriceSeer: Evaluating Large Language Models in Real-Time Stock Prediction

Stock prediction, a subject closely related to people's investment activities in fully dynamic and live environments, has been widely studied. Current large language models (LLMs) have shown remarkable potential in various domains, exhibiting expert-level performance through advanced reasoning and contextual understanding. In this paper, we introduce PriceSeer, a live, dynamic, and data-uncontaminated benchmark specifically designed for LLMs performing stock prediction tasks. Specifically, PriceSeer includes 110 U.S. stocks from 11 industrial sectors, with each containing 249 historical data points. Our benchmark implements both internal and external information expansion, where LLMs receive extra financial indicators, news, and fake news to perform stock price prediction. We evaluate six cutting-edge LLMs under different prediction horizons, demonstrating their potential in generating investment strategies after obtaining accurate price predictions for different sectors. Additionally, we provide analyses of LLMs' suboptimal performance in long-term predictions, including the vulnerability to fake news and specific industries. The code and evaluation data will be open-sourced at https://github.com/BobLiang2113/PriceSeer.

q-fin.ST

Beyond Hoeffding and Chernoff: Trading conclusiveness for advantages in quantum hypothesis testing

The ultimate limits of quantum state discrimination are often thought to be captured by asymptotic bounds that restrict the achievable error probabilities, notably the quantum Chernoff and Hoeffding bounds. Here we study hypothesis testing protocols that are permitted a probability of producing an inconclusive discrimination outcome, and investigate their performance when this probability is suitably constrained. We show that even by allowing an arbitrarily small probability of inconclusiveness, the limits imposed by the quantum Hoeffding and Chernoff bounds can be significantly exceeded. This completely circumvents the conventional trade-offs between error exponents in hypothesis testing while incurring only a vanishingly small overhead over conventional approaches. Such improvements over standard state discrimination are robust and can be obtained even when an exponentially vanishing probability of inconclusive outcomes is demanded. Relaxing the constraints on the inconclusive probability can enable even larger advantages, but this comes at a price. We show a 'strong converse' property of this setting: targeting error exponents beyond those achievable with vanishing inconclusiveness necessarily forces the probability of inconclusive outcomes to converge to one. By exactly quantifying the rate of this convergence, we give a complete characterisation of the trade-offs between error exponents and rates of conclusive outcome probabilities. Overall, our results provide a comprehensive asymptotic picture of how the allowance for inconclusive measurement outcomes reshapes optimal quantum hypothesis testing.

quant-ph

Converse bounds for quantum hypothesis exclusion: A divergence-radius approach

Hypothesis exclusion is an information-theoretic task in which an experimenter aims at ruling out a false hypothesis from a finite set of known candidates, and an error occurs if and only if the hypothesis being ruled out is the ground truth. For the tasks of quantum state exclusion and quantum channel exclusion -- where hypotheses are represented by quantum states and quantum channels, respectively -- efficiently computable upper bounds on the asymptotic error exponents were established in a recent work of the current authors [Ji et al., arXiv:2407.13728 (2024)], where the derivation was based on nonasymptotic analysis. In this companion paper of our previous work, we provide alternative proofs for the same upper bounds on the asymptotic error exponents of quantum state and channel exclusion, but using a conceptually different approach from the one adopted in the previous work. Specifically, we apply strong converse results for asymmetric binary hypothesis testing to distinguishing an arbitrary ``dummy'' hypothesis from each of the concerned candidates. This leads to the desired upper bounds in terms of divergence radii via a geometrically inspired argument.

quant-ph

Chiral supersolid and dissipative time crystal in Rydberg-dressed Bose-Einstein condensates with Raman-induced spin-orbit coupling

Spin-orbit coupling (SOC) is one of the crucial factors that affect the chiral symmetry of matter by causing the spatial symmetry breaking of the system. We find that Raman-induced SOC can induce a chiral supersolid phase with a helical antiskyrmion lattice in balanced Rydberg-dressed two-component Bose-Einstein condensates (BECs) in a harmonic trap by modulating the Raman coupling strength. This is in stark contrast to the mirror symmetric supersolid phase containing skyrmion-antiskyrmion lattice pair for the case of Rashba SOC. Two ground-state phase diagrams are presented as a function of the Rydberg interaction and the Raman-induced SOC. It is shown that the interplay among Raman-induced SOC, Rydberg interactions, and nonlinear contact interactions favors rich ground-state structures, including half-quantum vortex phase, stripe supersolid phase, toroidal stripe phase with a central Anderson-Toulouse coreless vortex, checkerboard supersolid phase, mirror symmetric supersolid phase, chiral supersolid phase and standing-wave supersolid phase. In addition, the effects of rotation and in-plane quadrupole magnetic field on the ground state of the system are analyzed. In these two cases, the chiral supersolid phase is broken and the ground state tends to form a miscible phase. Furthermore, we demonstrate that when the initial state is a chiral supersolid phase the rotating harmonic trapped system sustains dissipative continuous time crystal by studying the rotational dynamic behaviors of the system.

cond-mat.quant-gas

Entropic and operational characterizations of dynamic quantum resources

We offer new methods for characterizing general closed and convex quantum resource theories, including dynamic ones, based on entropic concepts and operational tasks. We propose a resource-theoretic generalization of the quantum conditional min-entropy, termed the free conditional min-entropy (FCME), in the sense that it quantifies an observer's ``subjective'' degree of uncertainty about a quantum system given that the observer's information processing is limited to free operations of the resource theory. Using this generalized concept, we provide a complete set of entropic conditions for free convertibility between quantum states or channels in any closed and convex quantum resource theory. We also derive an information-theoretic interpretation for the resource global robustness of a state or a channel in terms of a mutual-information-like quantity based on the FCME. Apart from this entropic approach, we characterize dynamic resources by also analyzing their performance in operational tasks. We construct operationally meaningful and complete sets of resource monotones with these tasks, which enable faithful tests of free convertibility between quantum channels. Finally, we show that every well-defined robustness-based measure of a channel can be interpreted as an operational advantage of the channel over free channels in a communication task.

quant-ph

MedOmni-45°: A Safety-Performance Benchmark for Reasoning-Oriented LLMs in Medicine

With the increasing use of large language models (LLMs) in medical decision-support, it is essential to evaluate not only their final answers but also the reliability of their reasoning. Two key risks are Chain-of-Thought (CoT) faithfulness -- whether reasoning aligns with responses and medical facts -- and sycophancy, where models follow misleading cues over correctness. Existing benchmarks often collapse such vulnerabilities into single accuracy scores. To address this, we introduce MedOmni-45 Degrees, a benchmark and workflow designed to quantify safety-performance trade-offs under manipulative hint conditions. It contains 1,804 reasoning-focused medical questions across six specialties and three task types, including 500 from MedMCQA. Each question is paired with seven manipulative hint types and a no-hint baseline, producing about 27K inputs. We evaluate seven LLMs spanning open- vs. closed-source, general-purpose vs. medical, and base vs. reasoning-enhanced models, totaling over 189K inferences. Three metrics -- Accuracy, CoT-Faithfulness, and Anti-Sycophancy -- are combined into a composite score visualized with a 45 Degrees plot. Results show a consistent safety-performance trade-off, with no model surpassing the diagonal. The open-source QwQ-32B performs closest (43.81 Degrees), balancing safety and accuracy but not leading in both. MedOmni-45 Degrees thus provides a focused benchmark for exposing reasoning vulnerabilities in medical LLMs and guiding safer model development.

cs.CV

Towards All-in-One Medical Image Re-Identification

Medical image re-identification (MedReID) is under-explored so far, despite its critical applications in personalized healthcare and privacy protection. In this paper, we introduce a thorough benchmark and a unified model for this problem. First, to handle various medical modalities, we propose a novel Continuous Modality-based Parameter Adapter (ComPA). ComPA condenses medical content into a continuous modality representation and dynamically adjusts the modality-agnostic model with modality-specific parameters at runtime. This allows a single model to adaptively learn and process diverse modality data. Furthermore, we integrate medical priors into our model by aligning it with a bag of pre-trained medical foundation models, in terms of the differential features. Compared to single-image feature, modeling the inter-image difference better fits the re-identification problem, which involves discriminating multiple images. We evaluate the proposed model against 25 foundation models and 8 large multi-modal language models across 11 image datasets, demonstrating consistently superior performance. Additionally, we deploy the proposed MedReID technique to two real-world applications, i.e., history-augmented personalized diagnosis and medical privacy protection. Codes and model is available at \href{https://github.com/tianyuan168326/All-in-One-MedReID-Pytorch}{https://github.com/tianyuan168326/All-in-One-MedReID-Pytorch}.

cs.CV

Postselected communication over quantum channels

The single-letter characterisation of the entanglement-assisted capacity of a quantum channel is one of the seminal results of quantum information theory. In this paper, we consider a modified communication scenario in which the receiver is allowed an additional, `inconclusive' measurement outcome, and we employ an error metric given by the error probability in decoding the transmitted message conditioned on a conclusive measurement result. We call this setting postselected communication and the ensuing highest achievable rates the postselected capacities. Here, we provide a precise single-letter characterisation of postselected capacities in the setting of entanglement assistance as well as the more general nonsignalling assistance, establishing that they are both equal to the channel's projective mutual information -- a variant of mutual information based on the Hilbert projective metric. We do so by establishing bounds on the one-shot postselected capacities, with a lower bound that makes use of a postselected teleportation-based protocol and an upper bound in terms of the postselected hypothesis testing relative entropy. As such, we obtain fundamental limits on a channel's ability to communicate even when this strong resource of postselection is allowed, implying limitations on communication even when the receiver has access to postselected closed timelike curves.

quant-ph

Incompatibility as a resource for programmable quantum instruments

Quantum instruments represent the most general type of quantum measurement, as they incorporate processes with both classical and quantum outputs. In many scenarios, it may be desirable to have some "on-demand" device that is capable of implementing one of many possible instruments whenever the experimenter desires. We refer to such objects as programmable instrument devices (PIDs), and this paper studies PIDs from a resource-theoretic perspective. A physically important class of PIDs are those that do not require quantum memories to implement, and these are naturally "free" in this resource theory. Additionally, these free objects correspond precisely to the class of unsteerable channel assemblages in the study of channel steering. The traditional notion of measurement incompatibility emerges as a resource in this theory since any PID controlling an incompatible family of instruments requires a quantum memory to build. We identify an incompatibility preorder between PIDs based on whether one can be transformed into another using processes that do not require additional quantum memories. Necessary and sufficient conditions are derived for when such transformations are possible based on how well certain guessing games can be played using a given PID. Ultimately our results provide an operational characterization of incompatibility, and they offer semi-device-independent tests for incompatibility in the most general types of quantum instruments.

quant-ph