SearcharxivSearch

arXiv subjects

Hongyu He

Publications and source records attributed to Hongyu He.

At least 19 recordsLinked to original sources

Quotient DAGs for Off-Policy Evaluation:Forward-Flow Importance Sampling and Exact Slate Propensities

Off-policy evaluation estimates how a target policy would perform using data collected by a different behavior policy, which is crucial when online testing is costly or risky, such as in recommendation or healthcare. Standard importance sampling reweights each logged trajectory, but it can treat details of the generation process as meaningful even when the evaluation target ignores them: for example, an autoregressive slate recommender may generate an ordered sequence of items while the reward and downstream estimator depend only on the unordered slate. This creates nuisance variance and a computational gap, since exact unordered slate propensities require summing over all generation orders. We introduce a quotient-DAG view that merges histories equivalent for evaluation and assigns weights using target-to-behavior forward-flow ratios on the merged graph. For slate recommendation under a set-sufficient next-item interface, this yields Forward-DP, a subset-DAG dynamic program that computes exact unordered propensities without factorial enumeration. The resulting propensity primitive enables practical propensity-based evaluation and model selection for context-dependent autoregressive slate loggers.

cs.LG

Navigating heterogeneous protein landscapes through geometry-aware smoothing

The evolutionary fitness landscape of biological molecules is extremely sparse and heterogeneous, with functional sequences forming isolated dense ``islands'' within a vast combinatorial space of largely non-functional variants. Protein sequences, in particular, exemplify this structure, yet most generative artificial intelligence models implicitly assume a homogeneous data distribution. We show that this assumption fundamentally breaks down in heterogeneous biological sequence spaces: fixed global noise levels impose a destructive trade-off, either oversmoothing dense functional clusters or fragmenting sparse regions and producing non-functional hallucinations. To address this limitation, we introduce \emph{Density-Dependent Smoothing} (DDS), a geometry-aware generative framework that adapts stochastic smoothing to the local density of the underlying sequence landscape. By inversely coupling diffusion noise to estimated sequence density, DDS enables gentle refinement in high-density functional regions while promoting controlled exploration across sparse regions. Implemented as a plug-in mechanism for discrete molecular sampling, DDS consistently outperforms state-of-the-art diffusion and autoregressive models across antibody repertoires, therapeutic antibody design, antimicrobial peptide generation and coronavirus antibody design. Together, these results show that fixed global smoothing assumptions fundamentally limit generative modeling in sparse biological sequence spaces, and that geometry-aware smoothing removes this constraint, enabling reliable exploration and design previously unattainable with fixed-noise generative models.

cs.CE

AI-generated data contamination erodes pathological variability and diagnostic reliability

Generative artificial intelligence (AI) is rapidly populating medical records with synthetic content, creating a feedback loop where future models are increasingly at risk of training on uncurated AI-generated data. However, the clinical consequences of this AI-generated data contamination remain unexplored. Here, we show that in the absence of mandatory human verification, this self-referential cycle drives a rapid erosion of pathological variability and diagnostic reliability. By analysing more than 800,000 synthetic data points across clinical text generation, vision-language reporting, and medical image synthesis, we find that models progressively converge toward generic phenotypes regardless of the model architecture. Specifically, rare but critical findings, including pneumothorax and effusions, vanish from the synthetic content generated by AI models, while demographic representations skew heavily toward middle-aged male phenotypes. Crucially, this degradation is masked by false diagnostic confidence; models continue to issue reassuring reports while failing to detect life-threatening pathology, with false reassurance rates tripling to 40%. Blinded physician evaluation confirms that this decoupling of confidence and accuracy renders AI-generated documentation clinically useless after just two generations. We systematically evaluate three mitigation strategies, finding that while synthetic volume scaling fails to prevent collapse, mixing real data with quality-aware filtering effectively preserves diversity. Ultimately, our results suggest that without policy-mandated human oversight, the deployment of generative AI threatens to degrade the very healthcare data ecosystems it relies upon.

cs.CY

Interference-governed electromagnetic-thermal coupling and heat transport in pulse EUV-irradiated multilayer nanofilms

Mo-Si multilayer mirrors are central to extreme ultraviolet lithography, where nanoscale optical interference and heat accumulation together constrain reflectivity and operational stability. Here we develop an analytical electromagnetic-thermal coupling model that directly links transfer-matrix-based interference-controlled energy deposition with transient heat conduction in EUV-irradiated multilayers. The model reveals a fundamental trade-off whereby increasing the multilayer period number enhances reflectivity but simultaneously elevates temperature by impeding heat dissipation. Interference-driven volumetric absorption further gives rise to pronounced axial temperature gradients and a post-pulse downward migration of the heat-flux maximum, a delayed-heating effect inaccessible to conventional surface-flux-based models. Systematic analysis establishes scaling laws connecting interfacial thermal resistance, beam size, and incident energy density to thermal confinement and temperature rise. By incorporating interfacial compaction kinetics, the model enables a quantitative assessment of mirror lifetime. This work offers a theoretical tool for thermal-optical co-design of multilayer nanostructures including EUV mirrors under pulsed irradiation across a wide spectral range.

physics.app-ph

HypoSpace: A Diagnostic Benchmark for Set-Valued Hypothesis Generation under Underdetermination and Sublinear Coverage Bounds

Many scientific problems are underdetermined: multiple distinct hypotheses are equally consistent with the same observations. In such settings, effective inference requires not only producing valid explanations, but also systematically exploring and covering the admissible hypothesis set. We introduce HypoSpace, a benchmark that treats large language models (LLMs) as samplers over finite hypothesis spaces and evaluates them on three metrics: Validity, Uniqueness, and Recovery. HypoSpace spans three structured domains (causal graph inference, gravity-constrained 3D voxel reconstruction, and Boolean genetic interaction modeling) with deterministic validators and exactly enumerable solution spaces, plus real-world anchored case studies. Empirically, HypoSpace reveals a capability- and scale-dependent coverage failure: models can maintain high Validity while exhibiting reduced Uniqueness and Recovery as admissible hypothesis spaces become larger or more combinatorial. We further show that the analysis on stratified decoding partially mitigates this collapse, demonstrating HypoSpace's utility as a diagnostic benchmark for set-valued inference. Code is available at: https://github.com/CTT-Pavilion/_HypoSpace.

cs.CL

Safety challenges of AI in medicine in the era of large language models

Recent advancements in artificial intelligence (AI), particularly in large language models (LLMs), have unlocked significant potential to enhance the quality and efficiency of medical care. By introducing a novel way to interact with AI and data through natural language, LLMs offer new opportunities for medical practitioners, patients, and researchers. However, as AI and LLMs become more powerful and especially achieve superhuman performance in some medical tasks, public concerns over their safety have intensified. These concerns about AI safety have emerged as the most significant obstacles to the adoption of AI in medicine. In response, this review examines emerging risks in AI utilization during the LLM era. First, we explore LLM-specific safety challenges from functional and communication perspectives, addressing issues across data collection, model training, and real-world application. We then consider inherent safety problems shared by all AI systems, along with additional complications introduced by LLMs. Last, we discussed how safety issues of using AI in clinical practice and healthcare system operation would undermine trust among patient, clinicians and the public, and how to build confidence in these systems. By emphasizing the development of safe AI, we believe these technologies can be more rapidly and reliably integrated into everyday medical practice to benefit both patients and clinicians.

cs.CY

LLMvsSmall Model? Large Language Model Based Text Augmentation Enhanced Personality Detection Model

Personality detection aims to detect one's personality traits underlying in social media posts. One challenge of this task is the scarcity of ground-truth personality traits which are collected from self-report questionnaires. Most existing methods learn post features directly by fine-tuning the pre-trained language models under the supervision of limited personality labels. This leads to inferior quality of post features and consequently affects the performance. In addition, they treat personality traits as one-hot classification labels, overlooking the semantic information within them. In this paper, we propose a large language model (LLM) based text augmentation enhanced personality detection model, which distills the LLM's knowledge to enhance the small model for personality detection, even when the LLM fails in this task. Specifically, we enable LLM to generate post analyses (augmentations) from the aspects of semantic, sentiment, and linguistic, which are critical for personality detection. By using contrastive learning to pull them together in the embedding space, the post encoder can better capture the psycho-linguistic information within the post representations, thus improving personality detection. Furthermore, we utilize the LLM to enrich the information of personality labels for enhancing the detection performance. Experimental results on the benchmark datasets demonstrate that our model outperforms the state-of-the-art methods on personality detection.

cs.CL

Projection of Elliptic Orbits and Branching Laws

Let $G$ be a Lie group, and $H\subset G$ a closed subgroup. Let $\pi$ be an irreducible unitary representation of $G$. In this paper, we briefly discuss the orbit method and its application to the branching problem $\pi|_{H}$. We use the Gan-Gross-Prasad branching law for $(G, H)= ( U(p,q), U(p, q-1) )$ as an example to illustrate the relation between $\pro_{\f u(p, q-1)}^{\f u(p,q)} \mc O(\lambda)$ and the branching law of the discrete series $D_{\lambda}|_{U(p,q-1)}$ for $\lambda$ an regular elliptic element. We also discuss some results regarding branching laws and wave front sets. The presentation of this paper does not follow the historical timeline of development.

math.RT

SuperMask: Generating High-resolution object masks from multi-view, unaligned low-resolution MRIs

Three-dimensional segmentation in magnetic resonance images (MRI), which reflects the true shape of the objects, is challenging since high-resolution isotropic MRIs are rare and typical MRIs are anisotropic, with the out-of-plane dimension having a much lower resolution. A potential remedy to this issue lies in the fact that often multiple sequences are acquired on different planes. However, in practice, these sequences are not orthogonal to each other, limiting the applicability of many previous solutions to reconstruct higher-resolution images from multiple lower-resolution ones. We propose a weakly-supervised deep learning-based solution to generating high-resolution masks from multiple low-resolution images. Our method combines segmentation and unsupervised registration networks by introducing two new regularizations to make registration and segmentation reinforce each other. Finally, we introduce a multi-view fusion method to generate high-resolution target object masks. The experimental results on two datasets show the superiority of our methods. Importantly, the advantage of not using high-resolution images in the training process makes our method applicable to a wide variety of MRI segmentation tasks.

eess.IV

A Thin Fundamental Set for SL(2, Z)

Let $\Gamma=SL(2, \mathbb Z)$ and $G=SL(2, \mathbb R)$. Let $g=kan$ be the Iwasawa decomposition. Let $\epsilon$ be a small positive number. In this paper, we construct a fundamental set $\mathcal F_{\epsilon}$ such that the $k$-component of $ g \in \mathcal F_{\epsilon}$ is within the $\epsilon$-distance from the identity. We further prove an inequality for the $L^2$-norm of functions on $G/\Gamma$.

math.NT

FaceGuard: Proactive Deepfake Detection

Existing deepfake-detection methods focus on passive detection, i.e., they detect fake face images via exploiting the artifacts produced during deepfake manipulation. A key limitation of passive detection is that it cannot detect fake faces that are generated by new deepfake generation methods. In this work, we propose FaceGuard, a proactive deepfake-detection framework. FaceGuard embeds a watermark into a real face image before it is published on social media. Given a face image that claims to be an individual (e.g., Nicolas Cage), FaceGuard extracts a watermark from it and predicts the face image to be fake if the extracted watermark does not match well with the individual's ground truth one. A key component of FaceGuard is a new deep-learning-based watermarking method, which is 1) robust to normal image post-processing such as JPEG compression, Gaussian blurring, cropping, and resizing, but 2) fragile to deepfake manipulation. Our evaluation on multiple datasets shows that FaceGuard can detect deepfakes accurately and outperforms existing methods.

cs.CV

How Can Datacenters Join the Smart Grid to Address the Climate Crisis? Using simulation to explore power and cost effects of direct participation in the energy market

Amidst the climate crisis, the massive introduction of renewable energy sources has brought tremendous challenges to both the power grid and its surrounding markets. As datacenters have become ever-larger and more powerful, they play an increasingly significant role in the energy arena. With their unique characteristics, datacenters have been proved to be well-suited for regulating the power grid yet currently provide little, if any, such active response. This problem is due to issues such as unsuitability of the market design, high complexity of the currently proposed solutions, as well as the potential risks thereof. This work aims to provide individual datacenters with insights on the feasibility and profitability of directly participating in the energy market. By modelling the power system of datacenters, and by conducting simulations on real-world datacenter traces, we demonstrate the substantial financial incentive for individual datacenters to directly participate in both the day-ahead and the balancing markets. In turn, we suggest a new short-term, direct scheme of market participation for individual datacenters in place of the current long-term, inactive participation. Furthermore, we develop a novel proactive DVFS scheduling algorithm that can both reduce energy consumption and save energy costs during the market participation of datacenters. Also, in developing this scheduler, we propose an innovative combination of machine learning methods and the DVFS technology that can provide the power grid with indirect demand response (DR). Our experimental results strongly support that individual datacenters can and should directly participate in the energy market both to save their energy costs and to curb their energy consumption, whilst providing the power grid with indirect DR.

cs.DC

Tab2Know: Building a Knowledge Base from Tables in Scientific Papers

Tables in scientific papers contain a wealth of valuable knowledge for the scientific enterprise. To help the many of us who frequently consult this type of knowledge, we present Tab2Know, a new end-to-end system to build a Knowledge Base (KB) from tables in scientific papers. Tab2Know addresses the challenge of automatically interpreting the tables in papers and of disambiguating the entities that they contain. To solve these problems, we propose a pipeline that employs both statistical-based classifiers and logic-based reasoning. First, our pipeline applies weakly supervised classifiers to recognize the type of tables and columns, with the help of a data labeling system and an ontology specifically designed for our purpose. Then, logic-based reasoning is used to link equivalent entities (via sameAs links) in different tables. An empirical evaluation of our approach using a corpus of papers in the Computer Science domain has returned satisfactory performance. This suggests that ours is a promising step to create a large-scale KB of scientific knowledge.

cs.AI

On the growth of Rankin-Selberg L-functions for $SL(2)$

In this paper, we establish bounds of the Rankin-Selberg $L$-function for $SL(2)$ using the supnorm of the Eisenstein series and a purely representation theoretic index over the real group. Consequently, we obtain a subconvexity bound $L(\frac{1}{2}+ it, f_1 \times f_2) \leq C (1+ |t|)^{\frac{5}{6}+\epsilon}$ for two Maass cusp forms of $SL(2, \mathbb Z)$.

math.RT

Certain L2-norms on automorphic representations of SL(2)

Let $\Gamma$ be a non-uniform lattice in $SL(2, \mathbb R)$. In this paper, we study various $L^2$-norms of automorphic representations of $SL(2, \mathbb R)$. We bound these norms with intrinsic norms defined on the representation. Comparison of these norms will help us understand the growth of $L$-functions in a systematic way.

math.RT

Certain L^2-norm and Asymptotic bounds of Whittaker Function for GL(n)

Whittaker functions of $GL(n, \mathbb R)$ , are most known for its role in the Fourier-Whittaker expansion of cusp forms. Their behavior in the Siegel set, in large, is well-understood. In this paper, we insert into the literature some potentially useful properties of Whittaker function over the group $GL(n, \mathbb R)$ and the mirobolic group $P_n$. We proved the square integrabilty of the Whittaker functions with respect to certain measures, extending a theorem of Jacquet and Shalika . For principal series representations, we gave various asymptotic bounds of smooth Whittaker functions over the whole group $GL(n, \mathbb R)$. Due to the lack of good terminology, we use whittaker functions to refer to $K$-finite or smooth vectors in the Whittaker model.

math.RT

Representation of ax+b group and Dirichlet Series

Let $G$ be the $ax+b$ group. There are essentially two irreducible infinite dimensional unitary representations of $G$, $(\mu, L^2(\mathbb R^+))$ and $(\mu^*, L^2(\mathbb R^+))$. In this paper, we give various characterizations about smooth vectors of $\mu$ and their Mellin transforms. Let $\f d$ be a linear sum of delta distributions supported on the the positive integers $\mathbb Z^+$. We study the Mellin transform of the matrix coefficients $\mu_{ \f d, f}(a)$ with $f$ smooth. We express these Mellin transforms in terms of the Dirichlet series $L(s, \f d)$. We determine a sufficient condition such that the generalized matrix coefficient $\mu_{\f d, f}$ is a locally integrable function and estimate the $L^2$-norms of $\mu_{\f d, f}$ over the Siegel set. We further derive an inequality which may potentially be used to study the Dirichlet series $L(s, \f d)$.

math.RT

OpenHI2 -- Open source histopathological image platform

Transition from conventional to digital pathology requires a new category of biomedical informatic infrastructure which could facilitate delicate pathological routine. Pathological diagnoses are sensitive to many external factors and is known to be subjective. Only systems that can meet strict requirements in pathology would be able to run along pathological routines and eventually digitized the study area, and the developed platform should comply with existing pathological routines and international standards. Currently, there are a number of available software tools which can perform histopathological tasks including virtual slide viewing, annotating, and basic image analysis, however, none of them can serve as a digital platform for pathology. Here we describe OpenHI2, an enhanced version Open Histopathological Image platform which is capable of supporting all basic pathological tasks and file formats; ready to be deployed in medical institutions on a standard server environment or cloud computing infrastructure. In this paper, we also describe the development decisions for the platform and propose solutions to overcome technical challenges so that OpenHI2 could be used as a platform for histopathological images. Further addition can be made to the platform since each component is modularized and fully documented. OpenHI2 is free, open-source, and available at https://gitlab.com/BioAI/OpenHI.

q-bio.QM