SearcharxivSearch

arXiv subjects

Yang Han

Publications and source records attributed to Yang Han.

At least 37 records · Page 2Linked to original sources

REAL: Benchmarking Abilities of Large Language Models for Housing Transactions and Services

The development of large language models (LLMs) has greatly promoted the progress of chatbot in multiple fields. There is an urgent need to evaluate whether LLMs can play the role of agent in housing transactions and services as well as humans. We present Real Estate Agent Large Language Model Evaluation (REAL), the first evaluation suite designed to assess the abilities of LLMs in the field of housing transactions and services. REAL comprises 5,316 high-quality evaluation entries across 4 topics: memory, comprehension, reasoning and hallucination. All these entries are organized as 14 categories to assess whether LLMs have the knowledge and ability in housing transactions and services scenario. Additionally, the REAL is used to evaluate the performance of most advanced LLMs. The experiment results indicate that LLMs still have significant room for improvement to be applied in the real estate field.

cs.AI

Counterfactual Uncertainty Quantification of Factual Estimand of Efficacy from Before-and-After Treatment Repeated Measures Randomized Controlled Trials

This article quantifies the uncertainty reduction achievable for \textit{counterfactual} estimand, and cautions against potential bias when the estimand uses Digital Twins. Posed by Neyman (1923a) who showed unbiased \textit{point estimation} from designed \textit{factual} experiments is possible, \textit{counterfactual} uncertainty quantification (CUQ) remained an open challenge for about one hundred years. The $Rx: C$ \textit{counterfactual} efficacy we focus on is the ideal estimand for comparing treatment $Rx$ with control $C$, the expected outcome differential if each patient received \textit{both} $Rx$ and $C$. Enabled by our new statistical modeling principle called ETZ, we show CUQ is achievable in Randomized Controlled Trials (RCTs) with \textit{Before-and-After} Repeated Measures, common in many therapeutic areas. The CUQ we are able to achieve typically has lower variability than factual UQ. We caution against using predictors with measurement error, which violates regression assumptions and can cause \textit{attenuation} bias in estimating treatment effects. For traditional medicine and population-averaged targeted therapy, counterfactual point estimation remains unbiased. However, in both Real Human and Digital Twin approaches, estimating effects in \emph{subgroups} may suffer attenuation bias.

stat.ML

Reward Generalization in RLHF: A Topological Perspective

Existing alignment methods share a common topology of information flow, where reward information is collected from humans, modeled with preference learning, and used to tune language models. However, this shared topology has not been systematically characterized, nor have its alternatives been thoroughly explored, leaving the problems of low data efficiency and unreliable generalization unaddressed. As a solution, we introduce a theory of reward generalization in reinforcement learning from human feedback (RLHF), focusing on the topology of information flow at both macro and micro levels. At the macro level, we portray the RLHF information flow as an autoencoding process over behavior distributions, formalizing the RLHF objective of distributional consistency between human preference and model behavior. At the micro level, we present induced Bayesian networks to model the impact of dataset topologies on reward generalization. Combining analysis on both levels, we propose reward modeling from tree-structured preference information. It is shown to reduce reward uncertainty by up to $Θ(\log n/\log\log n)$ times compared to baselines, where $n$ is the dataset size. Validation on three NLP tasks shows that it achieves an average win rate of 65% against baselines, thus improving reward generalization for free via topology design, while reducing the amount of data requiring annotation.

cs.LG

Reverse-Speech-Finder: A Neural Network Backtracking Architecture for Generating Alzheimer's Disease Speech Samples and Improving Diagnosis Performance

This study introduces Reverse-Speech-Finder (RSF), a groundbreaking neural network backtracking architecture designed to enhance Alzheimer's Disease (AD) diagnosis through speech analysis. Leveraging the power of pre-trained large language models, RSF identifies and utilizes the most probable AD-specific speech markers, addressing both the scarcity of real AD speech samples and the challenge of limited interpretability in existing models. RSF's unique approach consists of three core innovations: Firstly, it exploits the observation that speech markers most probable of predicting AD, defined as the most probable speech-markers (MPMs), must have the highest probability of activating those neurons (in the neural network) with the highest probability of predicting AD, defined as the most probable neurons (MPNs). Secondly, it utilizes a speech token representation at the input layer, allowing backtracking from MPNs to identify the most probable speech-tokens (MPTs) of AD. Lastly, it develops an innovative backtracking method to track backwards from the MPNs to the input layer, identifying the MPTs and the corresponding MPMs, and ingeniously uncovering novel speech markers for AD detection. Experimental results demonstrate RSF's superiority over traditional methods such as SHAP and Integrated Gradients, achieving a 3.5% improvement in accuracy and a 3.2% boost in F1-score. By generating speech data that encapsulates novel markers, RSF not only mitigates the limitations of real data scarcity but also significantly enhances the robustness and accuracy of AD diagnostic models. These findings underscore RSF's potential as a transformative tool in speech-based AD detection, offering new insights into AD-related linguistic deficits and paving the way for more effective non-invasive early intervention strategies.

cs.LG

The High Voltage Splitter board for the JUNO SPMT system

The Jiangmen Underground Neutrino Observatory (JUNO) in southern China is designed to study neutrinos from nuclear reactors and natural sources to address fundamental questions in neutrino physics. Achieving its goals requires continuous operation over a 20-year period. The small photomultiplier tube (small PMT or SPMT) system is a subsystem within the experiment composed of 25600 3-inch PMTs and their associated readout electronics. The High Voltage Splitter (HVS) is the first board on the readout chain of the SPMT system and services the PMTs by providing high voltage for biasing and by decoupling the generated physics signal from the high-voltage bias for readout, which is then fed to the front-end board. The necessity to handle high voltage, manage a large channel count, and operate stably for 20 years imposes significant constraints on the physical design of the HVS. This paper serves as a comprehensive documentation of the HVS board: its role in the SPMT readout system, the challenges in its design, performance and reliability metrics, and the methods employed for production and quality control.

physics.ins-det

Large Language Models to Accelerate Organic Chemistry Synthesis

Chemical synthesis, as a foundational methodology in the creation of transformative molecules, exerts substantial influence across diverse sectors from life sciences to materials and energy. Current chemical synthesis practices emphasize laborious and costly trial-and-error workflows, underscoring the urgent need for advanced AI assistants. Nowadays, large language models (LLMs), typified by GPT-4, have been introduced as an efficient tool to facilitate scientific research. Here, we present Chemma, a fully fine-tuned LLM with 1.28 million pairs of Q&A about reactions, as an assistant to accelerate organic chemistry synthesis. Chemma surpasses the best-known results in multiple chemical tasks, e.g., single-step retrosynthesis and yield prediction, which highlights the potential of general AI for organic chemistry. Via predicting yields across the experimental reaction space, Chemma significantly improves the reaction exploration capability of Bayesian optimization. More importantly, integrated in an active learning framework, Chemma exhibits advanced potential for autonomous experimental exploration and optimization in open reaction spaces. For an unreported Suzuki-Miyaura cross-coupling reaction of cyclic aminoboronates and aryl halides for the synthesis of $α$-Aryl N-heterocycles, the human-AI collaboration successfully explored suitable ligand and solvent (1,4-dioxane) within only 15 runs, achieving an isolated yield of 67%. These results reveal that, without quantum-chemical calculations, Chemma can comprehend and extract chemical insights from reaction data, in a manner akin to human experts. This work opens avenues for accelerating organic chemistry synthesis with adapted large language models.

physics.chem-ph

Magnetic Preference Optimization: Achieving Last-iterate Convergence for Language Model Alignment

Self-play methods have demonstrated remarkable success in enhancing model capabilities across various domains. In the context of Reinforcement Learning from Human Feedback (RLHF), self-play not only boosts Large Language Model (LLM) performance but also overcomes the limitations of traditional Bradley-Terry (BT) model assumptions by finding the Nash equilibrium (NE) of a preference-based, two-player constant-sum game. However, existing methods either guarantee only average-iterate convergence, incurring high storage and inference costs, or converge to the NE of a regularized game, failing to accurately reflect true human preferences. In this paper, we introduce Magnetic Preference Optimization (MPO), a novel approach capable of achieving last-iterate convergence to the NE of the original game, effectively overcoming the limitations of existing methods. Building upon Magnetic Mirror Descent (MMD), MPO attains a linear convergence rate, making it particularly suitable for fine-tuning LLMs. To ensure our algorithm is both theoretically sound and practically viable, we present a simple yet effective implementation that adapts the theoretical insights to the RLHF setting. Empirical results demonstrate that MPO can significantly enhance the performance of LLMs, highlighting the potential of self-play methods in alignment.

cs.CL

Unravelling Causal Genetic Biomarkers of Alzheimer's Disease via Neuron to Gene-token Backtracking in Neural Architecture: A Groundbreaking Reverse-Gene-Finder Approach

Alzheimer's Disease (AD) affects over 55 million people globally, yet the key genetic contributors remain poorly understood. Leveraging recent advancements in genomic foundation models, we present the innovative Reverse-Gene-Finder technology, a ground-breaking neuron-to-gene-token backtracking approach in a neural network architecture to elucidate the novel causal genetic biomarkers driving AD onset. Reverse-Gene-Finder comprises three key innovations. Firstly, we exploit the observation that genes with the highest probability of causing AD, defined as the most causal genes (MCGs), must have the highest probability of activating those neurons with the highest probability of causing AD, defined as the most causal neurons (MCNs). Secondly, we utilize a gene token representation at the input layer to allow each gene (known or novel to AD) to be represented as a discrete and unique entity in the input space. Lastly, in contrast to the existing neural network architectures, which track neuron activations from the input layer to the output layer in a feed-forward manner, we develop an innovative backtracking method to track backwards from the MCNs to the input layer, identifying the Most Causal Tokens (MCTs) and the corresponding MCGs. Reverse-Gene-Finder is highly interpretable, generalizable, and adaptable, providing a promising avenue for application in other disease scenarios.

cs.LG

From Generalist to Specialist: A Survey of Large Language Models for Chemistry

Large Language Models (LLMs) have significantly transformed our daily life and established a new paradigm in natural language processing (NLP). However, the predominant pretraining of LLMs on extensive web-based texts remains insufficient for advanced scientific discovery, particularly in chemistry. The scarcity of specialized chemistry data, coupled with the complexity of multi-modal data such as 2D graph, 3D structure and spectrum, present distinct challenges. Although several studies have reviewed Pretrained Language Models (PLMs) in chemistry, there is a conspicuous absence of a systematic survey specifically focused on chemistry-oriented LLMs. In this paper, we outline methodologies for incorporating domain-specific chemistry knowledge and multi-modal information into LLMs, we also conceptualize chemistry LLMs as agents using chemistry tools and investigate their potential to accelerate scientific research. Additionally, we conclude the existing benchmarks to evaluate chemistry ability of LLMs. Finally, we critically examine the current challenges and identify promising directions for future research. Through this comprehensive survey, we aim to assist researchers in staying at the forefront of developments in chemistry LLMs and to inspire innovative applications in the field.

physics.chem-ph

Multi-Calorimetry in Light-based Neutrino Detectors

Neutrino detectors are among the largest photon detection instruments, built to capture scarce photons upon energy deposition. Many discoveries in neutrino physics, including the neutrino itself, are inseparable from the advances in photon detection technology, particularly in photo-sensors and readout electronics, to yield ever higher precision and richer detection information. The measurement of the energy of neutrinos, referred to as calorimetry, can be achieved in two distinct approaches: photon-counting, where single-photon can be counted digitally, and photon-integration, where multi-photons are aggregated and estimated via analogue signals. The energy is pursued today to reach permille level systematics control precision in ever-vast volumes, exemplified by experiments like JUNO. The unprecedented precision brings to the foreground the systematics due to calorimetric response entanglements in energy, position and time that were negligible in the past, thus driving further innovation in calorimetry. This publication describes a novel articulation that detectors can be endowed with multiple photon detection systems. This multi-calorimetry approach opens the notion of dual-calorimetry detector, consisting of both photon-counting and photon-integration systems, as a cost-effective evolution from the single calorimetry setups used over several decades for most experiments so far. The dual-calorimetry design exploits unique response synergies between photon-counting and photon-integration systems, including correlations and cancellations in calorimetric responses, to maximise the mitigation of response entanglements, thereby yielding permille-level high-precision calorimetry.

hep-ex

SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific Research

Recently, there has been growing interest in using Large Language Models (LLMs) for scientific research. Numerous benchmarks have been proposed to evaluate the ability of LLMs for scientific research. However, current benchmarks are mostly based on pre-collected objective questions. This design suffers from data leakage problem and lacks the evaluation of subjective Q/A ability. In this paper, we propose SciEval, a comprehensive and multi-disciplinary evaluation benchmark to address these issues. Based on Bloom's taxonomy, SciEval covers four dimensions to systematically evaluate scientific research ability. In particular, we design a "dynamic" subset based on scientific principles to prevent evaluation from potential data leakage. Both objective and subjective questions are included in SciEval. These characteristics make SciEval a more effective benchmark for scientific research ability evaluation of LLMs. Comprehensive experiments on most advanced LLMs show that, although GPT-4 achieves SOTA performance compared to other LLMs, there is still substantial room for improvement, especially for dynamic questions. The codes and data are publicly available on https://github.com/OpenDFM/SciEval.

cs.CL

J2N -- Nominal Adjective Identification and its Application

This paper explores the challenges posed by nominal adjectives (NAs) in natural language processing (NLP) tasks, particularly in part-of-speech (POS) tagging. We propose treating NAs as a distinct POS tag, "JN," and investigate its impact on POS tagging, BIO chunking, and coreference resolution. Our study shows that reclassifying NAs can improve the accuracy of syntactic analysis and structural understanding in NLP. We present experimental results using Hidden Markov Models (HMMs), Maximum Entropy (MaxEnt) models, and Spacy, demonstrating the feasibility and potential benefits of this approach. Additionally we finetuned a bert model to identify the NA in untagged text.

cs.CL

AlignSum: Data Pyramid Hierarchical Fine-tuning for Aligning with Human Summarization Preference

Text summarization tasks commonly employ Pre-trained Language Models (PLMs) to fit diverse standard datasets. While these PLMs excel in automatic evaluations, they frequently underperform in human evaluations, indicating a deviation between their generated summaries and human summarization preferences. This discrepancy is likely due to the low quality of fine-tuning datasets and the limited availability of high-quality human-annotated data that reflect true human preference. To address this challenge, we introduce a novel human summarization preference alignment framework AlignSum. This framework consists of three parts: Firstly, we construct a Data Pymarid with extractive, abstractive, and human-annotated summary data. Secondly, we conduct the Gaussian Resampling to remove summaries with extreme lengths. Finally, we implement the two-stage hierarchical fine-tuning with Data Pymarid after Gaussian Resampling. We apply AlignSum to PLMs on the human-annotated CNN/DailyMail and BBC XSum datasets. Experiments show that with AlignSum, PLMs like BART-Large surpass 175B GPT-3 in both automatic and human evaluations. This demonstrates that AlignSum significantly enhances the alignment of language models with human summarization preferences.

cs.CL

Research on fusing topological data analysis with convolutional neural network

Convolutional Neural Network (CNN) struggle to capture the multi-dimensional structural information of complex high-dimensional data, which limits their feature learning capability. This paper proposes a feature fusion method based on Topological Data Analysis (TDA) and CNN, named TDA-CNN. This method combines numerical distribution features captured by CNN with topological structure features captured by TDA to improve the feature learning and representation ability of CNN. TDA-CNN divides feature extraction into a CNN channel and a TDA channel. CNN channel extracts numerical distribution features, and the TDA channel extracts topological structure features. The two types of features are fused to form a combined feature representation, with the importance weights of each feature adaptively learned through an attention mechanism. Experimental validation on datasets such as Intel Image, Gender Images, and Chinese Calligraphy Styles by Calligraphers demonstrates that TDA-CNN improves the performance of VGG16, DenseNet121, and GoogleNet networks by 17.5%, 7.11%, and 4.45%, respectively. TDA-CNN demonstrates improved feature clustering and the ability to recognize important features. This effectively enhances the model's decision-making ability.

cs.CV

Entropies of Serre functors for higher hereditary algebras

For a higher hereditary algebra, we calculate its upper (lower) Serre dimension, the entropy and polynomial entropy of Serre functor, and the Hochschild (co)homology entropy of Serre quasi-functor. These invariants are given by its Calabi-Yau dimension for a higher representation-finite algebra, and by its global dimension and the spectral radius and polynomial growth rate of its Coxeter matrix for a higher representation-infinite algebra. For this, we prove the Yomdin type inequality on Hochschild homology entropy for a finite dimensional elementary algebra of finite global dimension. Our calculations imply that the Kikuta and Ouchi's question on relations between entropy and Hochschild (co)homology entropy has positive answer, and the Gromov-Yomdin type equalities on entropy and Hochschild (co)homology entropy hold, for the Serre functor on perfect derived category and the Serre quasi-functor on perfect dg module category of an indecomposable elementary higher hereditary algebra.

math.RT

Minimum Area Confidence Set Optimality for Simultaneous Confidence Bands for Percentiles in Linear Regression

Simultaneous confidence bands (SCBs) for percentiles in linear regression are valuable tools with many applications. In this paper, we propose a novel criterion for comparing SCBs for percentiles, termed the Minimum Area Confidence Set (MACS) criterion. This criterion utilizes the area of the confidence set for the pivotal quantities, which are generated from the confidence set of the unknown parameters. Subsequently, we employ the MACS criterion to construct exact SCBs over any finite covariate intervals and to compare multiple SCBs of different forms. This approach can be used to determine the optimal SCBs. It is discovered that the area of the confidence set for the pivotal quantities of an asymmetric SCB is uniformly and can be very substantially smaller than that of the corresponding symmetric SCB. Therefore, under the MACS criterion, exact asymmetric SCBs should always be preferred. Furthermore, a new computationally efficient method is proposed to calculate the critical constants of exact SCBs for percentiles. A real data example on drug stability study is provided for illustration.

stat.ME

Melting of excitonic insulator phase by an intense terahertz pulse in Ta$_2$NiSe$_5$

In this study, the optical response to a terahertz pulse was investigated in the transition metal chalcogenide Ta$_2$NiSe$_5$, a candidate excitonic insulator. First, by irradiating a terahertz pulse with a relatively weak electric field (0.3 MV/cm), the spectral changes in reflectivity near the absorption edge due to third-order optical nonlinearity were measured and the absorption peak characteristic of the excitonic phase just below the interband transition was identified. Next, by irradiating a strong terahertz pulse with a strong electric field of 1.65 MV/cm, the absorption of the excitonic phase was found to be reduced, and a Drude-like response appeared in the mid-infrared region. These responses can be interpreted as carrier generation by exciton dissociation induced by the electric field, resulting in the partial melting of the excitonic phase and metallization. The presence of a distinct threshold electric field for carrier generation indicates exciton dissociation via quantum-tunnelling processes. The spectral change due to metallization by the electric field is significantly different from that due to the strong optical excitation across the gap, which can be explained by the different melting mechanisms of the excitonic phase in the two types of excitations.

cond-mat.str-el

Identifying outliers in astronomical images with unsupervised machine learning

Astronomical outliers, such as unusual, rare or unknown types of astronomical objects or phenomena, constantly lead to the discovery of genuinely unforeseen knowledge in astronomy. More unpredictable outliers will be uncovered in principle with the increment of the coverage and quality of upcoming survey data. However, it is a severe challenge to mine rare and unexpected targets from enormous data with human inspection due to a significant workload. Supervised learning is also unsuitable for this purpose since designing proper training sets for unanticipated signals is unworkable. Motivated by these challenges, we adopt unsupervised machine learning approaches to identify outliers in the data of galaxy images to explore the paths for detecting astronomical outliers. For comparison, we construct three methods, which are built upon the k-nearest neighbors (KNN), Convolutional Auto-Encoder (CAE)+ KNN, and CAE + KNN + Attention Mechanism (attCAE KNN) separately. Testing sets are created based on the Galaxy Zoo image data published online to evaluate the performance of the above methods. Results show that attCAE KNN achieves the best recall (78%), which is 53% higher than the classical KNN method and 22% higher than CAE+KNN. The efficiency of attCAE KNN (10 minutes) is also superior to KNN (4 hours) and equal to CAE+KNN(10 minutes) for accomplishing the same task. Thus, we believe it is feasible to detect astronomical outliers in the data of galaxy images in an unsupervised manner. Next, we will apply attCAE KNN to available survey datasets to assess its applicability and reliability.

cs.CV