Searcharxiv⌕ Search

arXiv subjects

Ge Liu

Publications and source records attributed to Ge Liu.

66 records · Page 4Linked to original sources

Neural P$^3$M: A Long-Range Interaction Modeling Enhancer for Geometric GNNs

Geometric graph neural networks (GNNs) have emerged as powerful tools for modeling molecular geometry. However, they encounter limitations in effectively capturing long-range interactions in large molecular systems. To address this challenge, we introduce Neural P$^3$M, a versatile enhancer of geometric GNNs to expand the scope of their capabilities by incorporating mesh points alongside atoms and reimaging traditional mathematical operations in a trainable manner. Neural P$^3$M exhibits flexibility across a wide range of molecular systems and demonstrates remarkable accuracy in predicting energies and forces, outperforming on benchmarks such as the MD22 dataset. It also achieves an average improvement of 22% on the OE62 dataset while integrating with various architectures.

cs.LG↗

Paper Copilot: A Self-Evolving and Efficient LLM System for Personalized Academic Assistance

As scientific research proliferates, researchers face the daunting task of navigating and reading vast amounts of literature. Existing solutions, such as document QA, fail to provide personalized and up-to-date information efficiently. We present Paper Copilot, a self-evolving, efficient LLM system designed to assist researchers, based on thought-retrieval, user profile and high performance optimization. Specifically, Paper Copilot can offer personalized research services, maintaining a real-time updated database. Quantitative evaluation demonstrates that Paper Copilot saves 69.92\% of time after efficient deployment. This paper details the design and implementation of Paper Copilot, highlighting its contributions to personalized academic support and its potential to streamline the research process.

cs.CL↗

Off-Policy Evaluation from Logged Human Feedback

Learning from human feedback has been central to recent advances in artificial intelligence and machine learning. Since the collection of human feedback is costly, a natural question to ask is if the new feedback always needs to collected. Or could we evaluate a new model with the human feedback on responses of another model? This motivates us to study off-policy evaluation from logged human feedback. We formalize the problem, propose both model-based and model-free estimators for policy values, and show how to optimize them. We analyze unbiasedness of our estimators and evaluate them empirically. Our estimators can predict the absolute values of evaluated policies, rank them, and be optimized.

cs.LG↗

Experimental Design for Active Transductive Inference in Large Language Models

One emergent ability of large language models (LLMs) is that query-specific examples can be included in the prompt at inference time. In this work, we use active learning for adaptive prompt design and call it Active In-context Prompt Design (AIPD). We design the LLM prompt by adaptively choosing few-shot examples from a training set to optimize performance on a test set. The training examples are initially unlabeled and we obtain the label of the most informative ones, which maximally reduces uncertainty in the LLM prediction. We propose two algorithms, GO and SAL, which differ in how the few-shot examples are chosen. We analyze these algorithms in linear models: first GO and then use its equivalence with SAL. We experiment with many different tasks in small, medium-sized, and large language models; and show that GO and SAL outperform other methods for choosing few-shot examples in the LLM prompt at inference time.

cs.LG↗

Pessimistic Off-Policy Multi-Objective Optimization

Multi-objective optimization is a type of decision making problems where multiple conflicting objectives are optimized. We study offline optimization of multi-objective policies from data collected by an existing policy. We propose a pessimistic estimator for the multi-objective policy values that can be easily plugged into existing formulas for hypervolume computation and optimized. The estimator is based on inverse propensity scores (IPS), and improves upon a naive IPS estimator in both theory and experiments. Our analysis is general, and applies beyond our IPS estimators and methods for optimizing them. The pessimistic estimator can be optimized by policy gradients and performs well in all of our experiments.

cs.LG↗

Maximum n-times Coverage for Vaccine Design

We introduce the maximum $n$-times coverage problem that selects $k$ overlays to maximize the summed coverage of weighted elements, where each element must be covered at least $n$ times. We also define the min-cost $n$-times coverage problem where the objective is to select the minimum set of overlays such that the sum of the weights of elements that are covered at least $n$ times is at least $τ$. Maximum $n$-times coverage is a generalization of the multi-set multi-cover problem, is NP-complete, and is not submodular. We introduce two new practical solutions for $n$-times coverage based on integer linear programming and sequential greedy optimization. We show that maximum $n$-times coverage is a natural way to frame peptide vaccine design, and find that it produces a pan-strain COVID-19 vaccine design that is superior to 29 other published designs in predicted population coverage and the expected number of peptides displayed by each individual's HLA molecules.

q-bio.QM↗

Information Condensing Active Learning

We introduce Information Condensing Active Learning (ICAL), a batch mode model agnostic Active Learning (AL) method targeted at Deep Bayesian Active Learning that focuses on acquiring labels for points which have as much information as possible about the still unacquired points. ICAL uses the Hilbert Schmidt Independence Criterion (HSIC) to measure the strength of the dependency between a candidate batch of points and the unlabeled set. We develop key optimizations that allow us to scale our method to large unlabeled sets. We show significant improvements in terms of model accuracy and negative log likelihood (NLL) on several image datasets compared to state of the art batch mode AL methods for deep learning.

cs.LG↗

Maximizing Overall Diversity for Improved Uncertainty Estimates in Deep Ensembles

The inaccuracy of neural network models on inputs that do not stem from the training data distribution is both problematic and at times unrecognized. Model uncertainty estimation can address this issue, where uncertainty estimates are often based on the variation in predictions produced by a diverse ensemble of models applied to the same input. Here we describe Maximize Overall Diversity (MOD), a straightforward approach to improve ensemble-based uncertainty estimates by encouraging larger overall diversity in ensemble predictions across all possible inputs that might be encountered in the future. When applied to various neural network ensembles, MOD significantly improves predictive performance for out-of-distribution test examples without sacrificing in-distribution performance on 38 Protein-DNA binding regression datasets, 9 UCI datasets, and the IMDB-Wiki image dataset. Across many Bayesian optimization tasks, the performance of UCB acquisition is also greatly improved by leveraging MOD uncertainty estimates.

cs.LG↗

Data Efficient Training for Reinforcement Learning with Adaptive Behavior Policy Sharing

Deep Reinforcement Learning (RL) is proven powerful for decision making in simulated environments. However, training deep RL model is challenging in real world applications such as production-scale health-care or recommender systems because of the expensiveness of interaction and limitation of budget at deployment. One aspect of the data inefficiency comes from the expensive hyper-parameter tuning when optimizing deep neural networks. We propose Adaptive Behavior Policy Sharing (ABPS), a data-efficient training algorithm that allows sharing of experience collected by behavior policy that is adaptively selected from a pool of agents trained with an ensemble of hyper-parameters. We further extend ABPS to evolve hyper-parameters during training by hybridizing ABPS with an adapted version of Population Based Training (ABPS-PBT). We conduct experiments with multiple Atari games with up to 16 hyper-parameter/architecture setups. ABPS achieves superior overall performance, reduced variance on top 25% agents, and equivalent performance on the best agent compared to conventional hyper-parameter tuning with independent training, even though ABPS only requires the same number of environmental interactions as training a single agent. We also show that ABPS-PBT further improves the convergence speed and reduces the variance.

cs.LG↗

Multi-party quantum privacy comparison of size based on d-level GHZ states

Quantum privacy comparison(QPC) plays an important role in secret ballot elections, private auctions and so on. To date, many multi-party QPC(MQPC) protocols have been proposed to compare the equality of $k(k\geq 3)$ participants. However, there are few examples of MQPC used to compare the sizes or values of their privacies. In this paper, we propose a MQPC protocol by which any $k(k\geq 3)$ participants can compare the sizes of their privacies with executing the protocol just once. The proposed MQPC protocol takes the $d-level$ GHZ states as quantum resources, and a semi-honest $TP$ is introduced to help the participants to determine the relationship of their privacies. Further more, only single-particle unitary transformations and measurements are involved, and the participants need not to share common secrets with each other beforehand which makes the proposed protocol much more efficient. Analysis shows that our protocol is secure against internal and external attack in theory.

quant-ph↗

Further Investigation on Classical Multiparty computation using Quantum Resources

The tremendous development of cloud computing and network technology makes it possible for multiple people with limited resources to complete a large-scale computing with the help of cloud servers. In order to protect the privacy of clients, secure multiparty computation plays an important role in the process of computing. Recently, Clementi et al[\textcolor[rgb]{0.00,0.07,1.00}{Phys. Rev. A {\bf 96}, 062317(2017)}] proposed a secure multiparty computation protocol using quantum resources. In their protocol, utilizing only linear classical computing and limited manipulation of quantum information, a method of computing $n-variable$ symmetric Boolean function $f(x_1, x_2, \cdots, x_n)$ with degree 2 is proposed, and all clients can jointly compute $f(x_1, x_2, \cdots, x_n)$ without revealing their private inputs with the help of a sever. They proposed an open problem: are there more simple nonlinear functions like the one presented by them that can be used as subroutines for larger computation protocols? We will give the answer to this question in this paper. Inspired by Clementi et al's work, we continue to explore the quantum realization of Boolean functions. First, we demonstrate a way to compute a class of $n-variable$ symmetric Boolean function $f_n^k$ by using single-particle quantum state $|0\rangle$ and single-particle unitary operations $U_k$. Second, we show that each $n-variable$ symmetric Boolean function can be represented by the linear combination of $f_n^k(k=0,1,\cdots,n)$ and each function $f_n^k(2\leq k\leq n)$ can be used to perform secure multiparty computation. Third, we propose an universal quantum implementation method for arbitrary $n-variable$ symmetric Boolean function $f(x_1, x_2, \cdots, x_n)$. Finally, we demonstrate our secure multiparty computation protocol on IBM quantum cloud platform.

quant-ph↗

Teleportation of an arbitrary multipartite state via photonic Faraday rotation

We propose a practical scheme for deterministically teleporting an arbitrary multipartite state, either product or entangled, using Faraday rotation of the photonic polarization. Our scheme, based on the input-output process of single-photon pulses regarding cavities, works in low-Q cavities and only involves virtual excitation of the atoms, which is insensitive to both cavity decay and atomic spontaneous emission. Besides, the Bell-state measurement is accomplished by the Faraday rotation plus product-state measurements, which could much relax the experimental difficulty to realize the Bell-state measurement by the CNOT operation.

quant-ph↗