Searcharxiv⌕ Search

arXiv subjects

Ying Fan

Publications and source records attributed to Ying Fan.

At least 37 records · Page 2Linked to original sources

Dimensionality reduction of networked systems with separable coupling-dynamics: theory and applications

Complex dynamical systems are prevalent in various domains, but their analysis and prediction are hindered by their high dimensionality and nonlinearity. Dimensionality reduction techniques can simplify the system dynamics by reducing the number of variables, but most existing methods do not account for networked systems with separable coupling-dynamics, where the interaction between nodes can be decomposed into a function of the node state and a function of the neighbor state. Here, we present a novel dimensionality reduction framework that can effectively capture the global dynamics of these networks by projecting them onto a low-dimensional system. We derive the reduced system's equation and stability conditions, and propose an error metric to quantify the reduction accuracy. We demonstrate our framework on two examples of networked systems with separable coupling-dynamics: a modified susceptible-infected-susceptible model with direct infection and a modified Michaelis-Menten model with activation and inhibition. We conduct numerical experiments on synthetic and empirical networks to validate and evaluate our framework, and find a good agreement between the original and reduced systems. We also investigate the effects of different network structures and parameters on the system dynamics and the reduction error. Our framework offers a general and powerful tool for studying complex dynamical networks with separable coupling-dynamics.

math.DS↗

DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models

Learning from human feedback has been shown to improve text-to-image models. These techniques first learn a reward function that captures what humans care about in the task and then improve the models based on the learned reward function. Even though relatively simple approaches (e.g., rejection sampling based on reward scores) have been investigated, fine-tuning text-to-image models with the reward function remains challenging. In this work, we propose using online reinforcement learning (RL) to fine-tune text-to-image models. We focus on diffusion models, defining the fine-tuning task as an RL problem, and updating the pre-trained text-to-image diffusion models using policy gradient to maximize the feedback-trained reward. Our approach, coined DPOK, integrates policy optimization with KL regularization. We conduct an analysis of KL regularization for both RL fine-tuning and supervised fine-tuning. In our experiments, we show that DPOK is generally superior to supervised fine-tuning with respect to both image-text alignment and image quality. Our code is available at https://github.com/google-research/google-research/tree/master/dpok.

cs.LG↗

Impact of resource availability and conformity effect on sustainability of common-pool resources

Sustainability of common-pool resources hinges on the interplay between human and environmental systems. However, there is still a lack of a novel and comprehensive framework for modelling extraction of common-pool resources and cooperation of human agents that can account for different factors that shape the system behavior and outcomes. In particular, we still lack a critical value for ensuring resource sustainability under different scenarios. In this paper, we present a novel framework for studying resource extraction and cooperation in human-environmental systems for common-pool resources. We explore how different factors, such as resource availability and conformity effect, influence the players' decisions and the resource outcomes. We identify critical values for ensuring resource sustainability under various scenarios. We demonstrate the observed phenomena are robust to the complexity and assumptions of the models and discuss implications of our study for policy and practice, as well as the limitations and directions for future research.

econ.TH↗

Disruptive papers in science are losing impact

The impact and originality are two critical dimensions for evaluating scientific publications, measured by citation and disruption metrics respectively. Despite the extensive effort made to understand the statistical properties and evolution of each of these metrics, the relations between the two remain unclear. In this paper, we study the evolution during last 70 years of the correlation between scientific papers' citation and disruption, finding surprisingly a decreasing trend from positive to negative correlations over the years. Consequently, during the years, there are fewer and fewer disruptive works among the highly cited papers. These results suggest that highly disruptive studies nowadays attract less attention from the scientific community. The analysis on papers' references supports this trend, showing that papers citing older references, less popular references and diverse references become to have less citations. Possible explanations for the less attention phenomenon could be due to the increasing information overload in science, and citations become more and more prominent for impact. This is supported by the evidence that research fields with more papers have a more negative correlation between citation and disruption. Finally, we show the generality of our findings by analyzing and comparing six disciplines.

cs.DL↗

Score-based Generative Modeling Secretly Minimizes the Wasserstein Distance

Score-based generative models are shown to achieve remarkable empirical performances in various applications such as image generation and audio synthesis. However, a theoretical understanding of score-based diffusion models is still incomplete. Recently, Song et al. showed that the training objective of score-based generative models is equivalent to minimizing the Kullback-Leibler divergence of the generated distribution from the data distribution. In this work, we show that score-based models also minimize the Wasserstein distance between them under suitable assumptions on the model. Specifically, we prove that the Wasserstein distance is upper bounded by the square root of the objective function up to multiplicative constants and a fixed constant offset. Our proof is based on a novel application of the theory of optimal transport, which can be of independent interest to the society. Our numerical experiments support our findings. By analyzing our upper bounds, we provide a few techniques to obtain tighter upper bounds.

cs.LG↗

Predicting the cascading dynamics in complex networks via the bimodal failure size distribution

Cascading failure as a systematic risk occurs in a wide range of real-world networks. Cascade size distribution is a basic and crucial characteristic of systemic cascade behaviors. Recent research works have revealed that the distribution of cascade sizes is a bimodal form indicating the existence of either very small cascades or large ones. In this paper, we aim to understand the properties and formation of such bimodal distribution of cascade sizes in complex networks, and further predict the final cascade size. We first find that the bimodal distribution of cascade sizes is ubiquitous in both synthetic and real networks. Moreover, the large cascade sizes distributed in the right peak of bimodal distribution are resulted from either the failure of nodes with high load at the first step of the cascade or multiple rounds of cascades triggered by the initial failure. Accordingly, we propose a hybrid load metric (HLM), which combines the load of the initial broken node and the load of failed nodes triggered by the initial failure, to predict the final size of cascading failures. Finally, we validate the effectiveness of HLM by computing the accuracy of identifying the cascades belonging to the right and left peaks of the bimodal distribution. The results show that HLM is a better predictor than commonly used network centrality metrics in both synthetic and real-world networks.

physics.soc-ph↗

Impactful scientists have higher tendency to involve collaborators in new topics

In scientific research, collaboration is one of the most effective ways to take advantage of new ideas, skills, resources, and for performing interdisciplinary research. Although collaboration networks have been intensively studied, the question of how individual scientists choose collaborators to study a new research topic remains almost unexplored. Here, we investigate the statistics and mechanisms of collaborations of individual scientists along their careers, revealing that, in general, collaborators are involved in significantly fewer topics than expected from controlled surrogate. In particular, we find that highly productive scientists tend to have higher fraction of single-topic collaborators, while highly cited, i.e., impactful, scientists have higher fraction of multi-topic collaborators. We also suggest a plausible mechanism for this distinction. Moreover, we investigate the cases where scientists involve existing collaborators into a new topic. We find that compared to productive scientists, impactful scientists show strong preference of collaboration with high impact scientists on a new topic. Finally, we validate our findings by investigating active scientists in different years and across different disciplines.

cs.DL↗

Academic mentees succeed in big groups, but thrive in small groups

Mentoring is a key component of scientific achievements, contributing to overall measures of career success for mentees and mentors. A common success metric in the scientific enterprise is acquiring a large research group, which is believed to indicate excellent mentorship and high-quality research. However, large, competitive groups might also amplify dropout rates, which are high especially among early career researchers. Here, we collect longitudinal genealogical data on mentor-mentee relations and their publication, and study the effects of a mentor's group on future academic survival and performance of their mentees. We find that mentees trained in large groups generally have better academic performance than mentees from small groups, if they continue working in academia after graduation. However, we also find two surprising results: Academic survival rate is significantly lower for (1) mentees from larger groups, and for (2) mentees with more productive mentors. These findings reveal that success of mentors has a negative effect on the academic survival rate of mentees, raising important questions about the definition of successful mentorship and providing actionable suggestions concerning career development.

physics.data-an↗

POEM: Out-of-Distribution Detection with Posterior Sampling

Out-of-distribution (OOD) detection is indispensable for machine learning models deployed in the open world. Recently, the use of an auxiliary outlier dataset during training (also known as outlier exposure) has shown promising performance. As the sample space for potential OOD data can be prohibitively large, sampling informative outliers is essential. In this work, we propose a novel posterior sampling-based outlier mining framework, POEM, which facilitates efficient use of outlier data and promotes learning a compact decision boundary between ID and OOD data for improved detection. We show that POEM establishes state-of-the-art performance on common benchmarks. Compared to the current best method that uses a greedy sampling strategy, POEM improves the relative performance by 42.0% and 24.2% (FPR95) on CIFAR-10 and CIFAR-100, respectively. We further provide theoretical insights on the effectiveness of POEM for OOD detection.

cs.LG↗

Model-based Reinforcement Learning for Continuous Control with Posterior Sampling

Balancing exploration and exploitation is crucial in reinforcement learning (RL). In this paper, we study model-based posterior sampling for reinforcement learning (PSRL) in continuous state-action spaces theoretically and empirically. First, we show the first regret bound of PSRL in continuous spaces which is polynomial in the episode length to the best of our knowledge. With the assumption that reward and transition functions can be modeled by Bayesian linear regression, we develop a regret bound of $\tilde{O}(H^{3/2}d\sqrt{T})$, where $H$ is the episode length, $d$ is the dimension of the state-action space, and $T$ indicates the total time steps. This result matches the best-known regret bound of non-PSRL methods in linear MDPs. Our bound can be extended to nonlinear cases as well with feature embedding: using linear kernels on the feature representation $ϕ$, the regret bound becomes $\tilde{O}(H^{3/2}d_ϕ\sqrt{T})$, where $d_ϕ$ is the dimension of the representation space. Moreover, we present MPC-PSRL, a model-based posterior sampling algorithm with model predictive control for action selection. To capture the uncertainty in models, we use Bayesian linear regression on the penultimate layer (the feature representation layer $ϕ$) of neural networks. Empirical results show that our algorithm achieves the state-of-the-art sample efficiency in benchmark continuous control tasks compared to prior model-based algorithms, and matches the asymptotic performance of model-free algorithms.

cs.LG↗

Real Negatives Matter: Continuous Training with Real Negatives for Delayed Feedback Modeling

One of the difficulties of conversion rate (CVR) prediction is that the conversions can delay and take place long after the clicks. The delayed feedback poses a challenge: fresh data are beneficial to continuous training but may not have complete label information at the time they are ingested into the training pipeline. To balance model freshness and label certainty, previous methods set a short waiting window or even do not wait for the conversion signal. If conversion happens outside the waiting window, this sample will be duplicated and ingested into the training pipeline with a positive label. However, these methods have some issues. First, they assume the observed feature distribution remains the same as the actual distribution. But this assumption does not hold due to the ingestion of duplicated samples. Second, the certainty of the conversion action only comes from the positives. But the positives are scarce as conversions are sparse in commercial systems. These issues induce bias during the modeling of delayed feedback. In this paper, we propose DElayed FEedback modeling with Real negatives (DEFER) method to address these issues. The proposed method ingests real negative samples into the training pipeline. The ingestion of real negatives ensures the observed feature distribution is equivalent to the actual distribution, thus reducing the bias. The ingestion of real negatives also brings more certainty information of the conversion. To correct the distribution shift, DEFER employs importance sampling to weigh the loss function. Experimental results on industrial datasets validate the superiority of DEFER. DEFER have been deployed in the display advertising system of Alibaba, obtaining over 6.0% improvement on CVR in several scenarios. The code and data in this paper are now open-sourced {https://github.com/gusuperstar/defer.git}.

cs.LG↗

The critical role of fresh teams in creating original and multi-disciplinary research

Teamwork is one of the most prominent features in modern science. It is now well-understood that the team size is an important factor that affects team creativity. However, the crucial question of how the character of research studies is influenced by the freshness of the team remains unclear. In this paper, we quantify the team freshness according to the absent of prior collaboration among team members. Our results suggest that fresher teams tend to produce works of higher originality and more multi-disciplinary impact. These effects are even magnified in larger teams. Furthermore, we find that freshness defined by new team members in a paper is a more effective indicator of research originality and multi-disciplinarity compared to freshness defined by new collaboration relations among team members. Finally, we show that career freshness of members also plays an important role in increasing the originality and multi-disciplinarity of produced papers.

physics.soc-ph↗

Search-based User Interest Modeling with Lifelong Sequential Behavior Data for Click-Through Rate Prediction

Rich user behavior data has been proven to be of great value for click-through rate prediction tasks, especially in industrial applications such as recommender systems and online advertising. Both industry and academy have paid much attention to this topic and propose different approaches to modeling with long sequential user behavior data. Among them, memory network based model MIMN proposed by Alibaba, achieves SOTA with the co-design of both learning algorithm and serving system. MIMN is the first industrial solution that can model sequential user behavior data with length scaling up to 1000. However, MIMN fails to precisely capture user interests given a specific candidate item when the length of user behavior sequence increases further, say, by 10 times or more. This challenge exists widely in previously proposed approaches. In this paper, we tackle this problem by designing a new modeling paradigm, which we name as Search-based Interest Model (SIM). SIM extracts user interests with two cascaded search units: (i) General Search Unit acts as a general search from the raw and arbitrary long sequential behavior data, with query information from candidate item, and gets a Sub user Behavior Sequence which is relevant to candidate item; (ii) Exact Search Unit models the precise relationship between candidate item and SBS. This cascaded search paradigm enables SIM with a better ability to model lifelong sequential behavior data in both scalability and accuracy. Apart from the learning algorithm, we also introduce our hands-on experience on how to implement SIM in large scale industrial systems. Since 2019, SIM has been deployed in the display advertising system in Alibaba, bringing 7.1\% CTR and 4.4\% RPM lift, which is significant to the business. Serving the main traffic in our real system now, SIM models user behavior data with maximum length reaching up to 54000, pushing SOTA to 54x.

cs.IR↗

Prediction Model Based on Integrated Political Economy System: The Case of US Presidential Election

This paper studies an integrated system of political and economic systems from a systematic perspective to explore the complex interaction between them, and specially analyzes the case of the US presidential election forecasting. Based on the signed association networks of industrial structure constructed by economic data, our framework simulates the diffusion and evolution of opinions during the election through a kinetic model called the Potts Model. Remarkably, we propose a simple and efficient prediction model for the US presidential election, and meanwhile inspire a new way to model the economic structure. Findings also highlight the close relationship between economic structure and political attitude. Furthermore, the case analysis in terms of network and economy demonstrates the specific features and the interaction between political tendency and industrial structure in a particular period, which is consistent with theories in politics and economics.

physics.soc-ph↗

A hyperbolic Embedding Model for Directed Networks

Network embedding is a fervid topic in current networks science and observes that most real complex systems can be embedded in hidden metrics space and emerge as the geometrical property, where the geometric distance between nodes determines the likelihood of links connected. Among those, hyperbolic space associated with the structural organization of many real complex systems, it has thus received extensive attention. However, the majority of methods and measurements, recently developed, less take these features into consideration for the asymmetry of links. Here, we discuss how to multiplex node information as an embedding foundation through identifying the bipartite structure of directed networks; and we proposed the generally mapping framework which hybrids the topological structure of complex networks, directed links and the hidden metrics space. By splitting the different properties of a node, possibilities between different types of nodes can be modeled. In addition to that, we apply this model to some real systems, including international trade networks and C.elegans neural networks. Results confirm that directed networks enable mapping into metrics space as well, and network embedding information can improve the scope of application of existing models.

physics.soc-ph↗

Efficient Model-Free Reinforcement Learning Using Gaussian Process

Efficient Reinforcement Learning usually takes advantage of demonstration or good exploration strategy. By applying posterior sampling in model-free RL under the hypothesis of GP, we propose Gaussian Process Posterior Sampling Reinforcement Learning(GPPSTD) algorithm in continuous state space, giving theoretical justifications and empirical results. We also provide theoretical and empirical results that various demonstration could lower expected uncertainty and benefit posterior sampling exploration. In this way, we combined the demonstration and exploration process together to achieve a more efficient reinforcement learning.

cs.LG↗

Increasing trend of scientists to switch between topics

We analyze the publication records of individual scientists, aiming to quantify the topic switching dynamics of scientists and its influence. For each scientist, the relations among her publications are characterized via shared references. We find that the co-citing network of the papers of a scientist exhibits a clear community structure where each major community represents a research topic. Our analysis suggests that scientists tend to have a narrow distribution of the number of topics. However, researchers nowadays switch more frequently between topics than those in the early days. We also find that high switching probability in early career (<12y) is associated with low overall productivity, while it is correlated with high overall productivity in latter career. Interestingly, the average citation per paper, however, is in all career stages negatively correlated with the switching probability. We propose a model with exploitation and exploration mechanisms that can explain the main observed features.

physics.soc-ph↗

Deep Interest Evolution Network for Click-Through Rate Prediction

Click-through rate~(CTR) prediction, whose goal is to estimate the probability of the user clicks, has become one of the core tasks in advertising systems. For CTR prediction model, it is necessary to capture the latent user interest behind the user behavior data. Besides, considering the changing of the external environment and the internal cognition, user interest evolves over time dynamically. There are several CTR prediction methods for interest modeling, while most of them regard the representation of behavior as the interest directly, and lack specially modeling for latent interest behind the concrete behavior. Moreover, few work consider the changing trend of interest. In this paper, we propose a novel model, named Deep Interest Evolution Network~(DIEN), for CTR prediction. Specifically, we design interest extractor layer to capture temporal interests from history behavior sequence. At this layer, we introduce an auxiliary loss to supervise interest extracting at each step. As user interests are diverse, especially in the e-commerce system, we propose interest evolving layer to capture interest evolving process that is relative to the target item. At interest evolving layer, attention mechanism is embedded into the sequential structure novelly, and the effects of relative interests are strengthened during interest evolution. In the experiments on both public and industrial datasets, DIEN significantly outperforms the state-of-the-art solutions. Notably, DIEN has been deployed in the display advertisement system of Taobao, and obtained 20.7\% improvement on CTR.

stat.ML↗