SearcharxivSearch

arXiv subjects

Yuren Zhou

Publications and source records attributed to Yuren Zhou.

At least 19 recordsLinked to original sources

Bayesian Latent Class Regression with Interpretable Binary Profiles

High-dimensional categorical data arise in diverse scientific domains and are often accompanied by covariates. Latent class regression models are routinely used in such settings, reducing dimensionality by assuming conditional independence of the categorical variables given a single latent class that depends on covariates through a logistic regression model. However, such methods become unreliable as the dimensionality increases. To address this, we propose Bayesian latent class regression with interpretable binary profiles (BLIP), a flexible family of models that introduces a binary latent-attribute layer between the covariate-dependent latent class and the observed categorical responses. BLIP satisfies key theoretical properties, including identifiability and posterior consistency, and we establish a Bayes oracle clustering property that ensures robustness against the curse of dimensionality. We develop efficient posterior computation methods, validate them through simulation studies, and use BLIP to infer regions of common profile in ecological data.

stat.ME

Bayesian Deep Generative Models for Multiplex Networks with Multiscale Overlapping Clusters

Our interest is in multiplex network data with multiple network samples observed across the same set of nodes. Examples originate from a variety of fields, including brain connectivity, international trade networks, and social networks, among others. Our goal is to infer a hierarchical structure of the nodes at a population level, while performing multi-resolution clustering of the individual replicates. To accomplish this, we propose a Bayesian hierarchical model, provide theoretical support in terms of identifiability and posterior consistency, and design efficient methods for posterior computation. We provide novel technical tools for proving model identifiability, which are of independent interest. Our proposed methodology is demonstrated through numerical simulation and an application to brain connectome data.

stat.ME

The origin of double-peaked narrow emission-line galaxies in MaNGA Survey

We select 36 double-peaked narrow emission-line galaxies (DPGs) from 10,010 unique galaxies in MaNGA survey. These DPGs show double-peaked Balmer lines and forbidden lines in the spectra. We use a double Gaussian model to separate the double-peaked profiles of each emission line into blue and red components ($λ_\text{blue}$ < $λ_\text{red}$), and analyze the spatially resolved kinematics and ionization mechanisms of each component. We find that in 35 out of 36 DPGs, the flux ratio between the blue and red components varies systematically along the major axes, while it keeps roughly a constant along the minor axes. The blue and red components of these DPGs exhibit similar distributions in both the value of line-of-sight velocity and the velocity dispersion. Additionally, 83.3% DPGs have both blue and red components located in the same ionization region in the [SII]-BPT diagram. Combining all these observational results, we suggest that the double-peaked emission line profiles in these 35 DPGs primarily originate from rotating discs. The remaining one galaxy shows clear outflow features. 8 out of 35 DPGs show symmetric line profiles that indicate undisturbed rotating discs, and the other 27 DPGs exhibit asymmetric profiles, suggesting dynamic disturbances in the rotating discs. Furthermore, we find that 58.3% DPGs experienced external processes, characterized by tidal features, companion galaxies, as well as gas-star misalignments. This fraction is about twice as much as that of the control sample, suggesting the origin of double-peaked emission line profiles is associated with external processes.

astro-ph.GA

Misaligned external gas acquisition boosts central black hole activities

One important question in active galactic nucleus (AGN) is how gas is brought down to the galaxy center. Both internal secular evolution (torque induced by non-axisymmetric galactic structures such as bars) and external processes (e.g. mergers or interactions) are expected to redistribute the angular momentum (AM) and transport gas inward. However, it is still under debate whether these processes can significantly affect AGN activities. Here we for the first time report that AGN fraction increases with the difference of kinematic position angles ($ΔPA\equiv|PA_{\mathrm{gas}}-PA_{\mathrm{star}}|$) between ionized gas ($PA_{\mathrm{gas}}$) and stellar disks ($PA_{\mathrm{star}}$) in blue and green galaxies, meanwhile this fraction remains roughly constant for red galaxies. Also the high luminosity AGN fraction increases with $ΔPA$ while the low luminosity AGN fraction is independent with $ΔPA$. These observational results support a scenario in which the interaction between accreted and pre-existing gas provides the AM loss mechanism, thereby the gas inflow fuels the central BH activities, and the AM loss efficiency is positively correlated with the $ΔPA$.

astro-ph.GA

Misaligned gas acquisition as a formation pathway of S0 galaxies

We analyze a sample of 753 S0 galaxies from the MPL-10 of MaNGA survey and investigate the gas-star kinematic misalignment and merger remnant fraction in galaxies with different morphological types. The misalign fraction in S0s is the highest among all the morphological types for both young (global $\mathrm{D}_n4000<1.6$, $\sim$15%) and old (global $\mathrm{D}_n4000>1.6$, $\sim$10%) galaxies. We compare the properties of misaligned S0s with other types of galaxies, finding: (i) misaligned S0s and misaligned spirals have higher bulge luminosity, higher B/T and larger Sérsic index compared to spirals; (ii) the misaligned S0s have lower bulge luminosity $M_r$ and smaller bulge size than merger remnant S0s, while aligned S0s have the widest coverage for these parameter distributions which are overlapped with both misaligned S0s and merger remnant S0s; (iii) misaligned S0s have lower stellar mass $M_*$ and more isolated environment than aligned S0s and merger remnant S0s; (iv) the young misaligned S0s have positive $\mathrm{D}_n4000$ radial gradient, while the aligned S0s and merger remnant S0s show negative $\mathrm{D}_n4000$ radial gradient. Combining all these observational results, we suggest misaligned gas acquisition as another efficient formation pathway for S0 galaxies. The redistribution of gas angular momentum during gas-gas collision between accreted and pre-existing gas leads to gas inflow and the growth of bulge component, meanwhile the lack of cold gas at the outskirts leads to fading of spiral arms.

astro-ph.GA

Drift Analysis with Fitness Levels for Elitist Evolutionary Algorithms

The fitness level method is a popular tool for analyzing the hitting time of elitist evolutionary algorithms. Its idea is to divide the search space into multiple fitness levels and estimate lower and upper bounds on the hitting time using transition probabilities between fitness levels. However, the lower bound generated by this method is often loose. An open question regarding the fitness level method is what are the tightest lower and upper time bounds that can be constructed based on transition probabilities between fitness levels. To answer this question, {\color{red} we combine drift analysis with fitness levels and define the tightest bound problem as a constrained multi-objective optimization problem subject to fitness levels.} The tightest metric bounds from fitness levels are constructed and proven for the first time. Then linear bounds are derived from metric bounds and a framework is established that can be used to develop different fitness level methods for different types of linear bounds. The framework is generic and promising, as it can be used to draw tight time bounds on both fitness landscapes without and with shortcuts. This is demonstrated in the example of the (1+1) EA maximizing the TwoMax1 function

cs.NE

DsMtGCN: A Direction-sensitive Multi-task framework for Knowledge Graph Completion

To solve the inherent incompleteness of knowledge graphs (KGs), numbers of knowledge graph completion (KGC) models have been proposed to predict missing links from known triples. Among those, several works have achieved more advanced results via exploiting the structure information on KGs with Graph Convolutional Networks (GCN). However, we observe that entity embeddings aggregated from neighbors in different directions are just simply averaged to complete single-tasks by existing GCN based models, ignoring the specific requirements of forward and backward sub-tasks. In this paper, we propose a Direction-sensitive Multi-task GCN (DsMtGCN) to make full use of the direction information, the multi-head self-attention is applied to specifically combine embeddings in different directions based on various entities and sub-tasks, the geometric constraints are imposed to adjust the distribution of embeddings, and the traditional binary cross-entropy loss is modified to reflect the triple uncertainty. Moreover, the competitive experiments results on several benchmark datasets verify the effectiveness of our model.

cs.AI

Contextual Dictionary Lookup for Knowledge Graph Completion

Knowledge graph completion (KGC) aims to solve the incompleteness of knowledge graphs (KGs) by predicting missing links from known triples, numbers of knowledge graph embedding (KGE) models have been proposed to perform KGC by learning embeddings. Nevertheless, most existing embedding models map each relation into a unique vector, overlooking the specific fine-grained semantics of them under different entities. Additionally, the few available fine-grained semantic models rely on clustering algorithms, resulting in limited performance and applicability due to the cumbersome two-stage training process. In this paper, we present a novel method utilizing contextual dictionary lookup, enabling conventional embedding models to learn fine-grained semantics of relations in an end-to-end manner. More specifically, we represent each relation using a dictionary that contains multiple latent semantics. The composition of a given entity and the dictionary's central semantics serves as the context for generating a lookup, thus determining the fine-grained semantics of the relation adaptively. The proposed loss function optimizes both the central and fine-grained semantics simultaneously to ensure their semantic consistency. Besides, we introduce two metrics to assess the validity and accuracy of the dictionary lookup operation. We extend several KGE models with the method, resulting in substantial performance improvements on widely-used benchmark datasets.

cs.AI

Location-based Activity Behavior Deviation Detection for Nursing Home using IoT Devices

With the advancement of the Internet of Things(IoT) and pervasive computing applications, it provides a better opportunity to understand the behavior of the aging population. However, in a nursing home scenario, common sensors and techniques used to track an elderly living alone are not suitable. In this paper, we design a location-based tracking system for a four-story nursing home - The Salvation Army, Peacehaven Nursing Home in Singapore. The main challenge here is to identify the group activity among the nursing home's residents and to detect if they have any deviated activity behavior. We propose a location-based deviated activity behavior detection system to detect deviated activity behavior by leveraging data fusion technique. In order to compute the features for data fusion, an adaptive method is applied for extracting the group and individual activity time and generate daily hybrid norm for each of the residents. Next, deviated activity behavior detection is executed by considering the difference between daily norm patterns and daily input data for each resident. Lastly, the deviated activity behavior among the residents are classified using a rule-based classification approach. Through the implementation, there are 44.4% of the residents do not have deviated activity behavior , while 37% residents involved in one deviated activity behavior and 18.6% residents have two or more deviated activity behaviors.

cs.CY

Clustering and Analysis of GPS Trajectory Data using Distance-based Features

The proliferation of smartphones has accelerated mobility studies by largely increasing the type and volume of mobility data available. One such source of mobility data is from GPS technology, which is becoming increasingly common and helps the research community understand mobility patterns of people. However, there lacks a standardized framework for studying the different mobility patterns created by the non-Work, non-Home locations of Working and Nonworking users on Workdays and Offdays using machine learning methods. We propose a new mobility metric, Daily Characteristic Distance, and use it to generate features for each user together with Origin-Destination matrix features. We then use those features with an unsupervised machine learning method, $k$-means clustering, and obtain three clusters of users for each type of day (Workday and Offday). Finally, we propose two new metrics for the analysis of the clustering results, namely User Commonality and Average Frequency. By using the proposed metrics, interesting user behaviors can be discerned and it helps us to better understand the mobility patterns of the users.

cs.LG

SDSS-IV MaNGA: Global Properties of Kinematically Misaligned Galaxies

We select 456 gas-star kinematically misaligned galaxies from the internal Product Launch-10 of MaNGA survey, including 74 star-forming (SF), 136 green-valley (GV) and 206 quiescent (QS) galaxies. We find that the distributions of difference between gas and star position angles for galaxies have three local peaks at $\sim0^{\circ}$, $90^{\circ}$, $180^{\circ}$. The fraction of misaligned galaxies peaks at $\log(M_*/M_{\odot})\sim10.5$ and declines to both low and high mass end. This fraction decreases monotonically with increasing SFR and sSFR. We compare the global parameters including gas kinematic asymmetry $V_{\mathrm{asym}}$, HI detection rate and mass fraction of molecular gas, effective radius $R_e$, Sérsic index $n$ as well as spin parameter $λ_{R_e}$ between misaligned galaxies and their control samples. We find that the misaligned galaxies have lower HI detection rate and molecular gas mass fraction, smaller size, higher Sérsic index and lower spin parameters than their control samples. The SF and GV misaligned galaxies are more asymmetric in gas velocity fields than their controls. These observational evidences point to the gas accretion scenario followed by angular momentum redistribution from gas-gas collision, leading to gas inflow and central star formation for the SF and GV misaligned galaxies. We propose three possible origins of the misaligned QS galaxies: (1) external gas accretion; (2) merger; (3) GV misaligned galaxies evolve into QS galaxies.

astro-ph.GA

Distributed Evolution Strategies for Black-box Stochastic Optimization

This work concerns the evolutionary approaches to distributed stochastic black-box optimization, in which each worker can individually solve an approximation of the problem with nature-inspired algorithms. We propose a distributed evolution strategy (DES) algorithm grounded on a proper modification to evolution strategies, a family of classic evolutionary algorithms, as well as a careful combination with existing distributed frameworks. On smooth and nonconvex landscapes, DES has a convergence rate competitive to existing zeroth-order methods, and can exploit the sparsity, if applicable, to match the rate of first-order methods. The DES method uses a Gaussian probability model to guide the search and avoids the numerical issue resulted from finite-difference techniques in existing zeroth-order methods. The DES method is also fully adaptive to the problem landscape, as its convergence is guaranteed with any parameter setting. We further propose two alternative sampling schemes which significantly improve the sampling efficiency while leading to similar performance. Simulation studies on several machine learning problems suggest that the proposed methods show much promise in reducing the convergence time and improving the robustness to parameter settings.

cs.NE

MMES: Mixture Model based Evolution Strategy for Large-Scale Optimization

This work provides an efficient sampling method for the covariance matrix adaptation evolution strategy (CMA-ES) in large-scale settings. In contract to the Gaussian sampling in CMA-ES, the proposed method generates mutation vectors from a mixture model, which facilitates exploiting the rich variable correlations of the problem landscape within a limited time budget. We analyze the probability distribution of this mixture model and show that it approximates the Gaussian distribution of CMA-ES with a controllable accuracy. We use this sampling method, coupled with a novel method for mutation strength adaptation, to formulate the mixture model based evolution strategy (MMES) -- a CMA-ES variant for large-scale optimization. The numerical simulations show that, while significantly reducing the time complexity of CMA-ES, MMES preserves the rotational invariance, is scalable to high dimensional problems, and is competitive against the state-of-the-arts in performing global optimization.

cs.NE

SDSS-IV MaNGA : spatial resolved properties of kinematically misaligned galaxies

We select 456 galaxies with kinematically misaligned gas and stellar components from 9546 parent galaxies in MaNGA, and classify them into 72 star-forming galaxies, 142 green-valley galaxies and 242 quiescent galaxies. Comparing the spatial resolved properties of the misaligned galaxies with control samples closely match in the D$_n$4000 and stellar velocity dispersion, we find that: (1) the misaligned galaxies have lower values in $V_{\rm gas}/σ_{\rm gas}$ and $V_{\rm star}/σ_{\rm star}$ (the ratio between ordered to random motion of gas and stellar components) across the entire galaxies than their control samples; (2) the star-forming and green-valley misaligned galaxies have enhanced central concentrated star formation than their control galaxies. The difference in stellar population between quiescent misaligned galaxies and control samples is small; (3) gas-phase metallicity of the green valley and quiescent misaligned galaxies are lower than the control samples. For the star forming misaligned galaxies, the difference in metallicity between the misaligned galaxies and their control samples strongly depends on how we select the control samples. All these observational results suggest external gas accretion influences the evolution of star forming and green valley galaxies, not only in kinematics/morphologies, but also in stellar populations. However, the quiescent misaligned galaxies have survived from different formation mechanisms.

astro-ph.GA

WiFi Fingerprint Clustering for Urban Mobility Analysis

In this paper, we present an unsupervised learning approach to identify the user points of interest (POI) by exploiting WiFi measurements from smartphone application data. Due to the lack of GPS positioning accuracy in indoor, sheltered, and high rise building environments, we rely on widely available WiFi access points (AP) in contemporary urban areas to accurately identify POI and mobility patterns, by comparing the similarity in the WiFi measurements. We propose a system architecture to scan the surrounding WiFi AP, and perform unsupervised learning to demonstrate that it is possible to identify three major insights, namely the indoor POI within a building, neighbourhood activity, and micro-mobility of the users. Our results show that it is possible to identify the aforementioned insights, with the fusion of WiFi and GPS, which are not possible to identify by only using GPS.

cs.LG

There Once Was a Really Bad Poet, It Was Automated but You Didn't Know It

Limerick generation exemplifies some of the most difficult challenges faced in poetry generation, as the poems must tell a story in only five lines, with constraints on rhyme, stress, and meter. To address these challenges, we introduce LimGen, a novel and fully automated system for limerick generation that outperforms state-of-the-art neural network-based poetry models, as well as prior rule-based poetry models. LimGen consists of three important pieces: the Adaptive Multi-Templated Constraint algorithm that constrains our search to the space of realistic poems, the Multi-Templated Beam Search algorithm which searches efficiently through the space, and the probabilistic Storyline algorithm that provides coherent storylines related to a user-provided prompt word. The resulting limericks satisfy poetic constraints and have thematically coherent storylines, which are sometimes even funny (when we are lucky).

cs.CL

Multiple-Perspective Clustering of Passive Wi-Fi Sensing Trajectory Data

Information about the spatiotemporal flow of humans within an urban context has a wide plethora of applications. Currently, although there are many different approaches to collect such data, there lacks a standardized framework to analyze it. The focus of this paper is on the analysis of the data collected through passive Wi-Fi sensing, as such passively collected data can have a wide coverage at low cost. We propose a systematic approach by using unsupervised machine learning methods, namely k-means clustering and hierarchical agglomerative clustering (HAC) to analyze data collected through such a passive Wi-Fi sniffing method. We examine three aspects of clustering of the data, namely by time, by person, and by location, and we present the results obtained by applying our proposed approach on a real-world dataset collected over five months.

cs.LG

Stochastic Recursive Momentum for Policy Gradient Methods

In this paper, we propose a novel algorithm named STOchastic Recursive Momentum for Policy Gradient (STORM-PG), which operates a SARAH-type stochastic recursive variance-reduced policy gradient in an exponential moving average fashion. STORM-PG enjoys a provably sharp $O(1/ε^3)$ sample complexity bound for STORM-PG, matching the best-known convergence rate for policy gradient algorithm. In the mean time, STORM-PG avoids the alternations between large batches and small batches which persists in comparable variance-reduced policy gradient methods, allowing considerably simpler parameter tuning. Numerical experiments depicts the superiority of our algorithm over comparative policy gradient algorithms.

stat.ML