SearcharxivSearch

arXiv subjects

Yu Lin

Publications and source records attributed to Yu Lin.

At least 55 records · Page 3Linked to original sources

Remaining Useful Life Modelling with an Escalator Health Condition Analytic System

The refurbishment of an escalator is usually linked with its design life as recommended by the manufacturer. However, the actual useful life of an escalator should be determined by its operating condition which is affected by the runtime, workload, maintenance quality, vibration, etc., rather than age only. The objective of this project is to develop a comprehensive health condition analytic system for escalators to support refurbishment decisions. The analytic system consists of four parts: 1) online data gathering and processing; 2) a dashboard for condition monitoring; 3) a health index model; and 4) remaining useful life prediction. The results can be used for a) predicting the remaining useful life of the escalators, in order to support asset replacement planning and b) monitoring the real-time condition of escalators; including alerts when vibration exceeds the threshold and signal diagnosis, giving an indication of possible root cause (components) of the alert signal.

stat.AP

Mean-variance hybrid portfolio optimization with quantile-based risk measure

This paper addresses the importance of incorporating various risk measures in portfolio management and proposes a dynamic hybrid portfolio optimization model that combines the spectral risk measure and the Value-at-Risk in the mean-variance formulation. By utilizing the quantile optimization technique and martingale representation, we offer a solution framework for these issues and also develop a closed-form portfolio policy when all market parameters are deterministic. Our hybrid model outperforms the classical continuous-time mean-variance portfolio policy by allocating a higher position of the risky asset in favorable market states and a less risky asset in unfavorable market states. This desirable property leads to promising numerical experiment results, including improved Sortino ratio and reduced downside risk compared to the benchmark models.

q-fin.PM

DBE-KT22: A Knowledge Tracing Dataset Based on Online Student Evaluation

Online education has gained an increasing importance over the last decade for providing affordable high-quality education to students worldwide. This has been further magnified during the global pandemic as more students switched to study online. The majority of online education tasks, e.g., course recommendation, exercise recommendation, or automated evaluation, depends on tracking students' knowledge progress. This is known as the \emph{Knowledge Tracing} problem in the literature. Addressing this problem requires collecting student evaluation data that can reflect their knowledge evolution over time. In this paper, we propose a new knowledge tracing dataset named Database Exercises for Knowledge Tracing (DBE-KT22) that is collected from an online student exercise system in a course taught at the Australian National University in Australia. We discuss the characteristics of the DBE-KT22 dataset and contrast it with the existing datasets in the knowledge tracing literature. Our dataset is available for public access through the Australian Data Archive platform.

cs.CY

Controllable Fake Document Infilling for Cyber Deception

Recent works in cyber deception study how to deter malicious intrusion by generating multiple fake versions of a critical document to impose costs on adversaries who need to identify the correct information. However, existing approaches are context-agnostic, resulting in sub-optimal and unvaried outputs. We propose a novel context-aware model, Fake Document Infilling (FDI), by converting the problem to a controllable mask-then-infill procedure. FDI masks important concepts of varied lengths in the document, then infills a realistic but fake alternative considering both the previous and future contexts. We conduct comprehensive evaluations on technical documents and news stories. Results show that FDI outperforms the baselines in generating highly believable fakes with moderate modification to protect critical information and deceive adversaries.

cs.AI

Graph Coloring via Neural Networks for Haplotype Assembly and Viral Quasispecies Reconstruction

Understanding genetic variation, e.g., through mutations, in organisms is crucial to unravel their effects on the environment and human health. A fundamental characterization can be obtained by solving the haplotype assembly problem, which yields the variation across multiple copies of chromosomes. Variations among fast evolving viruses that lead to different strains (called quasispecies) are also deciphered with similar approaches. In both these cases, high-throughput sequencing technologies that provide oversampled mixtures of large noisy fragments (reads) of genomes, are used to infer constituent components (haplotypes or quasispecies). The problem is harder for polyploid species where there are more than two copies of chromosomes. State-of-the-art neural approaches to solve this NP-hard problem do not adequately model relations among the reads that are important for deconvolving the input signal. We address this problem by developing a new method, called NeurHap, that combines graph representation learning with combinatorial optimization. Our experiments demonstrate substantially better performance of NeurHap in real and synthetic datasets compared to competing approaches.

q-bio.GN

Improving Contextual Representation with Gloss Regularized Pre-training

Though achieving impressive results on many NLP tasks, the BERT-like masked language models (MLM) encounter the discrepancy between pre-training and inference. In light of this gap, we investigate the contextual representation of pre-training and inference from the perspective of word probability distribution. We discover that BERT risks neglecting the contextual word similarity in pre-training. To tackle this issue, we propose an auxiliary gloss regularizer module to BERT pre-training (GR-BERT), to enhance word semantic similarity. By predicting masked words and aligning contextual embeddings to corresponding glosses simultaneously, the word similarity can be explicitly modeled. We design two architectures for GR-BERT and evaluate our model in downstream tasks. Experimental results show that the gloss regularizer benefits BERT in word-level and sentence-level semantic representation. The GR-BERT achieves new state-of-the-art in lexical substitution task and greatly promotes BERT sentence representation in both unsupervised and supervised STS tasks.

cs.CL

Dirac nodal lines in the quasi-one-dimensional ternary telluride TaPtTe$_5$

A Dirac nodal-line phase, as a quantum state of topological materials, usually occur in three-dimensional or at least two-dimensional materials with sufficient symmetry operations that could protect the Dirac band crossings. Here, we report a combined theoretical and experimental study on the electronic structure of the quasi-one-dimensional ternary telluride TaPtTe$_5$, which is corroborated as being in a robust nodal-line phase with fourfold degeneracy. Our angle-resolved photoemission spectroscopy measurements show that two pairs of linearly dispersive Dirac-like bands exist in a very large energy window, which extend from a binding energy of $\sim$ 0.75 eV to across the Fermi level. The crossing points are at the boundary of Brillouin zone and form Dirac-like nodal lines. Using first-principles calculations, we demonstrate the existing of nodal surfaces on the $k_y = \pm π$ plane in the absence of spin-orbit coupling (SOC), which are protected by nonsymmorphic symmetry in TaPtTe$_5$. When SOC is included, the nodal surfaces are broken into several nodal lines. By theoretical analysis, we conclude that the nodal lines along $Y$-$T$ and the ones connecting the $R$ points are non-trivial and protected by nonsymmorphic symmetry against SOC.

cond-mat.mtrl-sci

Cesium-involved electron transfer and electron-electron interaction in high-pressure metallic CsPbI3

Electron-phonon coupling was believed to govern the carrier transport in halide perovskites and related phases. Here we demonstrate that electron-electron interaction plays a direct and prominent role in the low-temperature electrical transport of compressed CsPbI3 and renders Fermi liquid (FL)-like behavior. By compressing δ-CsPbI3 to 80 GPa, an insulator-to-metal transition occurs, concomitant with the completion of a sluggish structural transition from the one-dimensional (1D) Pnma (δ) phase to a 3D Pmn21 (ε) phase. Deviation from FL behavior is observed in CsPbI3 upon entering the metallic ε phase, which progressively evolves into a FL-like state at 186 GPa. First-principles density functional theory calculations reveal that the enhanced electron-electron coupling is related to the Cs-involved electron transfer and sudden increase of the 5d state occupation of the high-pressure ε phase. Our study presents a promising strategy for tuning the electronic interaction in halide perovskites for realizing intriguing electronic states.

cond-mat.mtrl-sci

RepBin: Constraint-based Graph Representation Learning for Metagenomic Binning

Mixed communities of organisms are found in many environments (from the human gut to marine ecosystems) and can have profound impact on human health and the environment. Metagenomics studies the genomic material of such communities through high-throughput sequencing that yields DNA subsequences for subsequent analysis. A fundamental problem in the standard workflow, called binning, is to discover clusters, of genomic subsequences, associated with the unknown constituent organisms. Inherent noise in the subsequences, various biological constraints that need to be imposed on them and the skewed cluster size distribution exacerbate the difficulty of this unsupervised learning problem. In this paper, we present a new formulation using a graph where the nodes are subsequences and edges represent homophily information. In addition, we model biological constraints providing heterophilous signal about nodes that cannot be clustered together. We solve the binning problem by developing new algorithms for (i) graph representation learning that preserves both homophily relations and heterophily constraints (ii) constraint-based graph clustering method that addresses the problems of skewed cluster size distribution. Extensive experiments, on real and synthetic datasets, demonstrate that our approach, called RepBin, outperforms a wide variety of competing methods. Our constraint-based graph representation learning and clustering methods, that may be useful in other domains as well, advance the state-of-the-art in both metagenomics binning and graph representation learning.

q-bio.GN

From Aircraft Tracking Data to Network Delay Model: A Data-Driven Approach Considering En-Route Congestion

En-route congestion causes delays in air traffic networks and will become more prominent as air traffic demand will continue to increase yet airspace volume cannot grow. However, most existing studies on flight delay modeling do not consider en-route congestion explicitly. In this study, we propose a new flight delay model, Multi-layer Air Traffic Network Delay (MATND) model, to capture the impact of en-route congestion on flight delays over an air traffic network. This model is developed by a data-driven approach, taking aircraft tracking data and flight schedules as inputs to characterize a national air traffic network, as well as a system-level model approach, modeling the delay process based on queueing theory. The two approaches combined make the network delay model a close representation of reality and easy-to-implement for what-if scenario analysis. The proposed MATND model includes 1) a data-driven method to learn a network composed of airports, en-route congestion points, and air corridors from aircraft tracking data, 2) a stochastic and dynamic queuing network model to calculate flight delays and track their propagation at both airports and in en-route congestion areas, in which the delays are computed via a space-time decomposition method. Using one month of historical aircraft tracking data over China's air traffic network, MATND is tested and shows to give an accurate quantification of delays of the national air traffic network. "What-if" scenario analyses are conducted to demonstrate how the proposed model can be used for the evaluation of air traffic network improvement strategies, where the manipulation of reality at such a scale is impossible. Results show that MATND is computationally efficient, well suited for evaluating the impact of policy alternatives on system-wide delay at a macroscopic level.

eess.SY

SSNE: Effective Node Representation for Link Prediction in Sparse Networks

Graph embedding is gaining its popularity for link prediction in complex networks and achieving excellent performance. However, limited work has been done in sparse networks that represent most of real networks. In this paper, we propose a model, Sparse Structural Network Embedding (SSNE), to obtain node representation for link predication in sparse networks. The SSNE first transforms the adjacency matrix into the Sum of Normalized $H$-order Adjacency Matrix (SNHAM), and then maps the SNHAM matrix into a $d$-dimensional feature matrix for node representation via a neural network model. The mapping operation is proved to be an equivalent variation of singular value decomposition. Finally, we calculate nodal similarities for link prediction based on such feature matrix. By extensive testing experiments bases on synthetic and real sparse network, we show that the proposed method presents better link prediction performance in comparison of those of structural similarity indexes, matrix optimization and other graph embedding models.

cs.SI

Query-by-Sketch: Scaling Shortest Path Graph Queries on Very Large Networks

Computing shortest paths is a fundamental operation in processing graph data. In many real-world applications, discovering shortest paths between two vertices empowers us to make full use of the underlying structure to understand how vertices are related in a graph, e.g. the strength of social ties between individuals in a social network. In this paper, we study the shortest-path-graph problem that aims to efficiently compute a shortest path graph containing exactly all shortest paths between any arbitrary pair of vertices on complex networks. Our goal is to design an exact solution that can scale to graphs with millions or billions of vertices and edges. To achieve high scalability, we propose a novel method, Query-by-Sketch (QbS), which efficiently leverages offline labelling (i.e., precomputed labels) to guide online searching through a fast sketching process that summarizes the important structural aspects of shortest paths in answering shortest-path-graph queries. We theoretically prove the correctness of this method and analyze its computational complexity. To empirically verify the efficiency of QbS, we conduct experiments on 12 real-world datasets, among which the largest dataset has 1.7 billion vertices and 7.8 billion edges. The experimental results show that QbS can answer shortest-path graph queries in microseconds for million-scale graphs and less than half a second for billion-scale graphs.

cs.DB

SetConv: A New Approach for Learning from Imbalanced Data

For many real-world classification problems, e.g., sentiment classification, most existing machine learning methods are biased towards the majority class when the Imbalance Ratio (IR) is high. To address this problem, we propose a set convolution (SetConv) operation and an episodic training strategy to extract a single representative for each class, so that classifiers can later be trained on a balanced class distribution. We prove that our proposed algorithm is permutation-invariant despite the order of inputs, and experiments on multiple large-scale benchmark text datasets show the superiority of our proposed framework when compared to other SOTA methods.

cs.IR

A Highly Scalable Labelling Approach for Exact Distance Queries in Complex Networks

Answering exact shortest path distance queries is a fundamental task in graph theory. Despite a tremendous amount of research on the subject, there is still no satisfactory solution that can scale to billion-scale complex networks. Labelling-based methods are well-known for rendering fast response time to distance queries; however, existing works can only construct labelling on moderately large networks (million-scale) and cannot scale to large networks (billion-scale) due to their prohibitively large space requirements and very long preprocessing time. In this work, we present novel techniques to efficiently construct distance labelling and process exact shortest path distance queries for complex networks with billions of vertices and billions of edges. Our method is based on two ingredients: (i) a scalable labelling algorithm for constructing minimal distance labelling, and (ii) a querying framework that supports fast distance-bounded search on a sparsified graph. Thus, we first develop a novel labelling algorithm that can scale to graphs at the billion-scale. Then, we formalize a querying framework for exact distance queries, which combines our proposed highway cover distance labelling with distance-bounded searches to enable fast distance computation. To speed up the labelling construction process, we further propose a parallel labelling method that can construct labelling simultaneously for multiple landmarks. We evaluated the performance of the proposed methods on 12 real-world networks. The experiments show that the proposed methods can not only handle networks with billions of vertices, but also be up to 70 times faster in constructing labelling and save up to 90\% of labelling space. In particular, our method can answer distance queries on a billion-scale network of around 8B edges in less than 1ms, on average.

cs.DS

Multiplex Bipartite Network Embedding using Dual Hypergraph Convolutional Networks

A bipartite network is a graph structure where nodes are from two distinct domains and only inter-domain interactions exist as edges. A large number of network embedding methods exist to learn vectorial node representations from general graphs with both homogeneous and heterogeneous node and edge types, including some that can specifically model the distinct properties of bipartite networks. However, these methods are inadequate to model multiplex bipartite networks (e.g., in e-commerce), that have multiple types of interactions (e.g., click, inquiry, and buy) and node attributes. Most real-world multiplex bipartite networks are also sparse and have imbalanced node distributions that are challenging to model. In this paper, we develop an unsupervised Dual HyperGraph Convolutional Network (DualHGCN) model that scalably transforms the multiplex bipartite network into two sets of homogeneous hypergraphs and uses spectral hypergraph convolutional operators, along with intra- and inter-message passing strategies to promote information exchange within and across domains, to learn effective node embedding. We benchmark DualHGCN using four real-world datasets on link prediction and node classification tasks. Our extensive experiments demonstrate that DualHGCN significantly outperforms state-of-the-art methods, and is robust to varying sparsity levels and imbalanced node distributions.

cs.LG

A Gyrokinetic Simulation Model for Low Frequency Electromagnetic Fluctuations in Magnetized Plasmas

We present a new model for simulating the electromagnetic fluctuations with frequencies much lower than the ion cyclotron frequency in plasmas confined in general magnetic configurations. This novel model (termed as GK-E&B) employs nonlinear gyrokinetic equations formulated in terms of electromagnetic fields along with momentum balance equations for solving fields. It, thus, not only includes kinetic effects, such as wave-particle interaction and microscopic (ion Larmor radius scale) physics; but also is computationally more efficient than the conventional formulation described in terms of potentials. As a benchmark, we perform linear as well as nonlinear simulations of the kinetic Alfven wave; demonstrating physics in agreement with the analytical theories.

physics.plasm-ph

Asymptotics of the Charlier polynomials via difference equation methods

We derive uniform and non-uniform asymptotics of the Charlier polynomials by using difference equation methods alone. The Charlier polynomials are special in that they do not fit into the framework of the turning point theory, despite the fact that they are crucial in the Askey scheme. In this paper, asymptotic approximations are obtained respectively in the outside region, an intermediate region, and near the turning points. In particular, we obtain uniform asymptotic approximation at a pair of coalescing turning points with the aid of a local transformation. We also give a uniform approximation at the origin by applying the method of dominant balance and several matching techniques.

math.CA

Modeling Dynamic Heterogeneous Network for Link Prediction using Hierarchical Attention with Temporal RNN

Network embedding aims to learn low-dimensional representations of nodes while capturing structure information of networks. It has achieved great success on many tasks of network analysis such as link prediction and node classification. Most of existing network embedding algorithms focus on how to learn static homogeneous networks effectively. However, networks in the real world are more complex, e.g., networks may consist of several types of nodes and edges (called heterogeneous information) and may vary over time in terms of dynamic nodes and edges (called evolutionary patterns). Limited work has been done for network embedding of dynamic heterogeneous networks as it is challenging to learn both evolutionary and heterogeneous information simultaneously. In this paper, we propose a novel dynamic heterogeneous network embedding method, termed as DyHATR, which uses hierarchical attention to learn heterogeneous information and incorporates recurrent neural networks with temporal attention to capture evolutionary patterns. We benchmark our method on four real-world datasets for the task of link prediction. Experimental results show that DyHATR significantly outperforms several state-of-the-art baselines.

cs.SI