SearcharxivSearch

arXiv subjects

Weiqiang Lin

Publications and source records attributed to Weiqiang Lin.

6 recordsLinked to original sources

GenoBERT: A Language Model for Accurate Genotype Imputation

Genotype imputation enables dense variant coverage for genome-wide association and risk-prediction studies, yet conventional reference-panel methods remain limited by ancestry bias and reduced rare-variant accuracy. We present Genotype Bidirectional Encoder Representations from Transformers (GenoBERT), a transformer-based, reference-free framework that tokenizes phased genotypes and uses a self-attention mechanism to capture both short- and long-range linkage disequilibrium (LD) dependencies. Benchmarking on two independent datasets including the Louisiana Osteoporosis Study (LOS) and the 1000 Genomes Project (1KGP) across ancestry groups and multiple genotype missingness levels (5-50%) shows that GenoBERT achieves the highest overall accuracy compared to four baseline methods (Beagle5.4, SCDA, BiU-Net, and STICI). At practical sparsity levels (up to 25% missing), GenoBERT attains high overall imputation accuracy ($r^2 approx 0.98$) across datasets, and maintains robust performance ($r^2 > 0.90$) even at 50% missingness. Experimental results across different ancestries confirm consistent gains across datasets, with resilience to small sample sizes and weak LD. A 128-SNP (single-nucleotide polymorphism) context window (approximately 100 Kb) is validated through LD-decay analyses as sufficient to capture local correlation structures. By eliminating reference-panel dependence while preserving high accuracy, GenoBERT provides a scalable and robust solution for genotype imputation and a foundation for downstream genomic modeling.

q-bio.GN

A Non-Invasive Interpretable NAFLD Diagnostic Method Combining TCM Tongue Features

Non-alcoholic fatty liver disease (NAFLD) is a clinicopathological syndrome characterized by hepatic steatosis resulting from the exclusion of alcohol and other identifiable liver-damaging factors. It has emerged as a leading cause of chronic liver disease worldwide. Currently, the conventional methods for NAFLD detection are expensive and not suitable for users to perform daily diagnostics. To address this issue, this study proposes a non-invasive and interpretable NAFLD diagnostic method, the required user-provided indicators are only Gender, Age, Height, Weight, Waist Circumference, Hip Circumference, and tongue image. This method involves merging patients' physiological indicators with tongue features, which are then input into a fusion network named SelectorNet. SelectorNet combines attention mechanisms with feature selection mechanisms, enabling it to autonomously learn the ability to select important features. The experimental results show that the proposed method achieves an accuracy of 77.22\% using only non-invasive data, and it also provides compelling interpretability matrices. This study contributes to the early diagnosis of NAFLD and the intelligent advancement of TCM tongue diagnosis. The project mentioned in this paper is currently publicly available.

eess.IV

Social Media Brand Engagement as a Proxy for E-commerce Activities: A Case Study of Sina Weibo and JD

E-commerce platforms facilitate sales of products while product vendors engage in Social Media Activities (SMA) to drive E-commerce Platform Activities (EPA) of consumers, enticing them to search, browse and buy products. The frequency and timing of SMA are expected to affect levels of EPA, increasing the number of brand related queries, clickthrough, and purchase orders. This paper applies cross-sectional data analysis to explore such beliefs and demonstrates weak-to-moderate correlations between daily SMA and EPA volumes. Further correlation analysis, using 30-day rolling windows, shows a high variability in correlation of SMA-EPA pairs and calls into question the predictive potential of SMA in relation to EPA. Considering the moderate correlation of selected SMA and EPA pairs (e.g., Post-Orders), we investigate whether SMA features can predict changes in the EPA levels, instead of precise EPA daily volumes. We define such levels in terms of EPA distribution quantiles (2, 3, and 5 levels) over training data. We formulate the EPA quantile predictions as a multi-class categorization problem. The experiments with Random Forest and Logistic Regression show a varied success, performing better than random for the top quantiles of purchase orders and for the lowest quantile of search and clickthrough activities. Similar results are obtained when predicting multi-day cumulative EPA levels (1, 3, and 7 days). Our results have considerable practical implications but, most importantly, urge the common beliefs to be re-examined, seeking a stronger evidence of SMA effects on EPA.

cs.SI

Graded modules for Virasoro-like algebra

In this paper, we consider the classification of irreducible ${\bf Z}$- and ${\bf Z}^2$-graded modules with finite dimensional homogeneous subspaces over the Virasoro-like algebra. We first prove that such a module is a uniformly bounded module or a generalized highest weight module. Then we determine all generalized highest weight irreducible modules. As a consequence, we also determine all the modules with nonzero center. Finally, we prove that there does not exist any nontrivial ${\bf Z}$-graded modules of intermediate series.

math.RT

The classification of $\bf Z$-graded modules of the intermediate series over the $q$-analog Virasoro-like algebra

In this paper, we complete the classification of the {\bf Z}-graded modules of the intermediate series over the $q$-analog Virasoro-like algebra $L$. We first construct four classes of irreducible {\bf Z}-graded $L$-modules of the intermediate series. Then we prove that any {\bf Z}-graded $L$-modules of the intermediate series must be the direct sum of some trivial $L$-modules or one of the modules constructed by us.

math.RT

Classification of quasifinite representations with nonzero central charges for type $A_1$ EALA with coordinates in quantum torus

In this paper, we first construct a Lie algebra $L$ from rank 3 quantum torus, and show that it is isomorphic to the core of EALAs of type $A_1$ with coordinates in rank 2 quantum torus. Then we construct two classes of irreducible ${\bf Z}$-graded highest weight representations, and give the necessary and sufficient conditions for these representations to be quasifinite. Next, we prove that they exhaust all the generalized highest weight irreducible ${\bf Z}$-graded quasifinite representations. As a consequence, we determine all the irreducible ${\bf Z}$-graded quasifinite representations with nonzero central charges. Finally, we construct two classes of highest weight ${\bf Z}^2$-graded quasifinite representations by using these ${\bf Z}$-graded modules.

math.QA