SearcharxivSearch

arXiv subjects

Zicheng Xie

Publications and source records attributed to Zicheng Xie.

3 recordsLinked to original sources

An LLM-Powered Semantic Alignment Framework for Journal Recommendation

Journal recommendation is an important task in scholarly information systems. Existing approaches typically rely on supervised learning models, manually engineered features, or historical interaction data, which may limit their generalizability and interpretability. We propose an LLM-powered semantic alignment framework that formulates journal recommendation as a semantic matching problem between manuscript content and journal scope descriptions. The framework enables large language models (LLMs) to infer journal suitability directly from article titles, abstracts, keywords, and candidate journal information without task-specific training. Experiments are conducted using DeepSeek-V3 on a dataset of 23,609 articles from 49 journals in statistics and related fields. The proposed framework achieves Top-3, Top-5, and Top-10 accuracies of 40.23\%, 53.67\%, and 70.05\%, respectively. Additional analyses show that incorporating reference information generally improves recommendation performance and that recommendations remain highly stable across repeated runs, with an average Top-5 Jaccard similarity of 84\%. The framework also generates interpretable reasoning outputs that provide insights into the recommendation process. These findings demonstrate the potential of LLMs as a training-free and scalable paradigm for journal recommendation and scholarly decision support.

cs.IR

Learning Unified Representations from Heterogeneous Data for Robust Heart Rate Modeling

Heart rate prediction is vital for personalized health monitoring and fitness, while it frequently faces a critical challenge in real-world deployment: data heterogeneity. We classify it in two key dimensions: source heterogeneity from fragmented device markets with varying feature sets, and user heterogeneity reflecting distinct physiological patterns across individuals and activities. Existing methods either discard device-specific information, or fail to model user-specific differences, limiting their real-world performance. To address this, we propose a framework that learns latent representations agnostic to both heterogeneity,enabling downstream predictors to work consistently under heterogeneous data patterns. Specifically, we introduce a random feature dropout strategy to handle source heterogeneity, making the model robust to various feature sets. To manage user heterogeneity, we employ a history-aware attention module to capture long-term physiological traits and use a contrastive learning objective to build a discriminative representation space. To reflect the heterogeneous nature of real-world data, we created a new benchmark dataset, PARROTAO. Evaluations on both PARROTAO and the public FitRec dataset show that our model significantly outperforms existing baselines by 17.5% and 10.4% in terms of test MSE, respectively. Furthermore, analysis of the learned representations demonstrates their strong discriminative power,and two downstream application tasks confirm the practical value of our model.

cs.LG

Bi-SCORE for Weighted Bipartite Networks with Application in Knowledge Source Discovery

Community detection in citation networks offers a powerful approach to understanding knowledge flow and identifying core research areas within academic disciplines. This study focuses on knowledge source discovery in statistics by analyzing a weighted bipartite journal citation network constructed from 16,119 articles published in eight core journals from 2001 to 2023. To capture the inherent asymmetry of citation behavior, we explicitly preserve the bipartite structure of the network, distinguishing between citing and cited journals. For this task, we propose Bi-SCORE (Bipartite Spectral Clustering on Ratios-of-Eigenvectors), a computationally efficient and initialization-free spectral method designed for community detection in weighted bipartite networks with degree heterogeneity. We establish rigorous theoretical guarantees for the performance of Bi-SCORE under the weighted bipartite degree-corrected stochastic block model. Furthermore, simulation studies demonstrate its robustness across varying levels of sparsity and degree heterogeneity, where it outperforms existing methods. When applied to the real-world citation network, Bi-SCORE uncovers a six-community structure corresponding to key research areas in statistics, including applied statistics, methodology, theory, computation, and econometrics. These findings provide valuable insights into the intricate citation patterns and knowledge flow among statistical journals.

stat.ME