SearcharxivSearch

arXiv subjects

Xiaofei Chen

Publications and source records attributed to Xiaofei Chen.

8 recordsLinked to original sources

Transformer-based prediction of two-dimensional material electronic properties under elastic strain engineering

Strain engineering provides a powerful route for tuning the electronic properties of two-dimensional (2D) materials, but exploring the full multidimensional strain space with density functional theory (DFT) is computationally prohibitive due to the nonlinear coupling between normal and shear components. In this work, we introduce a Transformer-based, multi-target surrogate model framework that achieves DFT-level bandgap prediction accuracy, reaching a mean absolute error of 0.0103 eV while retaining full interpretability through attention-weight analysis. The learned self-attention map consistently identifies shear strain as the interaction center that influences both bandgap and phonon stability, an insight not readily captured by classical feature-importance metrics. This work establishes attention-based architectures as physically interpretable surrogate models for multi-property prediction, offering a generalizable strategy for accelerating deep elastic strain engineering in materials informatics.

cond-mat.mtrl-sci

Evaluation of Minimal Residual Disease as a Surrogate for Progression-Free Survival in Hematology Oncology Trials: A Meta-Analytic Review

Traditional health authority approval for oncology drugs is based on a clinical benefit endpoint, or a valid surrogate. In 1992 the FDA created the Accelerated Approval pathway to allow for earlier approval of therapies in serious conditions with an unmet medical need. This is accomplished typically by granting accelerated approval based on a surrogate endpoint that can be measured earlier than a traditional approval endpoint. Minimal residual disease (MRD) is a sensitive measure of residual cancer cells in hematology oncology after treatment, and is increasingly considered as a secondary or exploratory endpoint due to its prognostic potential for traditional clinical trial endpoints such as progression-free survival (PFS) and overall survival (OS). This work aims to evaluate MRD's surrogacy potential across several hematologic cancer indications while keeping the focus on follicular lymphoma (FL), using data from published studies. We examine individual-level and trial-level correlations extracted from previously published studies to elucidate the potential role of MRD in accelerating the drug approval process in hematology oncology trials.

stat.AP

Trade-offs in Image Generation: How Do Different Dimensions Interact?

Model performance in text-to-image (T2I) and image-to-image (I2I) generation often depends on multiple aspects, including quality, alignment, diversity, and robustness. However, models' complex trade-offs among these dimensions have rarely been explored due to (1) the lack of datasets that allow fine-grained quantification of these trade-offs, and (2) the use of a single metric for multiple dimensions. To bridge this gap, we introduce TRIG-Bench (Trade-offs in Image Generation), which spans 10 dimensions (Realism, Originality, Aesthetics, Content, Relation, Style, Knowledge, Ambiguity, Toxicity, and Bias), contains 40,200 samples, and covers 132 pairwise dimensional subsets. Furthermore, we develop TRIGScore, a VLM-as-judge metric that automatically adapts to various dimensions. Based on TRIG-Bench and TRIGScore, we evaluate 14 models across T2I and I2I tasks. In addition, we propose the Relation Recognition System to generate the Dimension Trade-off Map (DTM) that visualizes the trade-offs among model-specific capabilities. Our experiments demonstrate that DTM consistently provides a comprehensive understanding of the trade-offs between dimensions for each type of generative model. Notably, we show that the model's dimension-specific weaknesses can be mitigated through fine-tuning on DTM to enhance overall performance. Code is available at: https://github.com/fesvhtr/TRIG

cs.CV

Self-Arresting and Runaway Earthquakes:Nucleation, Propagation, Gutenberg-Richter law and Dragon-King Events

We develop a dissipation-based framework for earthquake rupture on homogeneous faults that explicitly separates the onset of unstable slip from the conditions required for self-sustained rupture propagation. This distinction explains the coexistence of self-arresting earthquakes and run-away ruptures (subshear and supershear events) observed in numerical simulations and empirical studies. We identify two distinct characteristic fault sizes: a nucleation radius controlling the instability of slip, and in general a larger propagation radius controlling whether an unstable rupture can be energetically sustained. Ruptures initiated above the nucleation scale but below the propagation scale spontaneously arrest. We further derive the Gutenberg-Richter law for self-arresting earthquakes by linking rupture physics to the fractal geometry of faulting. Finally, we interpret run-away ruptures as extreme events generated by an amplifying mechanism, consistent with the dragon-king concept. These results provide a unified physical basis for earthquake initiation, arrest, and seismicity statistics.

physics.geo-ph

Analyses and Concerns in Precision Medicine: A Statistical Perspective

This article explores the critical role of statistical analysis in precision medicine. It discusses how personalized healthcare is enhanced by statistical methods that interpret complex, multidimensional datasets, focusing on predictive modeling, machine learning algorithms, and data visualization techniques. The paper addresses challenges in data integration and interpretation, particularly with diverse data sources like electronic health records (EHRs) and genomic data. It also delves into ethical considerations such as patient privacy and data security. In addition, the paper highlights the evolution of statistical analysis in medicine, core statistical methodologies in precision medicine, and future directions in the field, emphasizing the integration of artificial intelligence (AI) and machine learning (ML).

cs.LG

Knowledge Boosting: Rethinking Medical Contrastive Vision-Language Pre-Training

The foundation models based on pre-training technology have significantly advanced artificial intelligence from theoretical to practical applications. These models have facilitated the feasibility of computer-aided diagnosis for widespread use. Medical contrastive vision-language pre-training, which does not require human annotations, is an effective approach for guiding representation learning using description information in diagnostic reports. However, the effectiveness of pre-training is limited by the large-scale semantic overlap and shifting problems in medical field. To address these issues, we propose the Knowledge-Boosting Contrastive Vision-Language Pre-training framework (KoBo), which integrates clinical knowledge into the learning of vision-language semantic consistency. The framework uses an unbiased, open-set sample-wise knowledge representation to measure negative sample noise and supplement the correspondence between vision-language mutual information and clinical knowledge. Extensive experiments validate the effect of our framework on eight tasks including classification, segmentation, retrieval, and semantic relatedness, achieving comparable or better performance with the zero-shot or few-shot settings. Our code is open on https://github.com/ChenXiaoFei-CS/KoBo.

cs.CV

PolarDB-IMCI: A Cloud-Native HTAP Database System at Alibaba

Cloud-native databases have become the de-facto choice for mission-critical applications on the cloud due to the need for high availability, resource elasticity, and cost efficiency. Meanwhile, driven by the increasing connectivity between data generation and analysis, users prefer a single database to efficiently process both OLTP and OLAP workloads, which enhances data freshness and reduces the complexity of data synchronization and the overall business cost. In this paper, we summarize five crucial design goals for a cloud-native HTAP database based on our experience and customers' feedback, i.e., transparency, competitive OLAP performance, minimal perturbation on OLTP workloads, high data freshness, and excellent resource elasticity. As our solution to realize these goals, we present PolarDB-IMCI, a cloud-native HTAP database system designed and deployed at Alibaba Cloud. Our evaluation results show that PolarDB-IMCI is able to handle HTAP efficiently on both experimental and production workloads; notably, it speeds up analytical queries up to $\times149$ on TPC-H (100 $GB$). PolarDB-IMCI introduces low visibility delay and little performance perturbation on OLTP workloads (< 5%), and resource elasticity can be achieved by scaling out in tens of seconds.

cs.DB

A Vector Autoregression Prediction Model for COVID-19 Outbreak

Since two people came down a county of north Seattle with positive COVID-19 (coronavirus-19) in 2019, the current total cases in the United States (U.S.) are over 12 million. Predicting the pandemic trend under effective variables is crucial to help find a way to control the epidemic. Based on available literature, we propose a validated Vector Autoregression (VAR) time series model to predict the positive COVID-19 cases. A real data prediction for U.S. is provided based on the U.S. coronavirus data. The key message from our study is that the situation of the pandemic will getting worse if there is no effective control.

stat.AP