Searcharxiv⌕ Search

arXiv subjects

Feng Shi

Publications and source records attributed to Feng Shi.

At least 37 records · Page 2Linked to original sources

A 2D Semantic-Aware Position Encoding for Vision Transformers

Vision transformers have demonstrated significant advantages in computer vision tasks due to their ability to capture long-range dependencies and contextual relationships through self-attention. However, existing position encoding techniques, which are largely borrowed from natural language processing, fail to effectively capture semantic-aware positional relationships between image patches. Traditional approaches like absolute position encoding and relative position encoding primarily focus on 1D linear position relationship, often neglecting the semantic similarity between distant yet contextually related patches. These limitations hinder model generalization, translation equivariance, and the ability to effectively handle repetitive or structured patterns in images. In this paper, we propose 2-Dimensional Semantic-Aware Position Encoding ($\text{SaPE}^2$), a novel position encoding method with semantic awareness that dynamically adapts position representations by leveraging local content instead of fixed linear position relationship or spatial coordinates. Our method enhances the model's ability to generalize across varying image resolutions and scales, improves translation equivariance, and better aggregates features for visually similar but spatially distant patches. By integrating $\text{SaPE}^2$ into vision transformers, we bridge the gap between position encoding and perceptual similarity, thereby improving performance on computer vision tasks.

cs.CV↗

LipidBERT: A Lipid Language Model Pre-trained on METiS de novo Lipid Library

In this study, we generate and maintain a database of 10 million virtual lipids through METiS's in-house de novo lipid generation algorithms and lipid virtual screening techniques. These virtual lipids serve as a corpus for pre-training, lipid representation learning, and downstream task knowledge transfer, culminating in state-of-the-art LNP property prediction performance. We propose LipidBERT, a BERT-like model pre-trained with the Masked Language Model (MLM) and various secondary tasks. Additionally, we compare the performance of embeddings generated by LipidBERT and PhatGPT, our GPT-like lipid generation model, on downstream tasks. The proposed bilingual LipidBERT model operates in two languages: the language of ionizable lipid pre-training, using in-house dry-lab lipid structures, and the language of LNP fine-tuning, utilizing in-house LNP wet-lab data. This dual capability positions LipidBERT as a key AI-based filter for future screening tasks, including new versions of METiS de novo lipid libraries and, more importantly, candidates for in vivo testing for orgran-targeting LNPs. To the best of our knowledge, this is the first successful demonstration of the capability of a pre-trained language model on virtual lipids and its effectiveness in downstream tasks using web-lab data. This work showcases the clever utilization of METiS's in-house de novo lipid library as well as the power of dry-wet lab integration.

cs.CL↗

Runtime Performance of Evolutionary Algorithms for the Chance-constrained Makespan Scheduling Problem

The Makespan Scheduling problem is an extensively studied NP-hard problem, and its simplest version looks for an allocation approach for a set of jobs with deterministic processing times to two identical machines such that the makespan is minimized. However, in real life scenarios, the actual processing time of each job may be stochastic around the expected value with a variance, under the influence of external factors, and the actual processing times of these jobs may be correlated with covariances. Thus within this paper, we propose a chance-constrained version of the Makespan Scheduling problem and investigate the theoretical performance of the classical Randomized Local Search and (1+1) EA for it. More specifically, we first study two variants of the Chance-constrained Makespan Scheduling problem and their computational complexities, then separately analyze the expected runtime of the two algorithms to obtain an optimal solution or almost optimal solution to the instances of the two variants. In addition, we investigate the experimental performance of the two algorithms for the two variants.

cs.NE↗

Galaxy and halo properties around cosmic filaments from Sloan Digital Sky Survey Data Release 7 and the ELUCID simulation

Using galaxies from the Sloan Digital Sky Survey Data Release 7 (SDSS DR7) along with haloes from the dark matter only constrained ELUCID (Exploring the Local Universe with the reConstructed Initial Density field) simulation, we examine the properties of galaxies and haloes with respect to their distance to cosmic filaments, determined by the medial-axis thinning technique of the COsmic Web Skeleton (COWS) method. Our findings suggest that galaxies or subhaloes grow in mass as they approach these filaments. Galaxies exhibit a redder colour and diminished specific star formation rates as they approach these filaments. Additionally, older subhaloes tend to be more common near the central regions of these filaments. Elliptical galaxies are more frequently found than spiral galaxies in the central regions of the filaments. Lower-mass galaxies typically display reduced sizes in proximity to filaments, whereas higher-mass galaxies tend to exhibit increased sizes when close to filaments. Moreover, the concentration and spin of the haloes grow as they approach the filaments. These findings support the notion that the large-scale structure of the universe, characterized by cosmic web structures, plays a vital role in shaping galaxy and halo properties.

astro-ph.CO↗

Populating Large N-body Simulations with LRGs Using Neural Networks

The analysis of state-of-the-art cosmological surveys like the Dark Energy Spectroscopic Instrument (DESI) survey requires high-resolution, large-volume simulations. However, the computational cost of hydrodynamical simulations at these scales is prohibitive. Instead, dark matter (DM)-only simulations are used, with galaxies populated a posteriori, typically via halo occupation distribution (HOD) models. While effective, HOD models are statistical in nature and lack full physical motivation. In this work, we explore using neural networks (NNs) to learn the complex, physically motivated relationships between DM haloes and galaxy properties. Trained on small-volume, high-resolution hydrodynamical simulations, our NN predicts galaxy properties in a larger DM-only simulation and determines which galaxies should be classified as luminous red galaxies (LRGs). Comparing the original LRG sample to the one generated by our NN, we find that, while the subhalo mass distributions are similar, our NN selects fewer low-mass subhaloes as LRG hosts, possibly due to the absence of baryonic feedback effects in DM-only simulations. This feedback could brighten or redden galaxies, altering their classification. Finally, we generate a new LRG sample by fitting an HOD model to the NN-generated LRG sample. We verify that both the HOD- and NN-generated samples preserve a set of bias parameter relations, which assume that the higher-order parameters, $b_{s2}$ and $b_{3\rm{nl}}$, are determined by the linear bias parameter $b_{1}$. These relations are commonly used to simplify clustering analyses.

astro-ph.CO↗

Cosmological distance forecasts for the CSST Galaxy Survey using BAO peaks

The measurement of cosmological distances using baryon acoustic oscillations (BAO) is crucial for studying the universe's expansion. The Chinese Space Station Telescope (CSST) galaxy redshift survey, with its vast volume and sky coverage, provides an opportunity to address key challenges in cosmology. However, redshift uncertainties in galaxy surveys can degrade both angular and radial distance estimates. In this study, we forecast the precision of BAO distance measurements using mock CSST galaxy samples, applying a two-point correlation function (2PCF) wedge approach to mitigate redshift errors. We simulate redshift uncertainties of $σ_0 = 0.003$ and $σ_0 = 0.006$, representative of expected CSST errors, and examine their effects on the BAO peak and distance scaling factors, $α_\perp$ and $α_\parallel$, across redshift bins within $0.0 < z \leqslant 1.0$. The wedge 2PCF method proves more effective in detecting the BAO peak compared to the monopole 2PCF, particularly for $σ_0 = 0.006$. Constraints on the BAO peaks show that $α_\perp$ is well constrained around 1.0, regardless of $σ_0$, with precision between 1% and 3% across redshift bins. In contrast, $α_\parallel$ measurements are more sensitive to increases in $σ_0$. For $σ_0 = 0.003$, the results remain close to the fiducial value, with uncertainties ranging between 4% and 9%; for $σ_0 = 0.006$, significant deviations from the fiducial value are observed. We also study the ability to measure parameters $(Ω_m, H_0r_\mathrm{d})$ using distance measurements, proving robust constraints as a cosmological probe under CSST-like redshift uncertainties.

astro-ph.CO↗

ChatZero:Zero-shot Cross-Lingual Dialogue Generation via Pseudo-Target Language

Although large language models(LLMs) show amazing capabilities, among various exciting applications discovered for LLMs fall short in other low-resource languages. Besides, most existing methods depend on large-scale dialogue corpora and thus building systems for dialogue generation in a zero-shot scenario remains a considerable challenge. To address this challenge, we propose a novel end-to-end zero-shot dialogue generation model ChatZero based on cross-lingual code-switching method. First, we construct code-switching language and pseudo-target language with placeholders. Then for cross-lingual semantic transfer, we employ unsupervised contrastive learning to minimize the semantics gap of the source language, code-switching language, and pseudo-target language that are mutually positive examples in the high dimensional semantic space. Experiments on the multilingual DailyDialog and DSTC7-AVSD datasets demonstrate that ChatZero can achieve more than 90\% of the original performance under the zero-shot case compared to supervised learning, and achieve state-of-the-art performance compared with other baselines.

cs.CL↗

Knowledge-Guided Prompt Learning for Lifespan Brain MR Image Segmentation

Automatic and accurate segmentation of brain MR images throughout the human lifespan into tissue and structure is crucial for understanding brain development and diagnosing diseases. However, challenges arise from the intricate variations in brain appearance due to rapid early brain development, aging, and disorders, compounded by the limited availability of manually-labeled datasets. In response, we present a two-step segmentation framework employing Knowledge-Guided Prompt Learning (KGPL) for brain MRI. Specifically, we first pre-train segmentation models on large-scale datasets with sub-optimal labels, followed by the incorporation of knowledge-driven embeddings learned from image-text alignment into the models. The introduction of knowledge-wise prompts captures semantic relationships between anatomical variability and biological processes, enabling models to learn structural feature embeddings across diverse age groups. Experimental findings demonstrate the superiority and robustness of our proposed method, particularly noticeable when employing Swin UNETR as the backbone. Our approach achieves average DSC values of 95.17% and 94.19% for brain tissue and structure segmentation, respectively. Our code is available at https://github.com/TL9792/KGPL.

eess.IV↗

Hunting imaging biomarkers in pulmonary fibrosis: Benchmarks of the AIIB23 challenge

Airway-related quantitative imaging biomarkers are crucial for examination, diagnosis, and prognosis in pulmonary diseases. However, the manual delineation of airway trees remains prohibitively time-consuming. While significant efforts have been made towards enhancing airway modelling, current public-available datasets concentrate on lung diseases with moderate morphological variations. The intricate honeycombing patterns present in the lung tissues of fibrotic lung disease patients exacerbate the challenges, often leading to various prediction errors. To address this issue, the 'Airway-Informed Quantitative CT Imaging Biomarker for Fibrotic Lung Disease 2023' (AIIB23) competition was organized in conjunction with the official 2023 International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI). The airway structures were meticulously annotated by three experienced radiologists. Competitors were encouraged to develop automatic airway segmentation models with high robustness and generalization abilities, followed by exploring the most correlated QIB of mortality prediction. A training set of 120 high-resolution computerised tomography (HRCT) scans were publicly released with expert annotations and mortality status. The online validation set incorporated 52 HRCT scans from patients with fibrotic lung disease and the offline test set included 140 cases from fibrosis and COVID-19 patients. The results have shown that the capacity of extracting airway trees from patients with fibrotic lung disease could be enhanced by introducing voxel-wise weighted general union loss and continuity loss. In addition to the competitive image biomarkers for prognosis, a strong airway-derived biomarker (Hazard ratio>1.5, p<0.0001) was revealed for survival prognostication compared with existing clinical measurements, clinician assessment and AI-based biomarkers.

eess.IV↗

CSST large-scale structure analysis pipeline: I. constructing reference mock galaxy redshift surveys

In this paper, we set out to construct a set of reference mock galaxy redshift surveys (MGRSs) for the future Chinese Space-station Survey Telescope (CSST) observation, where subsequent survey selection effects can be added and evaluated. This set of MGRSs is generated using the dark matter subhalos extracted from a high-resolution Jiutian $N$-body simulation of the standard $Λ$CDM cosmogony with $Ω_m=0.3111$, $Ω_Λ=0.6889$, and $σ_8=0.8102$. The simulation has a boxsize of $1~h^{-1} {\rm Gpc}$, and consists of $6144^3$ particles with mass resolution $3.723 \times 10^{8} h^{-1} M_\odot $. In order to take into account the effect of redshift evolution, we first use all 128 snapshots in the Jiutian simulation to generate a light-cone halo/subhalo catalog. Next, galaxy luminosities are assigned to the main and subhalo populations using the subhalo abundance matching (SHAM) method with the DESI $z$-band luminosity functions at different redshifts. Multi-band photometries, as well as images, are then assigned to each mock galaxy using a 3-dimensional parameter space nearest neighbor sampling of the DESI LS observational galaxies and groups. Finally, the CSST and DESI LS survey geometry and magnitude limit cuts are applied to generate the required MGRSs. As we have checked, this set of MGRSs can generally reproduce the observed galaxy luminosity/mass functions within 0.1 dex for galaxies with $L > 10^8 L_\odot$ (or $M_* > 10^{8.5} M_\odot$) and within 1-$σ$ level for galaxies with $L < 10^8L_\odot$ (or $M_* < 10^{8.5} M_\odot$). Together with the CSST slitless spectra and redshifts for our DESI LS seed galaxies that are under construction, we will set out to test various slitless observational selection effects in subsequent probes.

astro-ph.GA↗

21-cm foreground removal using AI and frequency-difference technique

The deep learning technique has been employed in removing foreground contaminants from 21 cm intensity mapping, but its effectiveness is limited by the large dynamic range of the foreground amplitude. In this study, we develop a novel foreground removal technique grounded in U-Net networks. The essence of this technique lies in introducing an innovative data preprocessing step specifically, utilizing the temperature difference between neighboring frequency bands as input, which can substantially reduce the dynamic range of foreground amplitudes by approximately two orders of magnitude. This reduction proves to be highly advantageous for the U-Net foreground removal. We observe that the HI signal can be reliably recovered, as indicated by the cross-correlation power spectra showing unity agreement at the scale of $k < 0.3 h^{-1}$Mpc in the absence of instrumental effects. Moreover, accounting for the systematic beam effects, our reconstruction displays consistent auto-correlation and cross-correlation power spectrum ratios at the $1σ$ level across scales $k \lesssim 0.1 h^{-1}$Mpc, with only a 10% reduction observed in the cross-correlation power spectrum at $k\simeq0.2 h^{-1}$Mpc. The effects of redshift-space distortion are also reconstructed successfully, as evidenced by the quadrupole power spectra matching. In comparison, our method outperforms the traditional Principal Component Analysis method, which derived cross-correlation ratios are underestimated by around 60%. We simulated various white noise levels in the map and found that the mean cross-correlation ratio $\bar{R}_\mathrm{cross} \gtrsim 0.8$ when the level of the thermal noise is smaller than or equal to that of the HI signal. We conclude that the proposed frequency-difference technique can significantly enhance network performance by reducing the amplitude range of foregrounds and aiding in the prevention of HI loss.

astro-ph.CO↗

Single Cells Are Spatial Tokens: Transformers for Spatial Transcriptomic Data Imputation

Spatially resolved transcriptomics brings exciting breakthroughs to single-cell analysis by providing physical locations along with gene expression. However, as a cost of the extremely high spatial resolution, the cellular level spatial transcriptomic data suffer significantly from missing values. While a standard solution is to perform imputation on the missing values, most existing methods either overlook spatial information or only incorporate localized spatial context without the ability to capture long-range spatial information. Using multi-head self-attention mechanisms and positional encoding, transformer models can readily grasp the relationship between tokens and encode location information. In this paper, by treating single cells as spatial tokens, we study how to leverage transformers to facilitate spatial tanscriptomics imputation. In particular, investigate the following two key questions: (1) $\textit{how to encode spatial information of cells in transformers}$, and (2) $\textit{ how to train a transformer for transcriptomic imputation}$. By answering these two questions, we present a transformer-based imputation framework, SpaFormer, for cellular-level spatial transcriptomic data. Extensive experiments demonstrate that SpaFormer outperforms existing state-of-the-art imputation algorithms on three large-scale datasets while maintaining superior computational efficiency.

q-bio.GN↗

(DarkAI) Mapping the large-scale density field of dark matter using artificial intelligence

Herein, we present a deep-learning technique for reconstructing the dark-matter density field from the redshift-space distribution of dark-matter halos. We built a UNet-architecture neural network and trained it using the COmoving Lagrangian Acceleration fast simulation, which is an approximation of the N-body simulation with $512^3$ particles in a box size of 500 Mpc $h^{-1}$. Further, we tested the resulting UNet model not only with training-like test samples but also with standard N-body simulations, such as the Jiutian simulation with $6144^3$ particles in a box size of 1000 Mpc $h^{-1}$ and the ELUCID simulation, which has a different cosmology. The real-space dark-matter density fields in the three simulations can be reconstructed reliably with only a small reduction of the cross-correlation power spectrum at 1% and 10% levels at $k=0.1$ and $0.3~h\mathrm{Mpc^{-1}}$, respectively. The reconstruction clearly helps to correct for redshift-space distortions and is unaffected by the different cosmologies between the training (Planck2018) and test samples (WMAP5). Furthermore, we tested the application of the UNet-reconstructed density field to obtain the velocity \& tidal field and found that this approach provides better results compared to the traditional approach based on the linear bias model, showing a 12.2% improvement in the correlation slope and a 21.1% reduction in the scatter between the predicted and true velocities. Thus, our method is highly efficient and has excellent extrapolation reliability beyond the training set. This provides an ideal solution for determining the three-dimensional underlying density field from the plentiful galaxy survey data.

astro-ph.CO↗

Optimizing Chance-Constrained Submodular Problems with Variable Uncertainties

Chance constraints are frequently used to limit the probability of constraint violations in real-world optimization problems where the constraints involve stochastic components. We study chance-constrained submodular optimization problems, which capture a wide range of optimization problems with stochastic constraints. Previous studies considered submodular problems with stochastic knapsack constraints in the case where uncertainties are the same for each item that can be selected. However, uncertainty levels are usually variable with respect to the different stochastic components in real-world scenarios, and rigorous analysis for this setting is missing in the context of submodular optimization. This paper provides the first such analysis for this case, where the weights of items have the same expectation but different dispersion. We present greedy algorithms that can obtain a high-quality solution, i.e., a constant approximation ratio to the given optimal solution from the deterministic setting. In the experiments, we demonstrate that the algorithms perform effectively on several chance-constrained instances of the maximum coverage problem and the influence maximization problem.

math.OC↗

Fragment and Integrate Network (FIN): A Novel Spatial-Temporal Modeling Based on Long Sequential Behavior for Online Food Ordering Click-Through Rate Prediction

Spatial-temporal information has been proven to be of great significance for click-through rate prediction tasks in online Location-Based Services (LBS), especially in mainstream food ordering platforms such as DoorDash, Uber Eats, Meituan, and Ele.me. Modeling user spatial-temporal preferences with sequential behavior data has become a hot topic in recommendation systems and online advertising. However, most of existing methods either lack the representation of rich spatial-temporal information or only handle user behaviors with limited length, e.g. 100. In this paper, we tackle these problems by designing a new spatial-temporal modeling paradigm named Fragment and Integrate Network (FIN). FIN consists of two networks: (i) Fragment Network (FN) extracts Multiple Sub-Sequences (MSS) from lifelong sequential behavior data, and captures the specific spatial-temporal representation by modeling each MSS respectively. Here both a simplified attention and a complicated attention are adopted to balance the performance gain and resource consumption. (ii) Integrate Network (IN) builds a new integrated sequence by utilizing spatial-temporal interaction on MSS and captures the comprehensive spatial-temporal representation by modeling the integrated sequence with a complicated attention. Both public datasets and production datasets have demonstrated the accuracy and scalability of FIN. Since 2022, FIN has been fully deployed in the recommendation advertising system of Ele.me, one of the most popular online food ordering platforms in China, obtaining 5.7% improvement on Click-Through Rate (CTR) and 7.3% increase on Revenue Per Mille (RPM).

cs.IR↗

Alternately Optimized Graph Neural Networks

Graph Neural Networks (GNNs) have greatly advanced the semi-supervised node classification task on graphs. The majority of existing GNNs are trained in an end-to-end manner that can be viewed as tackling a bi-level optimization problem. This process is often inefficient in computation and memory usage. In this work, we propose a new optimization framework for semi-supervised learning on graphs. The proposed framework can be conveniently solved by the alternating optimization algorithms, resulting in significantly improved efficiency. Extensive experiments demonstrate that the proposed method can achieve comparable or better performance with state-of-the-art baselines while it has significantly better computation and memory efficiency.

cs.LG↗

Towards Label Position Bias in Graph Neural Networks

Graph Neural Networks (GNNs) have emerged as a powerful tool for semi-supervised node classification tasks. However, recent studies have revealed various biases in GNNs stemming from both node features and graph topology. In this work, we uncover a new bias - label position bias, which indicates that the node closer to the labeled nodes tends to perform better. We introduce a new metric, the Label Proximity Score, to quantify this bias, and find that it is closely related to performance disparities. To address the label position bias, we propose a novel optimization framework for learning a label position unbiased graph structure, which can be applied to existing GNNs. Extensive experiments demonstrate that our proposed method not only outperforms backbone methods but also significantly mitigates the issue of label position bias in GNNs.

cs.LG↗

Mining fMRI Dynamics with Parcellation Prior for Brain Disease Diagnosis

To characterize atypical brain dynamics under diseases, prevalent studies investigate functional magnetic resonance imaging (fMRI). However, most of the existing analyses compress rich spatial-temporal information as the brain functional networks (BFNs) and directly investigate the whole-brain network without neurological priors about functional subnetworks. We thus propose a novel graph learning framework to mine fMRI signals with topological priors from brain parcellation for disease diagnosis. Specifically, we 1) detect diagnosis-related temporal features using a "Transformer" for a higher-level BFN construction, and process it with a following graph convolutional network, and 2) apply an attention-based multiple instance learning strategy to emphasize the disease-affected subnetworks to further enhance the diagnosis performance and interpretability. Experiments demonstrate higher effectiveness of our method than compared methods in the diagnosis of early mild cognitive impairment. More importantly, our method is capable of localizing crucial brain subnetworks during the diagnosis, providing insights into the pathogenic source of mild cognitive impairment.

eess.IV↗