SearcharxivSearch

arXiv subjects

Fucheng Zhong

Publications and source records attributed to Fucheng Zhong.

17 recordsLinked to original sources

Identification of gravitational lenses obscured by foreground light in the KiDS dataset using U-Nets and ResNets

*Context.* Many lensing images are often obscured by foreground light from the central galaxies, making them challenging to detect. *Aims.* To address the limitations of previous lens search efforts, particularly for samples with smaller $R_E$ or faint lensed images, we developed a composite convolutional neural network framework that utilizes both U-Net and ResNet architectures for feature extraction and classification. *Methods.* We propose a hybrid search method that combines U-Net and ResNet architectures to enhance the detection of foreground galaxy-obscured lenses. Our approach consists of two main stages: first, the U-Net model separates the foreground galaxy light from potential lensing signals, creating residual images that highlight the lensing features. Next, the ResNet module performs binary classification on these residual images to detect lensing signals. *Results.* We evaluated the hybrid search method with real observational data to demonstrate its effectiveness, achieving a recall of 71.5% and a 4.5% false positive rate at a confidence threshold of 0.6. Applying this method to over 638,398 galaxy samples from the Kilo-Degree Survey Data Release 4 and conducting thorough inspections, we identify 88 Class A, 322 Class B, and 1,758 Class C candidates. *Conclusions.* This hybrid approach significantly enhances the completeness of existing strong gravitational lensing searches and shows great potential for improving future astronomical surveys.

astro-ph.IM

Morphology classification for galaxies in the Kilo Degree Survey using a label-efficient self-supervised learning framework

Galaxy morphology classification is fundamental to understanding galaxy formation and evolution. The advent of large-scale sky surveys has produced an unprecedented volume of galaxy images, making traditional manual classification impractical. Although supervised deep learning can achieve high accuracy, it requires large labeled datasets that are time-consuming to construct. In contrast, unsupervised methods often show limited classification performance. To address this limitation, we propose a label-efficient self-supervised learning framework for galaxy morphology classification. Our method first learns robust morphological representations from 305,583 unlabeled KiDS galaxy images through contrastive learning, and then trains a classifier using only 5,000 human-labeled images. The classifier separates galaxies into five categories: elliptical, spiral, lenticular-disk, irregular, and "other." Using a ResNet-50 model with a crop size of 64x64 pixels, our approach achieves an overall test accuracy of up to 91.0% (90.5% +/- 0.2% on average) on the human-classified catalog. The corresponding F1 scores for elliptical, spiral, irregular, lenticular-disk, and "other" galaxies are 0.96, 0.86, 0.86, 0.95, and 0.92, respectively. We apply this pipeline to the Kilo-Degree Survey Data Release 5 and produce a publicly available morphology catalog of 310,583 galaxies. This is the first morphology catalog for KiDS galaxies and provides a valuable resource for future studies of galaxy evolution. Our results show that self-supervised learning can substantially reduce the need for manual labels while maintaining high classification accuracy, making it a promising and scalable approach for automated galaxy morphology classification in the era of large-scale surveys.

astro-ph.GA

Galaxy-Galaxy Strong Lensing simulation with the GPU acceleration across surveys and multi-bands

We present a GPU-accelerated, PyTorch tensor-based simulation framework designed to generate high-fidelity galaxy-galaxy strong lensing images. By integrating synthetic Spectral Energy Distribution (SEDs), the pipeline accurately models the redshift-dependent photometric properties of lens and source galaxies, ensuring physical consistency across multi-band observations. The framework incorporates key observational parameters, including Point Spread Functions (PSF), magnitude limits, and zero points, to replicate specific survey conditions, thereby enabling robust cross-survey joint analyses. As an application, we simulate multi-band images for KiDS, LSST, and Euclid using identical lens model parameters, and employ a deep learning network to evaluate image deblending performance. In particular, the simulation leverages PyTorch to ensure full auto-differentiability and GPU acceleration, making it a highly efficient tool for advanced deep learning algorithms that require gradient-based optimization beyond standard model training. Our framework achieves a speedup of approximately $\mathcal{O}(10^3)$ over traditional CPU-based pipelines, demonstrating the potential feasibility of joint gradient-based lens modeling across next-generation surveys.

astro-ph.GA

Cosmology with galaxy clusters using machine learning. Application to eROSITA Data

Context: We present the first Cosmological Parameter inferences from eROSITA X-ray observations of galaxy clusters using a Machine Learning algorithm. Methods: We train a Random Forest using mock catalogs of clusters from Magneticum multi-cosmology hydrodynamical simulations. We apply the trained ML algorithm to observed X-ray features (gas luminosity, mass, and temperature) at different redshifts from the eROSITA eFEDS and eRASS1 catalogs. Results: We obtain cosmological constraints with precision comparable to those from standard analyses, such as weak lensing and cluster abundances. We infer $Ω_{\rm m}=0.30^{+0.03}_{-0.02}$, $σ_8=0.81\pm0.01$, and $h_0=0.710\pm0.004$. The recovered parameters show no tension in the $Ω_{\rm m}-σ_8$ space, but a significant deviation of $h_0$ from the Planck estimates. These inferences remain rather stable against variations of the input observable set and parameter space coverage. These results indicate that correlations among intracluster properties contain cosmological information beyond that encoded in the cluster abundance alone, which can be captured by machine learning trained on multi-cosmology simulations. Conclusions: ML algorithms trained on multi-cosmology hydrodynamical simulations can effectively infer cosmological parameters directly from galaxy cluster data. This is a change of paradigm in the context of cosmological parameter inferences. This approach complements traditional cluster-count analyses and is particularly suited to large upcoming surveys, where systematic uncertainties in mass calibration may otherwise dominate the error budget. It also highlights the potential of large-scale X-ray surveys to deliver independent tests of the standard cosmological model.

astro-ph.CO

Using Deep Learning Methods to Detect for Ultra-diffuse Galaxies in KiDS

Ultra-diffuse Galaxies (UDGs) are a subset of Low Surface Brightness Galaxies (LSBGs), showing mean effective surface brightness fainter than $24\ \rm mag\ \rm arcsec^{-2}$ and a diffuse morphology, with effective radii larger than 1.5 kpc. Due to their elusiveness, traditional methods are challenging to be used over large sky areas. Here we present a catalog of ultra-diffuse galaxy (UDG) candidates identified in the full 1350 deg$^2$ area of the Kilo-Degree Survey (KiDS) using deep learning. In particular, we use a previously developed network for the detection of low surface brightness systems in the Sloan Digital Sky Survey \citep[LSBGnet,][]{su2024lsbgnet} and optimised for UDG detection. We train this new UDG detection network for KiDS (UDGnet-K), with an iterative approach, starting from a small-scale training sample. After training and validation, the UGDnet-K has been able to identify $\sim3300$ UDG candidates, among which, after visual inspection, we have selected 545 high-quality ones. The catalog contains independent re-discovery of previously confirmed UDGs in local groups and clusters (e.g NGC 5846 and Fornax), and new discovered candidates in about 15 local systems, for a total of 67 {\it bona fide} associations. Besides the value of the catalog {\it per se} for future studies of UDG properties, this work shows the effectiveness of an iterative approach to training deep learning tools in presence of poor training samples, due to the paucity of confirmed UDG examples, which we expect to replicate for upcoming all-sky surveys like Rubin Observatory, Euclid and the China Space Station Telescope.

astro-ph.GA

OpenGrok: Enhancing SNS Data Processing with Distilled Knowledge and Mask-like Mechanisms

This report details Lumen Labs' novel approach to processing Social Networking Service (SNS) data. We leverage knowledge distillation, specifically a simple distillation method inspired by DeepSeek-R1's CoT acquisition, combined with prompt hacking, to extract valuable training data from the Grok model. This data is then used to fine-tune a Phi-3-mini model, augmented with a mask-like mechanism specifically designed for handling the nuances of SNS data. Our method demonstrates state-of-the-art (SOTA) performance on several SNS data processing tasks, outperforming existing models like Grok, Phi-3, and GPT-4. We provide a comprehensive analysis of our approach, including mathematical formulations, engineering details, ablation studies, and comparative evaluations.

cs.LG

Enhancing Large Language Model Efficiencyvia Symbolic Compression: A Formal Approach Towards Interpretability

Large language models (LLMs) face significant token efficiency bottlenecks in code generation and logical reasoning tasks, a challenge that directly impacts inference cost and model interpretability. This paper proposes a formal framework based on symbolic compression,integrating combinatory logic, information-theoretic optimal encoding, and context-aware inference techniques to achieve a step-change improvement in token efficiency while preserving semantic integrity. We establish a mathematical framework within a functional programming paradigm, derive the quantitative relationship between symbolic density and model interpretability, and propose a differentiable compression factor metric to evaluate encoding efficiency. Furthermore, we leverage parameter-efficient fine-tuning (PEFT) techniques to achieve a low-cost application of the GAEL language. Experimental results show that this method achieves a 78.3% token compression rate in code generation tasks while improving logical traceability by 62% through structural explicitness. This research provides new theoretical tools for efficient inference in LLMs and opens a symbolic path for modelinterpretability research.

cs.AI

Chinese Stock Prediction Based on a Multi-Modal Transformer Framework: Macro-Micro Information Fusion

This paper proposes an innovative Multi-Modal Transformer framework (MMF-Trans) designed to significantly improve the prediction accuracy of the Chinese stock market by integrating multi-source heterogeneous information including macroeconomy, micro-market, financial text, and event knowledge. The framework consists of four core modules: (1) A four-channel parallel encoder that processes technical indicators, financial text, macro data, and event knowledge graph respectively for independent feature extraction of multi-modal data; (2) A dynamic gated cross-modal fusion mechanism that adaptively learns the importance of different modalities through differentiable weight allocation for effective information integration; (3) A time-aligned mixed-frequency processing layer that uses an innovative position encoding method to effectively fuse data of different time frequencies and solves the time alignment problem of heterogeneous data; (4) A graph attention-based event impact quantification module that captures the dynamic impact of events on the market through event knowledge graph and quantifies the event impact coefficient. We introduce a hybrid-frequency Transformer and Event2Vec algorithm to effectively fuse data of different frequencies and quantify the event impact. Experimental results show that in the prediction task of CSI 300 constituent stocks, the root mean square error (RMSE) of the MMF-Trans framework is reduced by 23.7% compared to the baseline model, the event response prediction accuracy is improved by 41.2%, and the Sharpe ratio is improved by 32.6%.

cs.LG

Transformer^-1: Input-Adaptive Computation for Resource-Constrained Deployment

Addressing the resource waste caused by fixed computation paradigms in deep learning models under dynamic scenarios, this paper proposes a Transformer$^{-1}$ architecture based on the principle of deep adaptivity. This architecture achieves dynamic matching between input features and computational resources by establishing a joint optimization model for complexity and computation. Our core contributions include: (1) designing a two-layer control mechanism, composed of a complexity predictor and a reinforcement learning policy network, enabling end-to-end optimization of computation paths; (2) deriving a lower bound theory for dynamic computation, proving the system's theoretical reach to optimal efficiency; and (3) proposing a layer folding technique and a CUDA Graph pre-compilation scheme, overcoming the engineering bottlenecks of dynamic architectures. In the ImageNet-1K benchmark test, our method reduces FLOPs by 42.7\% and peak memory usage by 34.1\% compared to the standard Transformer, while maintaining comparable accuracy ($\pm$0.3\%). Furthermore, we conducted practical deployment on the Jetson AGX Xavier platform, verifying the effectiveness and practical value of this method in resource-constrained environments. To further validate the generality of the method, we also conducted experiments on several natural language processing tasks and achieved significant improvements in resource efficiency.

cs.LG

MyGO Multiplex CoT: A Method for Self-Reflection in Large Language Models via Double Chain of Thought Thinking

Recent advancements in large language models (LLMs) have demonstrated their impressive abilities in various reasoning and decision-making tasks. However, the quality and coherence of the reasoning process can still benefit from enhanced introspection and self-reflection. In this paper, we introduce Multiplex CoT (Chain of Thought), a method that enables LLMs to simulate a form of self-review while reasoning, by initiating double Chain of Thought (CoT) thinking. Multiplex CoT leverages the power of iterative reasoning, where the model generates an initial chain of thought and subsequently critiques and refines this reasoning with a second round of thought generation. This recursive approach allows for more coherent, logical, and robust answers, improving the overall decision-making process. We demonstrate how this method can be effectively implemented using simple prompt engineering in existing LLM architectures, achieving an effect similar to that of the Learning-Refinement Model (LRM) without the need for additional training. Additionally, we present a practical guide for implementing the method in Google Colab, enabling easy integration into real-world applications.

cs.CL

Galaxy Spectra Networks (GaSNet). III. Generative pre-trained network for spectrum reconstruction, redshift estimate and anomaly detection

Classification of spectra (1) and anomaly detection (2) are fundamental steps to guarantee the highest accuracy in redshift measurements (3) in modern all-sky spectroscopic surveys. We introduce a new Galaxy Spectra Neural Network (GaSNet-III) model that takes advantage of generative neural networks to perform these three tasks at once with very high efficiency. We use two different generative networks, an autoencoder-like network and U-Net, to reconstruct the rest-frame spectrum (after redshifting). The autoencoder-like network operates similarly to the classical PCA, learning templates (eigenspectra) from the training set and returning modeling parameters. The U-Net, in contrast, functions as an end-to-end model and shows an advantage in noise reduction. By reconstructing spectra, we can achieve classification, redshift estimation, and anomaly detection in the same framework. Each rest-frame reconstructed spectrum is extended to the UV and a small part of the infrared (covering the blueshift of stars). Owing to the high computational efficiency of deep learning, we scan the chi-squared value for the entire type and redshift space and find the best-fitting point. Our results show that generative networks can achieve accuracy comparable to the classical PCA methods in spectral modeling with higher efficiency, especially achieving an average of $>98\%$ classification across all classes ($>99.9\%$ for star), and $>99\%$ (stars), $>98\%$ (galaxies) and $>93\%$ (quasars) redshift accuracy under cosmology research requirements. By comparing different peaks of chi-squared curves, we define the ``robustness'' in the scanned space, offering a method to identify potential ``anomalous'' spectra. Our approach provides an accurate and high-efficiency spectrum modeling tool for handling the vast data volumes from future spectroscopic sky surveys.

astro-ph.GA

Galaxy Spectra neural Network (GaSNet). II. Using Deep Learning for Spectral Classification and Redshift Predictions

Large sky spectroscopic surveys have reached the scale of photometric surveys in terms of sample sizes and data complexity. These huge datasets require efficient, accurate, and flexible automated tools for data analysis and science exploitation. We present the Galaxy Spectra Network/GaSNet-II, a supervised multi-network deep learning tool for spectra classification and redshift prediction. GaSNet-II can be trained to identify a customized number of classes and optimize the redshift predictions for classified objects in each of them. It also provides redshift errors, using a network-of-networks that reproduces a Monte Carlo test on each spectrum, by randomizing their weight initialization. As a demonstration of the capability of the deep learning pipeline, we use 260k Sloan Digital Sky Survey spectra from Data Release 16, separated into 13 classes including 140k galactic, and 120k extragalactic objects. GaSNet-II achieves 92.4% average classification accuracy over the 13 classes (larger than 90% for the majority of them), and an average redshift error of approximately 0.23% for galaxies and 2.1% for quasars. We further train/test the same pipeline to classify spectra and predict redshifts for a sample of 200k 4MOST mock spectra and 21k publicly released DESI spectra. On 4MOST mock data, we reach 93.4% accuracy in 10-class classification and an average redshift error of 0.55% for galaxies and 0.3% for active galactic nuclei. On DESI data, we reach 96% accuracy in (star/galaxy/quasar only) classification and an average redshift error of 2.8% for galaxies and 4.8% for quasars, despite the small sample size available. GaSNet-II can process ~40k spectra in less than one minute, on a normal Desktop GPU. This makes the pipeline particularly suitable for real-time analyses of Stage-IV survey observations and an ideal tool for feedback loops aimed at night-by-night survey strategy optimization.

astro-ph.IM

Galaxy-Galaxy Strong Lensing with U-Net (GGSL-UNet). I. Extracting 2-Dimensional Information from Multi-Band Images in Ground and Space Observations

We present a novel deep learning method to separately extract the two-dimensional flux information of the foreground galaxy (deflector) and background system (source) of Galaxy-Galaxy Strong Lensing events using U-Net (GGSL-Unet for short). In particular, the segmentation of the source image is found to enhance the performance of the lens modeling, especially for ground-based images. By combining mock lens foreground+background components with real sky survey noise to train the GGSL-Unet, we show it can correctly model the input image noise and extract the lens signal. However, the most important result of this work is that the GGSL-UNet can accurately reconstruct real ground-based lensing systems from the Kilo Degree Survey (KiDS) in one second. We also test the GGSL-UNet on space-based (HST) lenses from BELLS GALLERY, and obtain comparable accuracy of standard lens modeling tools. Finally, we calculate the magnitudes from the reconstructed deflector and source images and use this to derive photometric redshifts (photo-z), with the photo-z of the deflector well consistent with spectroscopic ones. This first work, demonstrates the great potential of the generative network for lens finding, image denoising, source segmentation, and decomposing and modeling of strong lensing systems. For the upcoming ground- and space-based surveys, the GGSL-UNet can provide high-quality images as well as geometry and redshift information for precise lens modeling, in combination with classical MCMC modeling for best accuracy in the galaxy-galaxy strong lensing analysis.

astro-ph.GA

Cosmology with Galaxy Cluster Properties using Machine Learning

[Abridged] Galaxy clusters are the most massive gravitationally-bound systems in the universe and are widely considered to be an effective cosmological probe. We propose the first Machine Learning method using galaxy cluster properties to derive unbiased constraints on a set of cosmological parameters, including Omega_m, sigma_8, Omega_b, and h_0. We train the machine learning model with mock catalogs including "measured" quantities from Magneticum multi-cosmology hydrodynamical simulations, like gas mass, gas bolometric luminosity, gas temperature, stellar mass, cluster radius, total mass, velocity dispersion, and redshift, and correctly predict all parameters with uncertainties of the order of ~14% for Omega_m, ~8% for sigma_8, ~6% for Omega_b, and ~3% for h_0. This first test is exceptionally promising, as it shows that machine learning can efficiently map the correlations in the multi-dimensional space of the observed quantities to the cosmological parameter space and narrow down the probability that a given sample belongs to a given cosmological parameter combination. In the future, these ML tools can be applied to cluster samples with multi-wavelength observations from surveys like LSST, CSST, Euclid, Roman in optical and near-infrared bands, and eROSITA in X-rays, to constrain both the cosmology and the effect of the baryonic feedback.

astro-ph.CO

Sommerfeld effect in freeze-in dark matter

If two annihilation products of dark matter (DM) particles are non-relativistic and coupled to a light force mediator, their plane wave functions are modified due to multiple exchanges of the force mediators. This gives rise to the Sommerfeld effect (SE). We consider the attractive and repulsive force SE on the relic density in different phases of freeze-in DM. We find that in the pure freeze-in region, the attractive/repulsive force SE slightly increases/decreases DM relic density by less than $20\%$ for TeV-scale DM. In the reannihilation region, if the portal coupling $κ$ is sufficiently large (by comparing the portal reaction rate to the Hubble rate), DM density will reach its equilibrium, and subsequently freeze out. Compared to the case without the SE, the presence of the attractive SE leads to an enlarged cross-section. As a result, a higher equilibrium value of DM density is reached, and a lower relic density is obtained after the subsequent freeze-out. However, the repulsive SE has the opposite influence. In the dark sector (DS) freeze-out region, also known as the middle flat plateau or ``mesa'' in the phase diagram, the SE has a significant impact on DM relic abundance. In this region, the attractive SE suppresses DM relic density by simultaneously enlarging the cross-section of the portal and DS internal interaction. In contrast, the repulsive SE will have the opposite effect. Finally, in the usual freeze-out region, DM relic density is suppressed or enhanced by an enlarged or reduced cross-section of the portal, respectively, due to the presence of the attractive or repulsive SE. In summary, when considering the constraint of producing correct DM relic abundance, the inclusion of SE in the portal reaction or DS internal reaction will modify the model parameters, resulting in a band-like possible parameter space.

hep-ph

Final bound-state formation effect on dark matter annihilation

If the annihilation products of dark matter (DM) are non-relativistic and couples directly to a light force mediator, the non-perturbation effect like final state bound state (FBS) formation and final state Sommerfeld (FSS) effect must be considered. Non-relativistic region of final particles will appear when there is small mass split between DM and products, so we study those effects in the degenerate region of mass (including kinematics forbidden case) using two specific models. We demonstrate that FBS effect will significantly modify the DM relic abundance comparing to the standard perturbation calculation in some mass split region. We emphasize that FBS effect is comparable to the FSS effect in those mass split. The conservation angular momentum are subtle considering FBS formation, in some cases there may be not $s$-wave, so we use two models exhibit the different partial wave FBS effect contribution. We also show that the FBS formation with vector boson emission process also contributes in DM relic abundance, and first calculate the $p$-wave FSS effect in the specific model.

hep-ph

Galaxy Spectra neural Networks (GaSNets). I. Searching for strong lens candidates in eBOSS spectra using Deep Learning

With the advent of new spectroscopic surveys from ground and space, observing up to hundreds of millions of galaxies, spectra classification will become overwhelming for standard analysis techniques. To prepare for this challenge, we introduce a family of deep learning tools to classify features in one-dimensional spectra. As the first application of these Galaxy Spectra neural Networks (GaSNets), we focus on tools specialized at identifying emission lines from strongly lensed star-forming galaxies in the eBOSS spectra. We first discuss the training and testing of these networks and define a threshold probability, PL, of 95% for the high quality event detection. Then, using a previous set of spectroscopically selected strong lenses from eBOSS, confirmed with HST, we estimate a completeness of ~80% as the fraction of lenses recovered above the adopted PL. We finally apply the GaSNets to ~1.3M spectra to collect a first list of ~430 new high quality candidates identified with deep learning applied to spectroscopy and visually graded as highly probable real events. A preliminary check against ground-based observations tentatively shows that this sample has a confirmation rate of 38%, in line with previous samples selected with standard (no deep learning) classification tools and follow-up by Hubble Space Telescope. This first test shows that machine learning can be efficiently extended to feature recognition in the wavelength space, which will be crucial for future surveys like 4MOST, DESI, Euclid, and the Chinese Space Station Telescope (CSST).

astro-ph.GA