SearcharxivSearch

arXiv subjects

Zhenyu Jin

Publications and source records attributed to Zhenyu Jin.

16 recordsLinked to original sources

Harnessing the Potential of Optimizing Data Mixtures via Bayesian Domain Reweighting

The performance of Large Language Models (LLMs) is fundamentally influenced by the distributional composition of multi-domain pre-training data. While manual heuristics were prevalent in early models, they increasingly fail to capture the intricate synergies between domains as data complexity grows. To overcome the issue, a dominant approach seeks to fit a proxy function mapping between domain weights and their corresponding validation losses, and then find the optimal domain weights to minimize validation losses. These methods rely on strong structural assumptions, such as rank invariance or scaling laws, which are often violated, resulting in non-negligible estimation bias. A promising approach is to directly optimize the weighting scheme from data. However, it suffers from unstable optimization trajectory and prohibitive computational overhead, limiting its potential to search better domain weights configurations. This paper presents a Bayesian domain weighting method to infer the weights from a Dirichlet distribution via introducing Gamma prior information learned from observations. Experimental results demonstrate that proposed method could achieve stable and efficient domain weights learning, and identifies optimal mixtures while consuming substantially less data than search-based function-fitting methods, revitalizing optimization-based domain weighting for large-scale applications.

cs.LG

Deep Learning with Magnetic Parameter Constraints for Short-Term Prediction of Solar Active Region Vector Magnetic Fields

Forecasting the dynamic evolution of solar magnetic fields is a critical technique for enabling space weather warnings. Addressing the limitations of existing methods in predicting all vector magnetic field components and in maintaining consistency with solar surface magnetic-field-related quantities, this study proposes a deep learning prediction method that integrates dynamic masks of active regions with multiple magnetic parameter constraints. By constructing a three-channel representation of vector magnetic fields, applying dynamic masks to enhance attention to strong-field regions, and incorporating multi-parameter magnetic parameter constraints, we developed an end-to-end short-term (12-hour) predictive model of solar vector magnetic field evolution. Using SDO/SHARP vector magnetogram data, the model predicts and analyses field evolution across all components. Quantitative evaluations demonstrate that our approach achieves horizon-averaged structural similarity index measure (SSIM) of 0.912 (per-hour range: 0.909--0.916) and correlation coefficient (CC) of 0.998 for the radial component Br (root-mean-square error (RMSE) 13.0--21.0 G); the horizontal components achieve Bphi SSIM 0.760--0.800 (CC 0.910--0.945, RMSE 38.5--50.0 G) and Btheta SSIM 0.728--0.750 (CC 0.895--0.920, RMSE 38.5--49.0 G). The model maintains unsigned magnetic flux prediction errors at 7.82% (95% confidence interval (CI): +/-0.11%). These results demonstrate strong image-domain performance together with consistency under the magnetic-parameter diagnostics used here, suggesting initial potential for supporting future space weather forecasting efforts.

astro-ph.SR

Machine Learning-based Separation of the He I 10830Å Chromospheric Signal: Quantitative Analysis of Chromosphere-Corona Intensity in the Quiet Sun

The He I 10830Å line, a crucial optically thin chromospheric line, is frequently used to study coronal heating and vertical coupling across the chromosphere-corona interface. However, its images are severely contaminated by the strong photospheric background signal, hindering the analysis of fine chromospheric structures. Given the morphological differences between the Active Region (AR) and the Quiet Sun (QS), we proposed separating the He I 10830Å chromospheric signal using two deep learning CNN models. Our model utilizes TiO images and cross-band learning to infer the He I 10830Å photospheric background. The output is combined with an exponential absorption model to achieve quantitative analysis of the pure chromospheric component. Joint analysis of Solar Dynamics Observatory (SDO) data and the separated QS structures reveals a strong spatial negative correlation between chromospheric He I 10830Å intensities(R approx -0.84 in 304Å ), and significant layered coupling with EUV (171, 193, and 304Å) radiation. Furthermore, strong He I 10830Å absorption areas are highly correlated with regions of strong magnetic fields, while 171Å radiative enhancement areas extend to the strong magnetic field edges and the mixed-polarity regions. These findings quantify the radiation intensity relationship between He I 10830Å and EUV bands in the Quiet Sun. It also demonstrates the differences in heating characteristics between unipolar and mixed-polarity magnetic fields.

astro-ph.SR

A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle

Most current paradigms in visual mechanistic interpretability (MI) remain confined to interpreting internal units of the vision model via heuristic methods (e.g., top-$K$ activation retrieval or optimization with regularization). In this work, we establish a theoretical distributional view for visual MI, which models the influence of a feature activation on the natural image distribution, thereby formulating a Kullback-Leibler (KL)-minimal optimization problem to model the MI task. Under this framework, statistical biases are identified within previous MI paradigms, which reveal that they may either be perceptually uninterpretable to humans (i.e., deviate from the natural image distribution), or mechanistically unfaithful to the vision models (i.e., unable to activate model features). To resolve the biases under the distributional view, we propose a model with a KL-minimal soft-constraint principle for visual MI that theoretically balances interpretability and faithfulness. We realize this principle via energy-guided diffusion posterior sampling. Extensive experiments validate the theoretical soundness of the proposed distributional view and demonstrate the practical effectiveness of our paradigm on the DINOv3 vision model.

cs.CV

Tracing the Thought of a Grandmaster-level Chess-Playing Transformer

While modern transformer neural networks achieve grandmaster-level performance in chess and other reasoning tasks, their internal computation process remains largely opaque. Focusing on Leela Chess Zero (LC0), we introduce a sparse decomposition framework to interpret its internal computation by decomposing its MLP and attention modules with sparse replacement layers, which capture the primary computation process of LC0. We conduct a detailed case study showing that these pathways expose rich, interpretable tactical considerations that are empirically verifiable. We further introduce three quantitative metrics and show that LC0 exhibits parallel reasoning behavior consistent with the inductive bias of its policy head architecture. To the best of our knowledge, this is the first work to decompose the internal computation of a transformer on both MLP and attention modules for interpretability. Combining sparse replacement layers and causal interventions in LC0 provides a comprehensive understanding of advanced tactical reasoning, offering critical insights into the underlying mechanisms of superhuman systems. Our code is available at https://github.com/JacklE0niden/Leela-SAEs.

cs.LG

SpecGen: Neural Spectral BRDF Generation via Spectral-Spatial Tri-plane Aggregation

Synthesizing spectral images across different wavelengths is essential for photorealistic rendering. Unlike conventional spectral uplifting methods that convert RGB images into spectral ones, we introduce SpecGen, a novel method that generates spectral bidirectional reflectance distribution functions (BRDFs) from a single RGB image of a sphere. This enables spectral image rendering under arbitrary illuminations and shapes covered by the corresponding material. A key challenge in spectral BRDF generation is the scarcity of measured spectral BRDF data. To address this, we propose the Spectral-Spatial Tri-plane Aggregation (SSTA) network, which models reflectance responses across wavelengths and incident-outgoing directions, allowing the training strategy to leverage abundant RGB BRDF data to enhance spectral BRDF generation. Experiments show that our method accurately reconstructs spectral BRDFs from limited spectral data and surpasses state-of-the-art methods in hyperspectral image reconstruction, achieving an improvement of 8 dB in PSNR. Codes and data will be released upon acceptance.

cs.CV

Using Neural Emulators and Hamiltonian Monte Carlo to constrain the Epoch of Reionization's History with the Ly$α$ Forest Power Spectrum

The Lyman-alpha (Ly$α$) forest at $z \sim 5$ offers a primary probe to constrain the history of the Epoch of Reionization (EoR), retaining thermal and ionization signatures imprinted by the reionization process. In this work, we present a new inference framework based on JAX that combines forward-modeled Ly$α$ forest observables with differentiable neural emulators and Hamiltonian Monte Carlo (HMC). We construct a dataset of 501 low-resolution simulations generated with user-defined reionization histories and compute a set of 1D Ly$α$ power spectra and model-dependent covariance matrices. We then train two independent neural emulators that achieve sub-percent errors across relevant scales and combine them with HMC to efficiently perform parameter estimation. We validate this framework by applying it to a suite of mock observations, demonstrating that the true parameters are reliably recovered. While this work is limited by the low resolution of the simulations used, our results highlight the potential of this method for inferring the reionization history from high-redshift Ly$α$ forest measurements. Future improvements in our reionization models will further enhance its ability to extract constraints from observational datasets.

astro-ph.CO

Compressive Imaging Reconstruction via Tensor Decomposed Multi-Resolution Grid Encoding

Compressive imaging (CI) reconstruction, such as snapshot compressive imaging (SCI) and compressive sensing magnetic resonance imaging (MRI), aims to recover high-dimensional images from low-dimensional compressed measurements. This process critically relies on learning an accurate representation of the underlying high-dimensional image. However, existing unsupervised representations may struggle to achieve a desired balance between representation ability and efficiency. To overcome this limitation, we propose Tensor Decomposed multi-resolution Grid encoding (GridTD), an unsupervised continuous representation framework for CI reconstruction. GridTD optimizes a lightweight neural network and the input tensor decomposition model whose parameters are learned via multi-resolution hash grid encoding. It inherently enjoys the hierarchical modeling ability of multi-resolution grid encoding and the compactness of tensor decomposition, enabling effective and efficient reconstruction of high-dimensional images. Theoretical analyses for the algorithm's Lipschitz property, generalization error bound, and fixed-point convergence reveal the intrinsic superiority of GridTD as compared with existing continuous representation models. Extensive experiments across diverse CI tasks, including video SCI, spectral SCI, and compressive dynamic MRI reconstruction, consistently demonstrate the superiority of GridTD over existing methods, positioning GridTD as a versatile and state-of-the-art CI reconstruction method.

eess.IV

NTIRE 2025 Challenge on Day and Night Raindrop Removal for Dual-Focused Images: Methods and Results

This paper reviews the NTIRE 2025 Challenge on Day and Night Raindrop Removal for Dual-Focused Images. This challenge received a wide range of impressive solutions, which are developed and evaluated using our collected real-world Raindrop Clarity dataset. Unlike existing deraining datasets, our Raindrop Clarity dataset is more diverse and challenging in degradation types and contents, which includes day raindrop-focused, day background-focused, night raindrop-focused, and night background-focused degradations. This dataset is divided into three subsets for competition: 14,139 images for training, 240 images for validation, and 731 images for testing. The primary objective of this challenge is to establish a new and powerful benchmark for the task of removing raindrops under varying lighting and focus conditions. There are a total of 361 participants in the competition, and 32 teams submitting valid solutions and fact sheets for the final testing phase. These submissions achieved state-of-the-art (SOTA) performance on the Raindrop Clarity dataset. The project can be found at https://lixinustc.github.io/CVPR-NTIRE2025-RainDrop-Competition.github.io/.

cs.CV

A High-Accuracy Alignment Approach for Solar Images of Different Wavelengths

Image alignment plays a crucial role in solar physics research, primarily involving translation, rotation, and scaling. \G{The different wavelength images of the chromosphere and transition region have structural complexity and differences in similarity, which poses a challenge to their alignment.} Therefore, a novel alignment approach based on dense optical flow (OF) and the RANSAC algorithm is proposed in this paper. \G{It takes the OF vectors of similar regions between images to be used as feature points for matching. Then, it calculates scaling, rotation, and translation.} The study selects three wavelengths for two groups of alignment experiments: the 304 Å of the Atmospheric Imaging Assembly (AIA), the 1216 Å of the Solar Disk Imager (SDI), and the 465 Å of the Solar Upper Transition Region Imager (SUTRI). Two methods are used to evaluate alignment accuracy: Monte Carlo simulation and Uncertainty Analysis Based on the Jacobian Matrix (UABJM). \G{The evaluation results indicate that this approach achieves sub-pixel accuracy in the alignment of AIA 304 Å and SDI 1216 Å, while demonstrating higher accuracy in the alignment of AIA 304 Å and SUTRI 465 Å, which have greater similarity.

astro-ph.IM

Improving image quality of the Solar Disk Imager (SDI) of the Lyman-alpha Solar Telescope (LST) onboard the ASO-S mission

The in-flight calibration and performance of the Solar Disk Imager (SDI), which is a pivotal instrument of the Lyman-alpha Solar Telescope (LST) onboard the Advanced Space-based Solar Observatory (ASO-S) mission, suggested a much lower spatial resolution than expected. In this paper, we developed the SDI point-spread function (PSF) and Image Bivariate Optimization Algorithm (SPIBOA) to improve the quality of SDI images. The bivariate optimization method smartly combines deep learning with optical system modeling. Despite the lack of information about the real image taken by SDI and the optical system function, this algorithm effectively estimates the PSF of the SDI imaging system directly from a large sample of observational data. We use the estimated PSF to conduct deconvolution correction to observed SDI images, and the resulting images show that the spatial resolution after correction has increased by a factor of more than three with respect to the observed ones. Meanwhile, our method also significantly reduces the inherent noise in the observed SDI images. The SPIBOA has now been successfully integrated into the routine SDI data processing, providing important support for the scientific studies based on the data. The development and application of SPIBOA also pave new ways to identify astronomical telescope systems and enhance observational image quality. Some essential factors and precautions in applying the SPIBOA method are also discussed.

astro-ph.SR

Neural network emulator to constrain the high-$z$ IGM thermal state from Lyman-$α$ forest flux auto-correlation function

We present a neural network emulator to constrain the thermal parameters of the intergalactic medium (IGM) at $\displaystyle{5.4}\le{z}\le{6.0}$ using the Lyman-$\displaystyleα$ (Ly$\displaystyleα$) forest flux auto-correlation function. Our auto-differentiable JAX-based framework accelerates the surrogate model generation process using approximately 100 sparsely sampled Nyx hydrodynamical simulations with varying combinations of thermal parameters, i.e., the temperature at mean density $\displaystyle{T}_{0}$, the slope of the temperature$\displaystyle-$density relation $\displaystyleγ$, and the mean transmission flux $\displaystyle{\langle}{F}{\rangle}$. We show that this emulator has a typical accuracy of 1.0% across the specified redshift range. Bayesian inference of the IGM thermal parameters, incorporating emulator uncertainty propagation, is further expedited using NumPyro Hamiltonian Monte Carlo. We compare both the inference results and computational cost of our framework with the traditional nearest-neighbor interpolation approach applied to the same set of mock Ly$α$ flux. By examining the credibility contours of the marginalized posteriors for $\displaystyle{T}_{0},γ,\text{and}{\langle}{F}{\rangle}$ obtained using the emulator, the statistical reliability of measurements is established through inference on 100 realistic mock data sets of the auto-correlation function.

astro-ph.CO

Research on fine co-focus adjustment method for segmented solar telescope

For segmented telescopes, achieving fine co-focus adjustment is essential for realizing co-phase adjustment and maintenance, which involves adjusting the millimeter-scale piston between segments to fall within the capture range of the co-phase detection system. CGST proposes using a SHWFS for piston detection during the co-focus adjustment stage. However, the residual piston after adjustment exceeds the capture range of the broadband PSF phasing algorithm$(\pm 30 μm) $, and the multi-wavelength PSF algorithm requires even higher precision in co-focus adjustment. To improve the co-focus adjustment accuracy of CGST, a fine co-focus adjustment based on cross-calibration is proposed. This method utilizes a high-precision detector to calibrate and fit the measurements from the SHWFS, thereby reducing the impact of atmospheric turbulence and systematic errors on piston measurement accuracy during co-focus adjustment. Simulation results using CGST demonstrate that the proposed method significantly enhances adjustment accuracy compared to the SHWFS detection method. Additionally, the residual piston after fine co-focus adjustment using this method falls within the capture range of the multi-wavelength PSF algorithm. To verify the feasibility of this method, experiments were conducted on an 800mm ring segmented mirror system, successfully achieving fine co-focus adjustment where the remaining piston of all segments fell within $\pm 15 μm$.

astro-ph.IM

Locating heating channels of the solar corona in a plage region with the aid of high-resolution 10830 Å filtergrams

In this paper, with a set of high-resolution He I 10830 Å filtergrams, we select an area in a plage, very likely an EUV moss area, as an interface layer to follow the clues of coronal heating channels down to the photosphere. The filtergrams are obtained from the 1-meter aperture New Vacuum Solar Telescope (NVST). We make a distinction between the darker and the brighter regions in the selected area and name the two regions enhanced absorption patches (EAPs) and low absorption patches (LAPs). With well-aligned, nearly simultaneous data from multiple channels of the AIA and the continuum of the HMI on board SDO, we compare the EUV/UV emissions, emission measure, mean temperature, and continuum intensity in the two kinds of regions. The following progress is made: 1) The mean EUV emissions over EAPs are mostly stronger than the corresponding emissions over LAPs except for the emission at 335 Å. The UV emissions at 1600 and 1700 Å fail to capture the difference between the two regions. 2) In the logarithmic temperature range of 5.6-6.2, EAPs have higher EUV emission measure than LAPs, but they have lower mean coronal temperature. 3) The mean continuum intensity over EAPs is lower. Based on the above progress, we suggest that the energy for coronal heating in the moss region can be traced down to some areas in intergranular lanes with enhanced density of both cool and hot material. The lower temperature over the EAPs is due to the greater fraction of cool material over there.

astro-ph.SR

High-resolution Solar Image Reconstruction Based on Non-rigid Alignment

Suppressing the interference of atmospheric turbulence and obtaining observation data with a high spatial resolution is an issue to be solved urgently for ground observations. One way to solve this problem is to perform a statistical reconstruction of short-exposure speckle images. Combining the rapidity of Shift-Add and the accuracy of speckle masking, this paper proposes a novel reconstruction algorithm-NASIR (Non-rigid Alignment based Solar Image Reconstruction). NASIR reconstructs the phase of the object image at each frequency by building a computational model between geometric distortion and intensity distribution and reconstructs the modulus of the object image on the aligned speckle images by speckle interferometry. We analyzed the performance of NASIR by using the correlation coefficient, power spectrum, and coefficient of variation of intensity profile (CVoIP) in processing data obtained by the NVST (1m New Vacuum Solar Telescope). The reconstruction experiments and analysis results show that the quality of images reconstructed by NASIR is close to speckle masking when the seeing is good, while NASIR has excellent robustness when the seeing condition becomes worse. Furthermore, NASIR reconstructs the entire field of view in parallel in one go, without phase recursion and block-by-block reconstruction, so its computation time is less than half that of speckle masking. Therefore, we consider NASIR is a robust and high-quality fast reconstruction method that can serve as an effective tool for data filtering and quick look.

astro-ph.IM

Objective Image Quality Assessment for High Resolution Photospheric Images by Median Filter Gradient Similarity

All next generation ground-based and space-based solar telescopes require a good quality assessment metric in order to evaluate their imaging performance. In this paper, a new image quality metric, the median filter gradient similarity (MFGS) is proposed for photospheric images. MFGS is a no-reference/blind objective image quality metric (IQM) by a measurement result between 0 and 1 and has been performed on short-exposure photospheric images captured by the New Vacuum Solar Telescope (NVST) of the Fuxian Solar Observatory and by the Solar Optical Telescope (SOT) onboard the Hinode satellite, respectively. The results show that: (1)the measured value of MFGS changes monotonically from 1 to 0 with degradation of image quality; (2)there exists a linear correlation between the measured values of MFGS and root-mean-square-contrast (RMS-contrast) of granulation; (3)MFGS is less affected by the image contents than the granular RMS-contrast. Overall, MFGS is a good alternative for the quality assessment of photospheric images.

astro-ph.IM