Searcharxiv⌕ Search

arXiv subjects

Hui Sun

Publications and source records attributed to Hui Sun.

At least 55 records · Page 3Linked to original sources

{\tt RapidGBM}: An Efficient Tool for Fermi-GBM Visibility Checking and Data Analysis with a Case Study of EP240617a

We have developed a lightweight tool, {\tt RapidGBM}, featuring a web-based interface and capabilities of rapid calculation of Fermi Gamma-ray Burst Monitor (GBM) visibilities and performance of basic data analysis. It has two key features: (1) it can immediately check the visibility of Fermi-GBM for new transients, and (2) it can check the light curve and perform spectral analysis after the hourly Time-Tagger Event data are released. The visibility check and the response matrix generation required for spectral analysis can be achieved through the historical pointing file after the orbit calculation, even when the real-time pointing file is not yet available. As a case study, we apply the tool to EP240617a, an X-ray transient triggered by Einstein Probe (EP). We demonstrate the workflow of visibility checking, data processing, and spectral analysis for this event. The results suggest that EP240617a can be classified as an X-ray-rich gamma-ray burst (XRR) and confirm the feasibility of using historical pointing files for rapid analysis. Further, we discuss possible physical interpretations of such events, including implications for jet launching and progenitor scenarios. Therefore, {\tt RapidGBM} is expected to assist EP Transient Advocates, Space-based multiband astronomical Variable Objects Monitor burst advocates, and other members of the community in cross checking high-energy transients. Based on prompt emission parameter relations (e.g. $E_{\rm p}$-$E_{γ,\rm iso}$), it can also help identify peculiar GRBs (e.g. long-short burst, magnetar giant flare, etc.) and provide useful references (e.g. more accurate $T_0$) for scheduling follow-up observations.

astro-ph.HE↗

Ovis2.5 Technical Report

We present Ovis2.5, a successor to Ovis2 designed for native-resolution visual perception and strong multimodal reasoning. Ovis2.5 integrates a native-resolution vision transformer that processes images at their native, variable resolutions, avoiding the degradation from fixed-resolution tiling and preserving both fine detail and global layout -- crucial for visually dense content like complex charts. To strengthen reasoning, we train the model to move beyond linear chain-of-thought and perform reflection -- including self-checking and revision. This advanced capability is exposed as an optional "thinking mode" at inference time, allowing users to trade latency for enhanced accuracy on difficult inputs. The model is trained via a comprehensive five-phase curriculum that progressively builds its skills. The process begins with foundational visual and multimodal pretraining, advances through large-scale instruction tuning, and culminates in alignment and reasoning enhancement using DPO and GRPO. To scale these upgrades efficiently, we employ multimodal data packing and hybrid parallelism, yielding a significant end-to-end speedup. We release two open-source models: Ovis2.5-9B and Ovis2.5-2B. The latter continues the "small model, big performance" philosophy of Ovis2, making it ideal for resource-constrained, on-device scenarios. On the OpenCompass multimodal leaderboard, Ovis2.5-9B averages 78.3, marking a substantial improvement over its predecessor, Ovis2-8B, and achieving state-of-the-art results among open-source MLLMs in the sub-40B parameter range; Ovis2.5-2B scores 73.9, establishing SOTA for its size. Beyond aggregate scores, Ovis2.5 achieves leading results on STEM benchmarks, exhibits strong capabilities on grounding and video tasks, and achieves open-source SOTA at its scale for complex chart analysis.

cs.CV↗

Spectral Hardening Reveals Afterglow Emergence in Long-Duration Fast X-ray Transients: A Case Study of GRB 250404A/EP250404a

The prompt emission and afterglow phases of gamma-ray bursts (GRBs) have been extensively studied, yet the transition between these two phases remains inadequately characterized due to limited multiwavelength observational coverage. Among the recent growing samples of fast X-ray transients observed by Einstein Probe (EP), a subgroup of GRBs are captured with long-duration X-ray emission, potentially containing featured evolution from prompt emission to the afterglow phase. In this Letter, we present a detailed analysis of GRB 250404A/EP250404a, a bright fast X-ray transient detected simultaneously by EP and the Fermi Gamma-ray Burst Monitor in X-rays and gamma rays. Its continuous X-ray emission reveals a long-duration tail, accompanied by distinct spectral evolution manifested by the spectral index $α_{\rm X}$ with an initial softening, followed by an evident hardening, eventually reaching a plateau at the value of $\sim$ -2. Early optical and near-infrared observations enable broadband modeling with forward- and reverse-shock components, confirming that the X-ray hardening signals the emergence of the external-shock afterglow. From this spectral hardening we infer that the prompt phase in soft X-rays lasted $\sim300\;\mathrm{s}$, which is more than 3 times longer than the gamma-ray $T_{90}$. This well-tracked soft-hard-flat spectral pattern provides a clear indication of afterglow emergence from the fading prompt emission and offers a practical criterion for identifying a distinct population of GRBs among fast X-ray transients, even when the detection of the gamma-ray counterpart or obvious temporal break is absent.

astro-ph.HE↗

Thermodynamic modeling of binaries in Cr-Fe-Mo-Nb-Ni supported by first-principles calculations

Thermodynamic descriptions of all binaries within the Cr-Fe-Mo-Nb-Ni system have been complied and, where necessary, remodeled. Notably, the Cr-Fe and Fe-Mo systems have been remodeled using comprehensive sublattice models for the topologically close-packed (TCP) phases of Laves_C14, sigma, and mu according to their Wyckoff positions. These refinements are supported by first-principles calculations based on density functional theory (DFT), in conjunction with available experimental data in the literature. The resulting models offer improved accuracy in describing the TCP phases. For instance, the predicted site occupancies of sigma in Cr-Fe show excellent agreement with experimental observations. The present work provides a robust foundation for CALPHAD modeling and the design of complex, multi-component materials, particularly those based on Fe-based and Ni-based alloys.

cond-mat.mtrl-sci↗

PMKLC: Parallel Multi-Knowledge Learning-based Lossless Compression for Large-Scale Genomics Database

Learning-based lossless compressors play a crucial role in large-scale genomic database backup, storage, transmission, and management. However, their 1) inadequate compression ratio, 2) low compression \& decompression throughput, and 3) poor compression robustness limit their widespread adoption and application in both industry and academia. To solve those challenges, we propose a novel \underline{P}arallel \underline{M}ulti-\underline{K}nowledge \underline{L}earning-based \underline{C}ompressor (PMKLC) with four crucial designs: 1) We propose an automated multi-knowledge learning-based compression framework as compressors' backbone to enhance compression ratio and robustness; 2) we design a GPU-accelerated ($s$,$k$)-mer encoder to optimize compression throughput and computing resource usage; 3) we introduce data block partitioning and Step-wise Model Passing (SMP) mechanisms for parallel acceleration; 4) We design two compression modes PMKLC-S and PMKLC-M to meet the complex application scenarios, where the former runs on a resource-constrained single GPU and the latter is multi-GPU accelerated. We benchmark PMKLC-S/M and 14 baselines (7 traditional and 7 leaning-based) on 15 real-world datasets with different species and data sizes. Compared to baselines on the testing datasets, PMKLC-S/M achieve the average compression ratio improvement up to 73.609\% and 73.480\%, the average throughput improvement up to 3.036$\times$ and 10.710$\times$, respectively. Besides, PMKLC-S/M also achieve the best robustness and competitive memory cost, indicating its greater stability against datasets with different probability distribution perturbations, and its strong ability to run on memory-constrained devices.

cs.LG↗

Batch Sample-wise Stochastic Optimal Control via Stochastic Maximum Principle

In this work, we study the stochastic optimal control problem (SOC) mainly from the probabilistic view point, i.e. via the Stochastic Maximum principle (SMP) \cite{Peng4}. We adopt the sample-wise backpropagation scheme proposed in \cite{Hui1} to solve the SOC problem under the strong convexity assumption. Importantly, in the Stochastic Gradient Descent (SGD) procedure, we use batch samples with higher order scheme in the forward SDE to improve the convergence rate in \cite{Hui1} from $\sim \mathcal{O}(\sqrt{\frac{N}{K} + \frac{1}{N}})$ to $\sim \mathcal{O}(\sqrt{\frac{1}{K} + \frac{1}{N^2}})$ and note that the main source of uncertainty originates from the scheme for the simulation of $Z$ term in the BSDE. In the meantime, we note the SGD procedure uses only the necessary condition of the SMP, while the batch simulation of the approximating solution of BSDEs allows one to obtain a more accurate estimate of the control $u$ that minimizes the Hamiltonian. We then propose a damped contraction algorithm to solve the SOC problem whose proof of convergence for a special case is attained under some appropriate assumption. We then show numerical results to check the first order convergence rate of the projection algorithm and analyze the convergence behavior of the damped contraction algorithm. Lastly, we briefly discuss how to incorporate the proposed scheme in solving practical problems especially when the Randomized Neural Networks are used. We note that in this special case, the error backward propagation can be avoided and parameter update can be achieved via purely algebraic computation (vector algebra) which will potentially improve the efficiency of the whole training procedure. Such idea will require further exploration and we will leave it as our future work.

math.OC↗

A TRPCA-Inspired Deep Unfolding Network for Hyperspectral Image Denoising via Thresholded t-SVD and Top-K Sparse Transformer

Hyperspectral images (HSIs) are often degraded by complex mixed noise during acquisition and transmission, making effective denoising essential for subsequent analysis. Recent hybrid approaches that bridge model-driven and data-driven paradigms have shown great promise. However, most of these approaches lack effective alternation between different priors or modules, resulting in loosely coupled regularization and insufficient exploitation of their complementary strengths. Inspired by tensor robust principal component analysis (TRPCA), we propose a novel deep unfolding network (DU-TRPCA) that enforces stage-wise alternation between two tightly integrated modules: low-rank and sparse. The low-rank module employs thresholded tensor singular value decomposition (t-SVD), providing a widely adopted convex surrogate for tensor low-rankness and has been demonstrated to effectively capture the global spatial-spectral structure of HSIs. The Top-K sparse transformer module adaptively imposes sparse constraints, directly matching the sparse regularization in TRPCA and enabling effective removal of localized outliers and complex noise. This tightly coupled architecture preserves the stage-wise alternation between low-rank approximation and sparse refinement inherent in TRPCA, while enhancing representational capacity through attention mechanisms. Extensive experiments on synthetic and real-world HSIs demonstrate that DU-TRPCA surpasses state-of-the-art methods under severe mixed noise, while offering interpretability benefits and stable denoising dynamics inspired by iterative optimization. Code is available at https://github.com/liangli97/TRPCA-Deep-Unfolding-HSI-Denoising.

cs.CV↗

Accelerating Bayesian Sampling for Massive Black Hole Binaries with Prior Constraints from Conditional Variational Autoencoder

A Conditional Variational Autoencoder (CVAE) model is employed for parameter inference on gravitational waves (GW) signals of massive black hole binaries, considering joint observations with a network of three space-based GW detectors. Our experiments show that the trained CVAE model can estimate the posterior distribution of source parameters in approximately one second, while the standard Bayesian sampling method, utilizing parallel computation across 16 CPU cores, takes an average of 20 hours for a GW signal instance. However, the sampling distributions from CVAE exhibit lighter tails, appearing broader when compared to the standard Bayesian sampling results. By using CVAE results to constrain the prior range for Bayesian sampling, the sampling time is reduced by a factor of $\sim$6 while maintaining the similar precision of the Bayesian results.

astro-ph.IM↗

The Soft X-ray Aspect of Gamma-ray Bursts in the Einstein Probe Era

The Einstein Probe (EP) satellite, dedicated at time-domain high-energy astrophysics and multi-messenger astronomy, was recently launched and successfully put into operation. The wide-field X-ray telescope (WXT, 0.5-4 keV) onboard has identified multiple gamma-ray burst (GRB) events, with an average duration of several hundred seconds. This duration is several times longer than the average duration of long gamma-ray bursts (LGRBs) detected by the Neil Gehrels Swift Observatory, which typically stands at several tens of seconds. Additionally, EP has detected some unknown X-ray transients whose connection to GRBs is uncertain, due to the absence of gamma-ray counterparts and efficient follow-up observation at multi-wavelengths. Several main factors could account for the longer time, including the Doppler effect of off-axis viewing, the spectral lag effect of the synchrotron spectrum of cooling electrons, and some unknown prolonged intrinsic X-ray activities. Our studies indicate that EP GRBs may primarily consist of off-axis viewed bursts, forming a unique population among the GRB zoo, yet the intrinsic origin for the specific bursts could not be excluded. By analyzing the statistical properties of the historical LGRB samples, we explored observable properties of on-axis and off-axis LGRBs in the soft X-ray band. The predicted characteristics of off-axis viewed GRBs, including the duration, energy fluence, low-energy spectral index, and the slopes of Amati and Yonetoku relations, could be tested with a larger sample of GRB events detected by EP in the future.

astro-ph.HE↗

EP240801a/XRF 240801B: An X-ray Flash Detected by the Einstein Probe and Implications of its Multiband Afterglow

We present multiband observations and analysis of EP240801a, a low-energy, extremely soft gamma-ray burst (GRB) discovered on August 1, 2024 by the Einstein Probe (EP) satellite, with a weak contemporaneous signal also detected by Fermi/GBM. Optical spectroscopy of the afterglow, obtained by GTC and Keck, identified the redshift of $z = 1.6734$. EP240801a exhibits a burst duration of 148 s in X-rays and 22.3 s in gamma-rays, with X-rays leading by 80.61 s. Spectral lag analysis indicates the gamma-ray signal arrived 8.3 s earlier than the X-rays. Joint spectral fitting of EP/WXT and Fermi/GBM data yields an isotropic energy $E_{γ,\rm{iso}} = (5.57^{+0.54}_{-0.50})\times 10^{51}\,\rm{erg}$, a peak energy $E_{\rm{peak}} = 14.90^{+7.08}_{-4.71}\,\rm{keV}$, a fluence ratio $\rm S(25-50\,\rm{keV})/S(50-100\,\rm{keV}) = 1.67^{+0.74}_{-0.46}$, classifying EP240801a as an X-ray flash (XRF). The host-galaxy continuum spectrum, inferred using Prospector, was used to correct its contribution for the observed outburst optical data. Unusual early $R$-band behavior and EP/FXT observations suggest multiple components in the afterglow. Three models are considered: two-component jet model, forward-reverse shock model and forward-shock model with energy injection. Both three provide reasonable explanations. The two-component jet model and the energy injection model imply a relatively small initial energy and velocity of the jet in the line of sight, while the forward-reverse shock model remains typical. Under the two-component jet model, EP240801a may resemble GRB 221009A (BOAT) if the bright narrow beam is viewed on-axis. Therefore, EP240801a can be interpreted as an off-beam (narrow) jet or an intrinsically weak GRB jet. Our findings provide crucial clues for uncovering the origin of XRFs.

astro-ph.HE↗

LEIA discovery of the longest-lasting and most energetic stellar X-ray flare ever detected

The Lobster Eye Imager for Astronomy (LEIA) detected a new X-ray transient on 2022 November 7, identified as a superflare event occurring on a nearby K-type giant star HD 251108. The flux increase was also detected in follow-up observations at X-ray, UV and optical wavelengths. The flare lasted for about 40 days in soft X-ray observations, reaching a peak luminosity of ~1.1 * 10^34 erg/s in 0.5-4.0 keV, which is roughly 60 times the quiescent luminosity. Optical brightening was observed for only one night. The X-ray light curve is well described by a double fast rise and exponential decay model, attributed to the cooling process of a loop arcade structure formed subsequent to the initial large loop with a half-length of ~1.9 * 10^12 cm. Time-resolved X-ray spectra were fitted by a four-temperature apec model (with three components being the quiescent background), showing significant evolution of plasma temperature and emission measure over time. The estimated energy released in the LEIA band is ~3 * 10^39 erg, suggesting that this is likely the most energetic X-ray stellar flare with the longest duration detected to date.

astro-ph.HE↗

Science objectives of the Einstein Probe mission

The Einstein Probe (EP) is an interdisciplinary mission of time-domain and X-ray astronomy. Equipped with a wide-field lobster-eye X-ray focusing imager, EP will discover cosmic X-ray transients and monitor the X-ray variability of known sources in 0.5-4 keV, at a combination of detecting sensitivity and cadence that is not accessible to the previous and current wide-field monitoring missions. EP can perform quick characterisation of transients or outbursts with a Wolter-I X-ray telescope onboard. In this paper, the science objectives of the Einstein Probe mission are presented. EP is expected to enlarge the sample of previously known or predicted but rare types of transients with a wide range of timescales. Among them, fast extragalactic transients will be surveyed systematically in soft X-rays, which include γ-ray bursts and their variants, supernova shock breakouts, and the predicted X-ray transients associated with binary neutron star mergers. EP will detect X-ray tidal disruption events and outbursts from active galactic nuclei, possibly at an early phase of the flares for some. EP will monitor the variability and outbursts of X-rays from white dwarfs, neutron stars and black holes in our and neighbouring galaxies at flux levels fainter than those detectable by the current instruments, and is expected to discover new objects. A large sample of stellar X-ray flares will also be detected and characterised. In the era of multi-messenger astronomy, EP has the potential of detecting the possible X-ray counterparts of gravitational wave events, neutrino sources, and ultra-high energy γ-ray and cosmic ray sources. EP is expected to help advance the studies of extreme objects/phenomena and their underlying physical processes revealed in the dynamic X-ray universe, as well as studies in other areas of X-ray astronomy.

astro-ph.HE↗

Triggering the Untriggered: The First Einstein Probe-Detected Gamma-Ray Burst 240219A and Its Implications

The Einstein Probe (EP) achieved its first detection and localization of a bright X-ray flare, EP240219a, on 2024 February 19, during its commissioning phase. Subsequent targeted searches triggered by the EP240219a alert identified a faint, untriggered gamma-ray burst (GRB) in the archived data of Fermi Gamma-ray Burst Monitor (GBM), Swift Burst Alert Telescope (BAT), and Insight-HXMT/HE. The EP Wide-field X-ray Telescope (WXT) light curve reveals a long duration of approximately 160 s with a slow decay, whereas the Fermi/GBM light curve shows a total duration of approximately 70 s. The peak in the Fermi/GBM light curve occurs slightly later with respect to the peak seen in the EP/WXT light curve. Our spectral analysis shows that a single cutoff power-law (PL) model effectively describes the joint EP/WXT--Fermi/GBM spectra in general, indicating coherent broad emission typical of GRBs. The model yielded a photon index of $\sim -1.70 \pm 0.05$ and a peak energy of $\sim 257 \pm 134$ keV. After detection of GRB 240219A, long-term observations identified several candidates in optical and radio wavelengths, none of which was confirmed as the afterglow counterpart during subsequent optical and near-infrared follow-ups. The analysis of GRB 240219A classifies it as an X-ray rich GRB (XRR) with a high peak energy, presenting both challenges and opportunities for studying the physical origins of X-ray flashes, XRRs, and classical GRBs. Furthermore, linking the cutoff PL component to nonthermal synchrotron radiation suggests that the burst is driven by a Poynting flux-dominated outflow.

astro-ph.HE↗

A Joint Learning Model with Variational Interaction for Multilingual Program Translation

Programs implemented in various programming languages form the foundation of software applications. To alleviate the burden of program migration and facilitate the development of software systems, automated program translation across languages has garnered significant attention. Previous approaches primarily focus on pairwise translation paradigms, learning translation between pairs of languages using bilingual parallel data. However, parallel data is difficult to collect for some language pairs, and the distribution of program semantics across languages can shift, posing challenges for pairwise program translation. In this paper, we argue that jointly learning a unified model to translate code across multiple programming languages is superior to separately learning from bilingual parallel data. We propose Variational Interaction for Multilingual Program Translation~(VIM-PT), a disentanglement-based generative approach that jointly trains a unified model for multilingual program translation across multiple languages. VIM-PT disentangles code into language-shared and language-specific features, using variational inference and interaction information with a novel lower bound, then achieves program translation through conditional generation. VIM-PT demonstrates four advantages: 1) captures language-shared information more accurately from various implementations and improves the quality of multilingual program translation, 2) mines and leverages the capability of non-parallel data, 3) addresses the distribution shift of program semantics across languages, 4) and serves as a unified model, reducing deployment complexity.

cs.SE↗

Circumventing cracking in grading 316L stainless steel to Monel400 through compositional modifications

In joining Fe-alloys and Cu-containing alloys to access the high strength of steels and corrosion resistance of Cu-alloy, cracking is widely observed due to the significant Cu microsegregation during the solidification process, resulting in an interdendritic Cu-rich liquid film at the end of solidification. By fabricating functionally graded materials (FGMs) that incorporate additional elements like Ni in the transition region between these terminal alloy classes, the hot cracking can be reduced. In the present work, the joining of stainless steel 316L (SS316L) and Monel400 by modifying the Ni concentration in the gradient region was studied. A new hot cracking criterion based on hybrid Scheil-equilibrium approach was developed and validated with monolithic multi-layer samples within the SS316L-Ni-Monel400 three-alloy system and an SS316L to 55/45 wt% SS316L/Ni to Monel400 FGM sample fabricated by direct energy deposition (DED) process. The new hot cracking criterion, based on the hybrid Scheil-equilibrium approach, is expected to help design FGM paths between other Fe-alloys and Cu-containing alloys as well.

cond-mat.mtrl-sci↗

X-ray Sources Classification Using Machine Learning: A Study with EP-WXT Pathfinder LEIA

X-ray observations play a crucial role in time-domain astronomy. The Einstein Probe (EP), a recently launched X-ray astronomical satellite, emerges as a forefront player in the field of time-domain astronomy and high-energy astrophysics. With a focus on systematic surveys in the soft X-ray band, EP aims to discover high-energy transients and monitor variable sources in the universe. To achieve these objectives, a quick and reliable classification of observed sources is essential. In this study, we developed a machine learning classifier for autonomous source classification using data from the EP-WXT Pathfinder Lobster Eye Imager for Astronomy (LEIA) and EP-WXT simulations. The proposed Random Forest classifier, built on selected features derived from light curves, energy spectra, and location information, achieves an accuracy of approximately 95% on EP simulation data and 98% on LEIA observational data. The classifier is integrated into the LEIA data processing pipeline, serving as a tool for manual validation and rapid classification during observations. This paper presents an efficient method for the classification of X-ray sources based on single observations, along with implications of most effective features for the task. This work facilitates rapid source classification for the EP mission and also provides valuable insights into feature selection and classification techniques for enhancing the efficiency and accuracy of X-ray source classification that can be adapted to other X-ray telescope data.

astro-ph.IM↗

Large Language Models for Link Stealing Attacks Against Graph Neural Networks

Graph data contains rich node features and unique edge information, which have been applied across various domains, such as citation networks or recommendation systems. Graph Neural Networks (GNNs) are specialized for handling such data and have shown impressive performance in many applications. However, GNNs may contain of sensitive information and susceptible to privacy attacks. For example, link stealing is a type of attack in which attackers infer whether two nodes are linked or not. Previous link stealing attacks primarily relied on posterior probabilities from the target GNN model, neglecting the significance of node features. Additionally, variations in node classes across different datasets lead to different dimensions of posterior probabilities. The handling of these varying data dimensions posed a challenge in using a single model to effectively conduct link stealing attacks on different datasets. To address these challenges, we introduce Large Language Models (LLMs) to perform link stealing attacks on GNNs. LLMs can effectively integrate textual features and exhibit strong generalizability, enabling attacks to handle diverse data dimensions across various datasets. We design two distinct LLM prompts to effectively combine textual features and posterior probabilities of graph nodes. Through these designed prompts, we fine-tune the LLM to adapt to the link stealing attack task. Furthermore, we fine-tune the LLM using multiple datasets and enable the LLM to learn features from different datasets simultaneously. Experimental results show that our approach significantly enhances the performance of existing link stealing attack tasks in both white-box and black-box scenarios. Our method can execute link stealing attacks across different datasets using only a single model, making link stealing attacks more applicable to real-world scenarios.

cs.LG↗

Convergence Analysis for A Stochastic Maximum Principle Based Data Driven Feedback Control Algorithm

This paper presents convergence analysis of a novel data-driven feedback control algorithm designed for generating online controls based on partial noisy observational data. The algorithm comprises a particle filter-enabled state estimation component, estimating the controlled system's state via indirect observations, alongside an efficient stochastic maximum principle type optimal control solver. By integrating weak convergence techniques for the particle filter with convergence analysis for the stochastic maximum principle control solver, we derive a weak convergence result for the optimization procedure in search of optimal data-driven feedback control. Numerical experiments are performed to validate the theoretical findings.

math.OC↗