SearcharxivSearch

arXiv subjects

Zhen Xie

Publications and source records attributed to Zhen Xie.

At least 19 recordsLinked to original sources

Identification of a Large-Scale Diffuse Gamma-Ray Structure in the Southern Galactic Hemisphere

We identify and characterize a large-scale diffuse gamma-ray structure in the Southern Galactic Hemisphere using 17 yr of Fermi-LAT data. An energy-dependent likelihood analysis, including alternative Galactic diffuse-emission models, isotropic emission, the Fermi bubbles, and resolved 4FGL sources, reveals an extended excess that persists across the tested background models and spans tens of degrees. The excess broadly follows the X-ray-defined southern eROSITA Bubble (eB) region, while also overlapping the projected southern extension of Loop I. Template fits favor a filled eB-like morphology over the adopted Wolleben Loop I shell geometry, making the structure a plausible gamma-ray counterpart of the southern eB, although Loop-I-related or other localized foreground emission cannot be excluded. Under the eB template, the southern component is fainter and softer than the northern large-scale component, with an integrated luminosity lower by a factor of about seven, broadly consistent with the eROSITA-bubble asymmetry. If interpreted as Galactic-scale outflow emission, its faint, soft spectrum may indicate aged particles and/or distributed reacceleration in the outer bubble. A hadronic interpretation is energetically demanding, whereas a leptonic inverse-Compton scenario is more economical but requires rapid transport and/or local reacceleration of high-energy electrons.

astro-ph.HE

Astrophysical origins of TeV features in the cosmic-ray lepton spectrum

Precise measurements of high-energy cosmic-ray electrons and positrons have revealed spectral structures that are difficult to capture with a single smooth power-law background. The rising positron fraction measured by PAMELA and AMS-02, together with the all-electron excess reported by ATIC and the high-precision all-electron spectrum measured by DAMPE, has motivated interpretations ranging from nearby astrophysical accelerators to dark-matter annihilation or decay. In this work we revisit the conventional diffuse electron background and the possible contribution from nearby pulsars in a common propagation framework. The diffuse component is modeled with GALPROP configurations calibrated by cosmic-ray nuclei and diffuse gamma-ray observations. We then use the Green-function solution for nearby discrete sources with radiative losses to study pulsar contributions with both burst-like and continuous injection histories, including the effect of stochastic inverse-Compton cooling on the propagated spectra. We also use the highest-energy DAMPE data points as an illustrative case to compare possible local-source contributions from pulsars and supernova-remnant-like burst sources. The spectral shape of such features provides a useful diagnostic for distinguishing physically plausible nearby-source features from more exotic interpretations.

astro-ph.HE

Nonparametric Empirical Bayes Confidence Intervals

Empirical Bayes methods can improve inference on unobservable individual effects by borrowing strength across units. This paper proposes nonparametric empirical Bayes confidence intervals (NP-EBCIs) for unobservable individual effects in a normal means model. The oracle intervals are constructed from posterior quantiles under a point-identified, fully nonparametric prior; feasible intervals replace these quantiles with nonparametric estimates. The NP-EBCIs are asymptotically exact in the sense that both their conditional and marginal coverage probabilities converge to the nominal level. The flexibility of this nonparametric construction has an unavoidable statistical cost. We demonstrate that posterior quantiles, unlike posterior means, inherit the severe ill-posedness of nonparametric deconvolution: the minimax optimal estimation rate is logarithmic. This logarithmic rate is minimax optimal for errors in the conditional coverage probability, and the resulting errors in the marginal coverage probability also vanish at the same logarithmic rate. Despite these slow asymptotic rates, simulations show that the NP-EBCIs remain close to nominal coverage when the prior is non-Gaussian, and deliver substantial length reductions relative to intervals that treat each unit in isolation.

econ.EM

Unveiling axion signals in galactic supernovae with future MeV telescopes

Axion-like particles (ALPs) produced via the Primakoff process in the cores of Galactic core-collapse supernovae (SNe) could convert into MeV-energy gamma-rays through interactions with the Milky Way's magnetic field. To evaluate the detection prospects for such signals, we perform sensitivity projections for next-generation MeV telescopes by combining hypothetical instrument responses with realistic background estimates. Our analysis incorporates detailed simulations of the expected ALP flux from nearby SNe, the energy-dependent conversion probability in Galactic magnetic fields, and the telescope's angular/energy resolution based on advanced detector designs. Background components are modeled using data from current MeV missions and extrapolated to future sensitivity regimes. Our simulations demonstrate that next-generation telescopes with improved effective areas and energy resolution could achieve sensitivity to photon-ALP couplings as low as gagamma approx 1.61 x 10^-13 GeV^-1 for ALP masses ma < 10^-9 eV in Galactic Center. These results indicate that future MeV missions will probe unexplored regions of ALP parameter space, with conservative estimates suggesting they could constrain gagamma values two orders of magnitude below current astrophysical limits. Such observations would provide the most stringent tests to date for axion-like particles as a dark matter candidate in the ultra-light mass regime.

astro-ph.HE

Anatomical Prior-Driven Framework for Autonomous Robotic Cardiac Ultrasound Standard View Acquisition

Cardiac ultrasound diagnosis is critical for cardiovascular disease assessment, but acquiring standard views remains highly operator-dependent. Existing medical segmentation models often yield anatomically inconsistent results in images with poor textural differentiation between distinct feature classes, while autonomous probe adjustment methods either rely on simplistic heuristic rules or black-box learning. To address these issues, our study proposed an anatomical prior (AP)-driven framework integrating cardiac structure segmentation and autonomous probe adjustment for standard view acquisition. A YOLO-based multi-class segmentation model augmented by a spatial-relation graph (SRG) module is designed to embed AP into the feature pyramid. Quantifiable anatomical features of standard views are extracted. Their priors are fitted to Gaussian distributions to construct probabilistic APs. The probe adjustment process of robotic ultrasound scanning is formalized as a reinforcement learning (RL) problem, with the RL state built from real-time anatomical features and the reward reflecting the AP matching. Experiments validate the efficacy of the framework. The SRG-YOLOv11s improves mAP50 by 11.3% and mIoU by 6.8% on the Special Case dataset, while the RL agent achieves a 92.5% success rate in simulation and 86.7% in phantom experiments.

cs.RO

ChartM$^3$: A Multi-Stage Code-Driven Pipeline for Constructing Multi-Dimensional and Multi-Step Visual Reasoning Data in Chart Comprehension

Complex chart understanding tasks demand advanced visual recognition and reasoning capabilities from multimodal large language models (MLLMs). However, current research provides limited coverage of complex chart scenarios and computation-intensive reasoning tasks prevalent in real-world applications. This study proposes an automated multi-stage code-driven pipeline for systematically generating visual reasoning datasets to address these limitations. The pipeline integrates retrieval-augmented generation (RAG) to retrieve professional chart templates and employs chain-of-thought (CoT) strategies to generate reasoning codes that simulate real data distributions, thereby driving chart rendering and question-related statistical computations. Through model-based evaluation, the pipeline enhances chart diversity and data quality. Using this framework, we construct ChartM$^3$, a multi-dimensional and multi-step dataset containing 38K charts and 142K Q&A pairs for training, along with 2,871 high-quality evaluation samples for enabling practical performance assessment. Supervised fine-tuning (SFT) and reinforcement learning (RL) experiments demonstrate that our dataset significantly improves reasoning capabilities and cross-domain generalization performance, enabling smaller models to achieve performance comparable to larger-scale models in complex chart comprehension.

cs.CV

Prospects for Observing the Microquasar SS 433 with the LACT Array

We investigate the observational capabilities of the upcoming LACT Cherenkov telescope array for the microquasar SS 433 through detailed simulations. Our results indicate that a detection significance of 5 sigma can be achieved with approximately 30 hours of observation. This exposure, coupled with LACT's excellent angular resolution, enables the spatial separation of the eastern and western jets. Furthermore, based on the LHAASO spectral and morphological findings, the array is expected to distinguish the central hadronic component after roughly 100 hours of observation. We also examine its ability to differentiate between the H.E.S.S. and LHAASO spectral models. These findings demonstrate LACT's strong potential to provide critical insights into particle acceleration in PeVatrons and the radiation mechanisms of microquasars.

astro-ph.HE

AirCache: Activating Inter-modal Relevancy KV Cache Compression for Efficient Large Vision-Language Model Inference

Recent advancements in Large Visual Language Models (LVLMs) have gained significant attention due to their remarkable reasoning capabilities and proficiency in generalization. However, processing a large number of visual tokens and generating long-context outputs impose substantial computational overhead, leading to excessive demands for key-value (KV) cache. To address this critical bottleneck, we propose AirCache, a novel KV cache compression method aimed at accelerating LVLMs inference. This work systematically investigates the correlations between visual and textual tokens within the attention mechanisms of LVLMs. Our empirical analysis reveals considerable redundancy in cached visual tokens, wherein strategically eliminating these tokens preserves model performance while significantly accelerating context generation. Inspired by these findings, we introduce an elite observation window for assessing the importance of visual components in the KV cache, focusing on stable inter-modal relevancy modeling with enhanced multi-perspective consistency. Additionally, we develop an adaptive layer-wise budget allocation strategy that capitalizes on the strength and skewness of token importance distribution, showcasing superior efficiency compared to uniform allocation. Comprehensive evaluations across multiple LVLMs and benchmarks demonstrate that our method achieves comparable performance to the full cache while retaining only 10% of visual KV cache, thereby reducing decoding latency by 29% to 66% across various batch size and prompt length of inputs. Notably, as cache retention rates decrease, our method exhibits increasing performance advantages over existing approaches.

cs.CV

Monitoring Vibrational Evolution in Jahn-Teller Effect by Raman Imaging

The Jahn-Teller effect (JTE) reduces the geometrical symmetry of a system with degenerate electronic states via vibronic coupling, playing a pivotal role in molecular and condensed systems. In this Letter, we propose that vibrational resolved tip-enhanced Raman scattering images can visualize the vibrational evolutions in JTE in real space. Taking an experimentally viable single zinc phthalocyanine (ZnPc) molecule as a proof-of-principle example, not only the degenerate vibrational splitting but also the overlooked vibration mixing caused by the JTE in its anionic form can be straightforwardly characterized by Raman images. Leveraging Raman images, the controllable configuration of JTE distortion with partial isotopic substitution could be further identified. These findings establish a practical protocol to monitor the detailed vibrational evolutions when a single molecule experiences JTE, opening a door for visualization of spontaneous symmetry breaking in molecular and solid-state systems.

cond-mat.mes-hall

GFormer: Accelerating Large Language Models with Optimized Transformers on Gaudi Processors

Heterogeneous hardware like Gaudi processor has been developed to enhance computations, especially matrix operations for Transformer-based large language models (LLMs) for generative AI tasks. However, our analysis indicates that Transformers are not fully optimized on such emerging hardware, primarily due to inadequate optimizations in non-matrix computational kernels like Softmax and in heterogeneous resource utilization, particularly when processing long sequences. To address these issues, we propose an integrated approach (called GFormer) that merges sparse and linear attention mechanisms. GFormer aims to maximize the computational capabilities of the Gaudi processor's Matrix Multiplication Engine (MME) and Tensor Processing Cores (TPC) without compromising model quality. GFormer includes a windowed self-attention kernel and an efficient outer product kernel for causal linear attention, aiming to optimize LLM inference on Gaudi processors. Evaluation shows that GFormer significantly improves efficiency and model performance across various tasks on the Gaudi processor and outperforms state-of-the-art GPUs.

cs.AR

Layout optimization and Performance of Large Array of imaging atmospheric Cherenkov Telescope (LACT)

Large Array of imaging atmospheric Cherenkov Telescope (LACT) is an array of 32 Cherenkov telescopes with 6-meter diameter mirrors to be constructed at the LHAASO site. In this work, we present a study on the layout optimization and performance analysis of LACT. We investigate two observation modes: large zenith angle observations for ultra-high energy events and small zenith angle observations for lower energy thresholds. For large zenith angles (60°), simulations show that an 8-telescope subarray can achieve an effective area of $3 ~\rm km^2$ and excellent angular resolution. For small zenith angles, we optimize the layout of 4-telescope cells and the full 32-telescope array. The threshold of the full array is about $200~\rm GeV$, which is particularly crucial for studying transient phenomena, including gamma-ray bursts (GRBs) and active galactic nuclei (AGNs). This study provides important guidance for the final LACT layout design and performance estimates under different observational conditions, demonstrating LACT's potential for deep observations of ultra-high energy \gray sources and morphological studies of PeVatrons, as well as time-domain \gray astronomy.

astro-ph.HE

CityLLaVA: Efficient Fine-Tuning for VLMs in City Scenario

In the vast and dynamic landscape of urban settings, Traffic Safety Description and Analysis plays a pivotal role in applications ranging from insurance inspection to accident prevention. This paper introduces CityLLaVA, a novel fine-tuning framework for Visual Language Models (VLMs) designed for urban scenarios. CityLLaVA enhances model comprehension and prediction accuracy through (1) employing bounding boxes for optimal visual data preprocessing, including video best-view selection and visual prompt engineering during both training and testing phases; (2) constructing concise Question-Answer sequences and designing textual prompts to refine instruction comprehension; (3) implementing block expansion to fine-tune large VLMs efficiently; and (4) advancing prediction accuracy via a unique sequential questioning-based prediction augmentation. Demonstrating top-tier performance, our method achieved a benchmark score of 33.4308, securing the leading position on the leaderboard. The code can be found: https://github.com/alibaba/AICITY2024_Track2_AliOpenTrek_CityLLaVA

cs.CV

Stochastic Wave Dark Matter with Fermi-LAT $γ$-ray Pulsar Timing Array

Pulsar timing arrays (PTAs) can detect disturbances in the fabric of spacetime on a galactic scale by monitoring the arrival time of pulses from millisecond pulsars (MSPs). Recent advancements have enabled the use of $γ$-ray radiation emitted by MSPs, in addition to radio waves, for PTA experiments. Wave dark matter (DM), a prominent class of DM candidates, can be detected with PTAs due to its periodic perturbations of the spacetime metric. In response to this development, we perform in this Letter a first analysis of applying the $γ$-ray PTA to detect the ultralight axion-like wave DM, with the data of Fermi Large Area Telescope (Fermi-LAT). Despite its much smaller collecting area, the Fermi-LAT $γ$-ray PTA demonstrates a promising sensitivity potential. We show that the upper limits not far from those of the dedicated radio-PTA projects can be achieved. Moreover, we initiate a cross-correlation analysis using the data of two Fermi-LAT pulsars. The cross-correlation of phases, while carrying key information on the source of the spacetime perturbations, has been ignored in the existing data analyses for the wave DM detection with PTAs. Our analysis indicates that taking this information into account can improve the sensitivity to wave DM by $\gtrsim 50\%$ at masses below $10^{-23}$ eV.

astro-ph.HE

Limits on the Primordial Black Holes Dark Matter with future MeV detectors

Primordial black holes (PBHs) are a compelling candidate for Dark Matter (DM). There remain significant parameter spaces to be explored despite current astrophysical observations have set strong limits. Utilizing advanced MeV observation instruments, we have statistically established the upper limit of Hawking radiation emitted by PBHs in DM-dense systems, such as galaxy clusters or dwarf galaxies. These results can set a stringent upper limit on the ratio of PBH to DM, expressed as $f_{\rm PBH}$. Our results highlight the efficacy of MeV observations in DM-dense environments. The constraints on $f_{\rm PBH}$ for PBHs in the mass range of $10^{16}-10^{17} ~\rm g$ can be improved significantly compared with the current observations.

astro-ph.HE

Transfer Learning Across Heterogeneous Features For Efficient Tensor Program Generation

Tuning tensor program generation involves searching for various possible program transformation combinations for a given program on target hardware to optimize the tensor program execution. It is already a complex process because of the massive search space and exponential combinations of transformations make auto-tuning tensor program generation more challenging, especially when we have a heterogeneous target. In this research, we attempt to address these problems by learning the joint neural network and hardware features and transferring them to the new target hardware. We extensively study the existing state-of-the-art dataset, TenSet, perform comparative analysis on the test split strategies and propose methodologies to prune the dataset. We adopt an attention-inspired approach for tuning the tensor programs enabling them to embed neural network and hardware-specific features. Our approach could prune the dataset up to 45\% of the baseline without compromising the Pairwise Comparison Accuracy (PCA). Further, the proposed methodology can achieve on-par or improved mean inference time with 25%-40% of the baseline tuning time across different networks and target hardware.

cs.PL

DeepSpeed4Science Initiative: Enabling Large-Scale Scientific Discovery through Sophisticated AI System Technologies

In the upcoming decade, deep learning may revolutionize the natural sciences, enhancing our capacity to model and predict natural occurrences. This could herald a new era of scientific exploration, bringing significant advancements across sectors from drug development to renewable energy. To answer this call, we present DeepSpeed4Science initiative (deepspeed4science.ai) which aims to build unique capabilities through AI system technology innovations to help domain experts to unlock today's biggest science mysteries. By leveraging DeepSpeed's current technology pillars (training, inference and compression) as base technology enablers, DeepSpeed4Science will create a new set of AI system technologies tailored for accelerating scientific discoveries by addressing their unique complexity beyond the common technical approaches used for accelerating generic large language models (LLMs). In this paper, we showcase the early progress we made with DeepSpeed4Science in addressing two of the critical system challenges in structural biology research.

cs.AI

A Comprehensive Performance Study of Large Language Models on Novel AI Accelerators

Artificial intelligence (AI) methods have become critical in scientific applications to help accelerate scientific discovery. Large language models (LLMs) are being considered as a promising approach to address some of the challenging problems because of their superior generalization capabilities across domains. The effectiveness of the models and the accuracy of the applications is contingent upon their efficient execution on the underlying hardware infrastructure. Specialized AI accelerator hardware systems have recently become available for accelerating AI applications. However, the comparative performance of these AI accelerators on large language models has not been previously studied. In this paper, we systematically study LLMs on multiple AI accelerators and GPUs and evaluate their performance characteristics for these models. We evaluate these systems with (i) a micro-benchmark using a core transformer block, (ii) a GPT- 2 model, and (iii) an LLM-driven science use case, GenSLM. We present our findings and analyses of the models' performance to better understand the intrinsic capabilities of AI accelerators. Furthermore, our analysis takes into account key factors such as sequence lengths, scaling behavior, sparsity, and sensitivity to gradient accumulation steps.

cs.PF

HPC-GPT: Integrating Large Language Model for High-Performance Computing

Large Language Models (LLMs), including the LLaMA model, have exhibited their efficacy across various general-domain natural language processing (NLP) tasks. However, their performance in high-performance computing (HPC) domain tasks has been less than optimal due to the specialized expertise required to interpret the model responses. In response to this challenge, we propose HPC-GPT, a novel LLaMA-based model that has been supervised fine-tuning using generated QA (Question-Answer) instances for the HPC domain. To evaluate its effectiveness, we concentrate on two HPC tasks: managing AI models and datasets for HPC, and data race detection. By employing HPC-GPT, we demonstrate comparable performance with existing methods on both tasks, exemplifying its excellence in HPC-related scenarios. Our experiments on open-source benchmarks yield extensive results, underscoring HPC-GPT's potential to bridge the performance gap between LLMs and HPC-specific tasks. With HPC-GPT, we aim to pave the way for LLMs to excel in HPC domains, simplifying the utilization of language models in complex computing applications.

cs.DC