SearcharxivSearch

arXiv subjects

Yanfei Jiang

Publications and source records attributed to Yanfei Jiang.

8 recordsLinked to original sources

PACHA: Probing AGN Coronae with High-redshift AGN

The X-ray emission of active galactic nuclei (AGN) is generally attributed to inverse Compton scattering of accretion-disk photons by hot electrons in a compact corona. In local AGN, directly constraining coronal properties is challenging because the high-energy cutoff often lies beyond the NuSTAR bandpass. High-redshift, luminous quasars enable systematic constraints on the high-energy cutoff, as cosmological redshift shifts the spectal cutoff into the observable hard X-ray band. We present first results from the ``Probing the AGN Coronae with High-redshift AGN'' (PACHA) project, based on quasi-simultaneous NuSTAR and XMM-Newton observations of 13 radio-quiet AGN at $z>1$. We constrain the high-energy cutoff and coronal temperature at 90\% confidence level for 10 and 9 sources, respectively. The sample exhibits a mean cutoff energy of $E_{\rm cut}=80.8\pm8.1$ keV and a mean coronal temperature of $kT_{\rm e}=18.4\pm1.6$ keV, both significantly lower than those measured in local {\it Swift}-BAT AGN, while the mean optical depth ($\tau=4.8\pm0.3$) is significantly higher. The uncertainties are at 1~$\sigma$. Combining our high-redshift sample with local AGN, we find a potential anti-correlation between cutoff energy and both X-ray luminosity and black hole mass, with no significant dependence on Eddington ratio. Within a hybrid coronal framework, the inferred temperatures lie well below the pair-production limits for purely thermal coronae, indicating a substantial efficient Compton cooling and/or non-thermal electron component. The detection of low coronal temperatures in high-luminosity AGN is broadly consistent with predictions from recent radiation MHD simulations that consider purely thermal electron populations, implying that non-thermal electrons may not be the primary drivers of the observed coronal properties in these systems.

astro-ph.HE

CoachLM: Automatic Instruction Revisions Improve the Data Quality in LLM Instruction Tuning

Instruction tuning is crucial for enabling Language Learning Models (LLMs) in responding to human instructions. The quality of instruction pairs used for tuning greatly affects the performance of LLMs. However, the manual creation of high-quality instruction datasets is costly, leading to the adoption of automatic generation of instruction pairs by LLMs as a popular alternative. To ensure the high quality of LLM-generated instruction datasets, several approaches have been proposed. Nevertheless, existing methods either compromise dataset integrity by filtering a large proportion of samples, or are unsuitable for industrial applications. In this paper, instead of discarding low-quality samples, we propose CoachLM, a novel approach to enhance the quality of instruction datasets through automatic revisions on samples in the dataset. CoachLM is trained from the samples revised by human experts and significantly increases the proportion of high-quality samples in the dataset from 17.7% to 78.9%. The effectiveness of CoachLM is further assessed on various real-world instruction test sets. The results show that CoachLM improves the instruction-following capabilities of the instruction-tuned LLM by an average of 29.9%, which even surpasses larger LLMs with nearly twice the number of parameters. Furthermore, CoachLM is successfully deployed in a data management system for LLMs at Huawei, resulting in an efficiency improvement of up to 20% in the cleaning of 40k real-world instruction pairs. We release various assets of CoachLM, including the training data, code and test set (https://github.com/lunyiliu/CoachLM).

cs.CL

Interpretable Online Log Analysis Using Large Language Models with Prompt Strategies

Automated log analysis is crucial in modern software-intensive systems for facilitating program comprehension throughout software maintenance and engineering life cycles. Existing methods perform tasks such as log parsing and log anomaly detection by providing a single prediction value without interpretation. However, given the increasing volume of system events, the limited interpretability of analysis results hinders analysts' comprehension of program status and their ability to take appropriate actions. Moreover, these methods require substantial in-domain training data, and their performance declines sharply (by up to 62.5%) in online scenarios involving unseen logs from new domains, a common occurrence due to rapid software updates. In this paper, we propose LogPrompt, a novel interpretable log analysis approach for online scenarios. LogPrompt employs large language models (LLMs) to perform online log analysis tasks via a suite of advanced prompt strategies tailored for log tasks, which enhances LLMs' performance by up to 380.7% compared with simple prompts. Experiments on nine publicly available evaluation datasets across two tasks demonstrate that LogPrompt, despite requiring no in-domain training, outperforms existing approaches trained on thousands of logs by up to 55.9%. We also conduct a human evaluation of LogPrompt's interpretability, with six practitioners possessing over 10 years of experience, who highly rated the generated content in terms of usefulness and readability (averagely 4.42/5). LogPrompt also exhibits remarkable compatibility with open-source and smaller-scale LLMs, making it flexible for practical deployment. Code of LogPrompt is available at https://github.com/lunyiliu/LogPrompt.

cs.SE

Knowledge-Prompted Estimator: A Novel Approach to Explainable Machine Translation Assessment

Cross-lingual Machine Translation (MT) quality estimation plays a crucial role in evaluating translation performance. GEMBA, the first MT quality assessment metric based on Large Language Models (LLMs), employs one-step prompting to achieve state-of-the-art (SOTA) in system-level MT quality estimation; however, it lacks segment-level analysis. In contrast, Chain-of-Thought (CoT) prompting outperforms one-step prompting by offering improved reasoning and explainability. In this paper, we introduce Knowledge-Prompted Estimator (KPE), a CoT prompting method that combines three one-step prompting techniques, including perplexity, token-level similarity, and sentence-level similarity. This method attains enhanced performance for segment-level estimation compared with previous deep learning models and one-step prompting approaches. Furthermore, supplementary experiments on word-level visualized alignment demonstrate that our KPE method significantly improves token alignment compared with earlier models and provides better interpretability for MT quality estimation. Code will be released upon publication.

cs.CL

Evaluating GPT's Programming Capability through CodeWars' Katas

In the burgeoning field of artificial intelligence (AI), understanding the capabilities and limitations of programming-oriented models is crucial. This paper presents a novel evaluation of the programming proficiency of Generative Pretrained Transformer (GPT) models, specifically GPT-3.5 and GPT-4, against coding problems of varying difficulty levels drawn from Codewars. The experiments reveal a distinct boundary at the 3kyu level, beyond which these GPT models struggle to provide solutions. These findings led to the proposal of a measure for coding problem complexity that incorporates both problem difficulty and the time required for solution. The research emphasizes the need for validation and creative thinking capabilities in AI models to better emulate human problem-solving techniques. Future work aims to refine this proposed complexity measure, enhance AI models with these suggested capabilities, and develop an objective measure for programming problem difficulty. The results of this research offer invaluable insights for improving AI programming capabilities and advancing the frontier of AI problem-solving abilities.

cs.AI

Future Simulations of Tidal Disruption Events

Tidal disruption events involve numerous physical processes (fluid dynamics, magnetohydrodynamics, radiation transport, self-gravity, general relativistic dynamics) in highly nonlinear ways, and, because TDEs are transients by definition, frequently in non-equilibrium states. For these reasons, numerical solution of the relevant equations can be an essential tool for studying these events. In this chapter, we present a summary of the key problems of the field for which simulations offer the greatest promise and identify the capabilities required to make progress on them. We then discuss what has been---and what cannot be---done with existing numerical methods. We close with an overview of what methods now under development may do to expand our ability to understand these events.

astro-ph.HE

Exploring the Low-Mass End of The M-Sigma Relation with Active Galaxies

We present new measurements of stellar velocity dispersions, using spectra obtained with the Keck Echellette Spectrograph and Imager (ESI) and the Magellan Echellette (MagE), for 76 Seyfert 1 galaxies from the recent catalogue of Greene & Ho. These objects were selected from the Sloan Digital Sky Survey (SDSS) to have estimated black hole (BH) masses below 2\times10^6 M\odot. Combining our results with previous ESI observations of similar objects, we obtain an expanded sample of 93 galaxies and examine the relation between BH mass and velocity dispersion (the M-Sigma relation) for active galaxies with low BH masses. The low-mass active galaxies tend to follow the extrapolation of the M-Sigma relation of inactive galaxies. Including results for active galaxies of higher BH mass from the literature, we find a zero pointα= 7.68\pm0.08 and slope of β= 3.32\pm0.22 for the M-Sigma relation [in the form log Mbh=α+βlog(σ*/200 km s-1)], with intrinsic scatter of 0.46\pm0.03 dex. This result is consistent, within the uncertainties, with the slope of the M-Sigma relation for reverberation-mapped active galaxies with BH masses from 10^6 to 10^9 M\odot. For the subset of our sample having morphological information from Hubble Space Telescope images, we examine the slope of the M-Sigma relation separately for subsamples of barred and unbarred host galaxies, and find no significant evidence for a difference in slope. We do find a mild offset between low-inclination and high-inclination disk galaxies, such that more highly inclined galaxies tend to have larger σ* at a given value of BH mass, presumably due to the contribution of disk rotation within the spectroscopic aperture. We also find that the velocity dispersion of the ionized gas trace the stellar velocity dispersion well for this large sample of low-mass Seyfert 1 galaxies.

astro-ph.CO

Star Formation in Quasar Disk

Using a version of the ZEUS code, we carry out two-dimensional simulations of self-gravitating shearing sheets, with application to QSO accretion disks at a few thousand Schwarzschild radii, corresponding to a few hundredths of a parsec for a 10^8 solar-mass black hole. Radiation pressure and optically thick radiative cooling are implemented via vertical averages. We determine dimensionless versions of the maximum surface density, accretion rate, and effective viscosity that can be sustained by density-wave turbulence without fragmentation. Where fragments do form, we study the final masses that result. The maximum Shakura-Sunyaev viscosity parameter is approximately 0.4. Fragmentation occurs when the cooling time is less than about twice the shearing time, as found by Gammie and others, but can also occur at very long cooling times in sheets that are strongly radiation-pressure dominated. For accretion at the Eddington rate onto a 10^8 solar-mass black hole, fragmentation occurs beyond four thousand Schwarzschild radii, r_s. Near this radius, initial fragment masses are several hundred suns, consistent with estimates from linear stability; final masses after merging increase with the size of the sheet, reaching several thousand suns in our largest simulations. With increasing black-hole mass at a fixed Eddington ratio, self-gravity prevails to smaller multiples of r_s, where radiation pressure is more important and the cooling time is longer compared to the dynamical time; nevertheless, fragmentation can occur and produces larger initial fragment masses. We also find energy conservation is likely to be a challenge for all eulerian codes in self-gravitating regimes where radiation pressure dominates.

astro-ph.HE