SearcharxivSearch

arXiv subjects

Zhang Hui

Publications and source records attributed to Zhang Hui.

3 recordsLinked to original sources

The $\beta$ Pictoris b Hill sphere transit campaign. Paper II: Searching for the signatures of the $\beta$ Pictoris exoplanets through time delay analysis of the $\delta$ Scuti pulsations

The $\beta$ Pictoris system is the closest known stellar system with directly detected gas giant planets, an edge-on circumstellar disc, and evidence of falling sublimating bodies and transiting exocomets. The inner planet, $\beta$ Pictoris c, has also been indirectly detected with radial velocity (RV) measurements. The star is a known $\delta$ Scuti pulsator, and the long-term stability of these pulsations opens up the possibility of indirectly detecting the gas giant planets through time delays of the pulsations due to a varying light travel time. We search for phase shifts in the $\delta$ Scuti pulsations consistent with the known planets $\beta$ Pictoris b and c and carry out an analysis of the stellar pulsations of $\beta$ Pictoris over a multi-year timescale. We used photometric data collected by the BRITE-Constellation, bRing, ASTEP, and TESS to derive a list of the strongest and most significant $\delta$ Scuti pulsations. We carried out an analysis with the open-source python package maelstrom to study the stability of the pulsation modes of $\beta$ Pictoris in order to determine the long-term trends in the observed pulsations. We did not detect the expected signal for $\beta$ Pictoris b or $\beta$ Pictoris c. The expected time delay is 6 seconds for $\beta$ Pictoris c and 24 seconds for $\beta$ Pictoris b. With simulations, we determined that the photometric noise in all the combined data sets cannot reach the sensitivity needed to detect the expected timing drifts. An analysis of the pulsational modes of $\beta$ Pictoris using maelstrom showed that the modes themselves drift on the timescale of a year, fundamentally limiting our ability to detect exoplanets around $\beta$ Pictoris via pulsation timing.

astro-ph.EP

Fine-tuning Language Models with Generative Adversarial Reward Modelling

Reinforcement Learning with Human Feedback (RLHF) has been demonstrated to significantly enhance the performance of large language models (LLMs) by aligning their outputs with desired human values through instruction tuning. However, RLHF is constrained by the expertise and productivity limitations of human evaluators. A response to this downside is to fall back to supervised fine-tuning (SFT) with additional carefully selected expert demonstrations. However, while this method has been proven to be effective, it invariably also leads to increased human-in-the-loop overhead. In this study, we propose another alternative approach: Reinforcement Learning with Generative Adversarial Feedback (RLGAF) to RLHF and SFT, which uses a generative adversarial training style to enable the LLMs to learn useful human expert demonstrations without being directly exposed to the training examples, thus enabling good generalization capabilities while preserving sample efficiency. Our preliminary findings indicate that RLGAF can help align LLMs outputs with competitive performance against RLHF and SFT, while not suffering from their respective inherent restrictions, suggesting promising avenues for further research on automating AI alignment.

cs.CL

Optimal Ternary Constant-Composition Codes with Weight Four and Distance Six

The sizes of optimal constant-composition codes of weight three have been determined by Chee, Ge and Ling with four cases in doubt. Group divisible codes played an important role in their constructions. In this paper, we study the problem of constructing optimal ternary constant-composition codes with Hamming weight four and minimum distance six. The problem is solved with a small number of lengths undetermined. The previously known results are those with code length no greater than 10.

cs.IT