SearcharxivSearch

arXiv subjects

Yiqian Xu

Publications and source records attributed to Yiqian Xu.

3 recordsLinked to original sources

MAGA-Bench: Machine-Augment-Generated Text via Alignment Detection Benchmark

Machine-Generated Text (MGT) is becoming increasingly difficult to distinguish from Human-Written Text (HWT). This trend has exacerbated malicious activities such as fake news and online fraud. The generalization ability of fine-tuned detectors relies heavily on dataset quality, and simply expanding the sources of MGT may become increasingly insufficient. Further augmentation of the generation process is required. Based on HC-Var's theory, enhancing the human-like alignment of MGT not only facilitates robustness testing of existing detectors but also boosts the generalization ability of detectors fine-tuned on such aligned MGT datasets. Therefore, we propose the \textbf{M}achine-\textbf{A}ugment-\textbf{G}enerated Text via \textbf{A}lignment (MAGA) Detection Benchmark. MAGA integrates several alignment methods, ranging from prompt construction to \textbf{G}enerator-\textbf{D}etector \textbf{A}dversarial \textbf{R}einforcement \textbf{L}earning (GDARL) and the reasoning process. In our experiments, the RoBERTa detector fine-tuned on MAGA achieves an average improvement of 4.60\% in generalization AUC. Conversely, the aligned MGTs in MAGA also lead to an average decrease of 8.13\% in the AUC of selected detectors. We hope the MAGA Benchmark will provide valuable insights for future research on the generalization ability of MGT detectors.

cs.CL

MIRAGE: Exploring How Large Language Models Perform in Complex Social Interactive Environments

Large Language Models (LLMs) have shown remarkable capabilities in environmental perception, reasoning-based decision-making, and simulating complex human behaviors, particularly in interactive role-playing contexts. This paper introduces the Multiverse Interactive Role-play Ability General Evaluation (MIRAGE), a comprehensive framework designed to assess LLMs' proficiency in portraying advanced human behaviors through murder mystery games. MIRAGE features eight intricately crafted scripts encompassing diverse themes and styles, providing a rich simulation. To evaluate LLMs' performance, MIRAGE employs four distinct methods: the Trust Inclination Index (TII) to measure dynamics of trust and suspicion, the Clue Investigation Capability (CIC) to measure LLMs' capability of conducting information, the Interactivity Capability Index (ICI) to assess role-playing capabilities and the Script Compliance Index (SCI) to assess LLMs' capability of understanding and following instructions. Our experiments indicate that even popular models like GPT-4 face significant challenges in navigating the complexities presented by the MIRAGE. The datasets and simulation codes are available in \href{https://github.com/lime728/MIRAGE}{github}.

cs.CL

TeV-PeV neutrino-nucleon cross section measurement with 5 years of IceCube data

We present a novel analysis method for the determination of the neutrino-nucleon Deep Inelastic Scattering (DIS) cross section in the TeV - PeV energy range utilizing neutrino absorption by the Earth. We analyze five years of data collected with the complete IceCube detector from May 2011 to May 2016. This analysis focuses on electromagnetic and hadronic showers (cascades) mainly induced by electron and tau neutrinos. The applied event selection features high background rejection (< 10% background contamination below 60 TeV, background free above 60 TeV) of atmospheric muons and high signal efficiency (~ 80%). The final neutrino sample consists of 4808 events, with 402 events above 10 TeV reconstructed energy. An unfolding method was applied to enable the mapping from reconstructed cascade parameters such as energy and zenith to true neutrino variables. The analysis was performed assuming isotropic astrophysical neutrino flux, in seven energy bins, and in two zenith bins ("down-going" from the south-hemisphere and "up-going" from the north-hemisphere). The ratio of down-going to up-going events (which are absorbed by the Earth at high energies) is sensitive to the neutrino-nucleon cross section but insensitive to the astrophysical neutrino flux uncertainties.

hep-ex