arXiv · 2509.18535
Trace Is In Sentences: Unbiased Lightweight ChatGPT-Generated Text Detector
Abstract
The widespread adoption of ChatGPT has raised concerns about its misuse, highlighting the need for robust detection of AI-generated text. Current word-level detectors are vulnerable to paraphrasing or simple prompts (PSP), suffer from biases induced by ChatGPT's word-level patterns (CWP) and training data content, degrade on modified text, and often require large models or online LLM interaction. To tackle these issues, we introduce a novel task to detect both original and PSP-modified AI-generated texts, and propose a lightweight framework that classifies texts based on their internal structure, which remains invariant under word-level changes. Our approach encodes sentence embeddings from pre-trained language models and models their relationships via attention. We employ contrastive learning to mitigate embedding biases from autoregressive generation and incorporate a causal graph with counterfactual methods to isolate structural features from topic-related biases. Experiments on two curated datasets, including abstract comparisons and revised life FAQs, validate the effectiveness of our method.
Explore related subjects
Keep this discovery
Mo Mu, Dianqiao Lei, Chang Li. 2025-09-23. Trace Is In Sentences: Unbiased Lightweight ChatGPT-Generated Text Detector. https://arxiv.org/abs/2509.18535
Cite the original work for its findings. Save a collection to share your selection of sources.