Searcharxiv⌕ Search

arXiv subjects

Trieu Hai Nguyen

Publications and source records attributed to Trieu Hai Nguyen.

4 recordsLinked to original sources

VietBinoculars: A Zero-Shot Approach for Detecting Vietnamese LLM-Generated Text

The rapid proliferation of Large Language Models has intensified the challenge of distinguishing LLM-generated text from human writing in non-English languages. This study introduces VietBinoculars, a zero-shot detection framework coupling PhoGPT-4B observer and performer models with calibrated global decision thresholds. By utilizing specialized Vietnamese BPE tokenization, the method eliminates byte-level fragmentation and probability dilution common in massive multilingual backbones. Evaluated across multi-domain benchmarks, VietBinoculars achieves an area under the ROC curve exceeding 0.99. Under optimal Youden's J thresholds and greedy decoding, detection accuracy reaches at least 98.78\%, while significantly outperforming baseline Binoculars, zero-shot detectors, and commercial tools on creative Capybara prompts. Even under a strict false positive rate constraint of 0.06\%, the detector maintains F1-scores between 83.15\% and 94.70\%. Detection performance consistently improves with sequence length, stabilizing at optimal accuracy for passages containing 450 to 550 tokens. Extended stress testing across 48 distinct model-decoding configurations and three post-generation rewriting strategies delineates practical operational boundaries. VietBinoculars exhibits robust resilience against single-pass paraphrasing and human-style revisions, but experiences notable performance degradation under high-entropy sampling and iterative double paraphrasing.

cs.CL↗

VietAIDetector: An Open-Source Zero-Shot Detector for Vietnamese AI-Generated Text

In recent years, distinguishing between AI-generated text and human-written text has remained a challenge. In this paper, we introduce VietAIDetector, an open-source tool designed specifically for detecting Vietnamese AI-generated text. It allows users to interact through a Gradio web interface with inputs ranging from raw Vietnamese text to common text file formats, including scanned documents and exceptionally long texts that exceed the context size of the employed Large Language Models (LLMs). The core component of the tool employs a Zero-Shot approach to detect AI-generated text without requiring domain-specific training data, building upon the previous VietBinoculars and Binoculars research. The tool is built upon a Vietnamese-specific language model and has been evaluated on out-of-domain datasets, demonstrating superior performance compared to existing methods primarily developed for English. Additionally, users can select optimal detection thresholds based on F1 score, accuracy, or TPR@0.05FPR requirements. The results are presented through the web interface, allowing users to easily review and verify suspicious texts or download them as a PDF report. The tool is publicly available at https://github.com/trieuntu/VietAIDetector

cs.CL↗

Clustering Vietnamese Conversations From Facebook Page To Build Training Dataset For Chatbot

The biggest challenge of building chatbots is training data. The required data must be realistic and large enough to train chatbots. We create a tool to get actual training data from Facebook messenger of a Facebook page. After text preprocessing steps, the newly obtained dataset generates FVnC and Sample dataset. We use the Retraining of BERT for Vietnamese (PhoBERT) to extract features of our text data. K-Means and DBSCAN clustering algorithms are used for clustering tasks based on output embeddings from PhoBERT$_{base}$. We apply V-measure score and Silhouette score to evaluate the performance of clustering algorithms. We also demonstrate the efficiency of PhoBERT compared to other models in feature extraction on the Sample dataset and wiki dataset. A GridSearch algorithm that combines both clustering evaluations is also proposed to find optimal parameters. Thanks to clustering such a number of conversations, we save a lot of time and effort to build data and storylines for training chatbot.

cs.CL↗

Analyze the Effects of Weighting Functions on Cost Function in the Glove Model

When dealing with the large vocabulary size and corpus size, the run-time for training Glove model is long, it can even be up to several dozen hours for data, which is approximately 500MB in size. As a result, finding and selecting the optimal parameters for the weighting function create many difficulties for weak hardware. Of course, to get the best results, we need to test benchmarks many times. In order to solve this problem, we derive a weighting function, which can save time for choosing parameters and making benchmarks. It also allows one to obtain nearly similar accuracy at the same given time without concern for experimentation.

cs.CL↗