Searcharxiv⌕ Search

arXiv subjects

Xuan Feng

Publications and source records attributed to Xuan Feng.

25 records · Page 2Linked to original sources

CHEAT: A Large-scale Dataset for Detecting ChatGPT-writtEn AbsTracts

The powerful ability of ChatGPT has caused widespread concern in the academic community. Malicious users could synthesize dummy academic content through ChatGPT, which is extremely harmful to academic rigor and originality. The need to develop ChatGPT-written content detection algorithms call for large-scale datasets. In this paper, we initially investigate the possible negative impact of ChatGPT on academia,and present a large-scale CHatGPT-writtEn AbsTract dataset (CHEAT) to support the development of detection algorithms. In particular, the ChatGPT-written abstract dataset contains 35,304 synthetic abstracts, with Generation, Polish, and Mix as prominent representatives. Based on these data, we perform a thorough analysis of the existing text synthesis detection algorithms. We show that ChatGPT-written abstracts are detectable, while the detection difficulty increases with human involvement.Our dataset is available in https://github.com/botianzhe/CHEAT.

cs.CL↗

Dual-Teacher De-biasing Distillation Framework for Multi-domain Fake News Detection

Multi-domain fake news detection aims to identify whether various news from different domains is real or fake and has become urgent and important. However, existing methods are dedicated to improving the overall performance of fake news detection, ignoring the fact that unbalanced data leads to disparate treatment for different domains, i.e., the domain bias problem. To solve this problem, we propose the Dual-Teacher De-biasing Distillation framework (DTDBD) to mitigate bias across different domains. Following the knowledge distillation methods, DTDBD adopts a teacher-student structure, where pre-trained large teachers instruct a student model. In particular, the DTDBD consists of an unbiased teacher and a clean teacher that jointly guide the student model in mitigating domain bias and maintaining performance. For the unbiased teacher, we introduce an adversarial de-biasing distillation loss to instruct the student model in learning unbiased domain knowledge. For the clean teacher, we design domain knowledge distillation loss, which effectively incentivizes the student model to focus on representing domain features while maintaining performance. Moreover, we present a momentum-based dynamic adjustment algorithm to trade off the effects of two teachers. Extensive experiments on Chinese and English datasets show that the proposed method substantially outperforms the state-of-the-art baseline methods in terms of bias metrics while guaranteeing competitive performance.

cs.CL↗

An Internet-wide Penetration Study on NAT Boxes via TCP/IP Side Channel

Network Address Translation (NAT) plays an essential role in shielding devices inside an internal local area network from direct malicious accesses from the public Internet. However, recent studies show the possibilities of penetrating NAT boxes in some specific circumstances. The penetrated NAT box can be exploited by attackers as a pivot to abuse the otherwise inaccessible internal network resources, leading to serious security consequences. In this paper, we aim to conduct an Internet-wide penetration testing on NAT boxes. The main difference between our study and the previous ones is that ours is based on the TCP/IP side channels. We explore the TCP/IP side channels in the research literature, and find that the shared-IPID side channel is the most suitable for NAT-penetration testing, as it satisfies the three requirements of our study: generality, ethics, and robustness. Based on this side channel, we develop an adaptive scanner that can accomplish the Internet-wide scanning in 5 days in a very non-aggressive manner. The evaluation shows that our scanner is effective in both the controlled network and the real network. Our measurement results reveal that more than 30,000 network middleboxes are potentially vulnerable to NAT penetration. They are distributed across 154 countries and 4,146 different organizations, showing that NAT-penetration poses a serious security threat.

cs.CR↗

Leveraging Urban Big Data for Informed Business Location Decisions: A Case Study of Starbucks in Tianhe District, Guangzhou City

With the development of the information age, cities provide a large amount of data that can be analyzed and utilized to facilitate the decision-making process. Urban big data and analytics are particularly valuable in the analysis of business location decisions, providing insight and supporting informed choices. By examining data relating to commercial locations, it becomes possible to analyze various spatial characteristics and derive the feasibility of different locations. This analytical approach contributes to effective decision-making and the formulation of robust location strategies. To illustrate this, the study focuses on Starbucks cafes in the Tianhe District of Guangzhou City, China. Utilizing data visualization maps, the spatial distribution characteristics and influencing factors of Starbucks locations are analyzed. By examining the geographical coordinates of Starbucks, main distribution characteristics are identified. Through this analysis, it explores the factors influencing the spatial layout of commercial store locations, using Starbucks as a case study. The findings offer valuable insights into the management of industrial layout and the location strategies of commercial businesses in urban environments, opening avenues for further research and development in this field.

cs.HC↗

Cell nucleus elastography with the adjoint-based inverse solver

Background and Objectives: The mechanics of the nucleus depends on cellular structures and architecture, and impact a number of diseases. Nuclear mechanics is yet rather complex due to heterogeneous distribution of dense heterochromatin and loose euchromatin domains, giving rise to spatially variable stiffness properties. Methods: In this study, we propose to use the adjoint-based inverse solver to identify for the first time the nonhomogeneous elastic property distribution of the nucleus. Inputs of the inverse solver are deformation fields measured with microscopic imaging in contracting cardiomyocytes. Results: The feasibility of the proposed method is first demonstrated using simulated data. Results indicate accurate identification of the assumed heterochromatin region, with a maximum relative error of less than 5%. We also investigate the influence of unknown Poisson's ratio on the reconstruction and find that variations of the Poisson's ratio in the range [0.3-0.5] result in uncertainties of less than 15% in the identified stiffness. Finally, we apply the inverse solver on actual deformation fields acquired within the nuclei of two cardiomyocytes. The obtained results are in good agreement with the density maps obtained from microscopy images. Conclusions: Overall, the proposed approach shows great potential for nuclear elastography, with promising value for emerging fields of mechanobiology and mechanogenetics.

physics.med-ph↗

AggPose: Deep Aggregation Vision Transformer for Infant Pose Estimation

Movement and pose assessment of newborns lets experienced pediatricians predict neurodevelopmental disorders, allowing early intervention for related diseases. However, most of the newest AI approaches for human pose estimation methods focus on adults, lacking publicly benchmark for infant pose estimation. In this paper, we fill this gap by proposing infant pose dataset and Deep Aggregation Vision Transformer for human pose estimation, which introduces a fast trained full transformer framework without using convolution operations to extract features in the early stages. It generalizes Transformer + MLP to high-resolution deep layer aggregation within feature maps, thus enabling information fusion between different vision levels. We pre-train AggPose on COCO pose dataset and apply it on our newly released large-scale infant pose estimation dataset. The results show that AggPose could effectively learn the multi-scale features among different resolutions and significantly improve the performance of infant pose estimation. We show that AggPose outperforms hybrid model HRFormer and TokenPose in the infant pose estimation dataset. Moreover, our AggPose outperforms HRFormer by 0.8 AP on COCO val pose estimation on average. Our code is available at github.com/SZAR-LAB/AggPose.

cs.CV↗

Understanding and Mitigating the Security Risks of Voice-Controlled Third-Party Skills on Amazon Alexa and Google Home

Virtual personal assistants (VPA) (e.g., Amazon Alexa and Google Assistant) today mostly rely on the voice channel to communicate with their users, which however is known to be vulnerable, lacking proper authentication. The rapid growth of VPA skill markets opens a new attack avenue, potentially allowing a remote adversary to publish attack skills to attack a large number of VPA users through popular IoT devices such as Amazon Echo and Google Home. In this paper, we report a study that concludes such remote, large-scale attacks are indeed realistic. More specifically, we implemented two new attacks: voice squatting in which the adversary exploits the way a skill is invoked (e.g., "open capital one"), using a malicious skill with similarly pronounced name (e.g., "capital won") or paraphrased name (e.g., "capital one please") to hijack the voice command meant for a different skill, and voice masquerading in which a malicious skill impersonates the VPA service or a legitimate skill to steal the user's data or eavesdrop on her conversations. These attacks aim at the way VPAs work or the user's mis-conceptions about their functionalities, and are found to pose a realistic threat by our experiments (including user studies and real-world deployments) on Amazon Echo and Google Home. The significance of our findings have already been acknowledged by Amazon and Google, and further evidenced by the risky skills discovered on Alexa and Google markets by the new detection systems we built. We further developed techniques for automatic detection of these attacks, which already capture real-world skills likely to pose such threats.

cs.CR↗