Searcharxiv⌕ Search

arXiv subjects

Hui Guo

Publications and source records attributed to Hui Guo.

At least 37 records · Page 2Linked to original sources

Tuning Bound States of Symmetry-Breaking Vortices via Unidirectional Charge Density Wave in a Transition-Metal Dichalcogenide Superconductor

The interplay between charge density wave (CDW) and superconducting vortex bound states are crucial for fundamental physics of superconductivity and advancing quantum nanotechnologies. However, the CDW-mediated modulation of vortex bound states, which opens up a new platform for vortex engineering, remains unexplored. Here, we report spatially anisotropic vortex states modulated by the unidirectional CDWs in a transition-metal dichalcogenide superconductor 1T''-NbTe2 using ultra-low-temperature scanning tunneling microscopy/spectroscopy. The stripe-like 3x1x3 CDW order exhibits a robust three-dimensional character across step edges and coexists with superconductivity below a critical temperature of 0.4 K. Under out-of-plane magnetic fields, we observe elliptical vortices whose elongation aligns with the CDW stripes, indicating strong coupling between vortex morphology and underlying electronic order. Remarkably, CDW domain boundaries induce abrupt changes in vortex orientation and vortex bound states, enabling controllable vortex states across CDW nanodomains. These findings establish a new pathway for manipulating superconducting vortex bound states via CDW coupling.

cond-mat.supr-con↗

Incorporating flexibility and resilience demand into capacity market considering the guidance on generation investment

The capacity market provides economic guidance for generation investment and ensures the adequacy of generation capability for power systems. With the rapidly increasing proportion of renewable energy, the adequacy of flexibility and resilience becomes more crucial for the secure operation of power systems. In this context, this paper incorporates the flexibility and resilience demand into the capacity market by formulating the capacity demand curves for ramping capability, inertia and recovery capabilities besides the generation capability. The guidance on generation investment of the capacity market is also taken into account by solving the generation investment equilibrium among generation companies with a Nash Cournot model employing an equivalent quadratic programming formulation. The overall problem is established as a trilevel game and an iterative algorithm is devised to formulate the capacity demand curves in the upper level based on Genco's investment acquired from the middle and lower levels. The case study further demonstrates that to incorporate flexibility and resilience demand into the capacity market could stimulate proper generation investment and ensure the adequacy of flexibility and resilience in power systems.

eess.SY↗

LinkTo-Anime: A 2D Animation Optical Flow Dataset from 3D Model Rendering

Existing optical flow datasets focus primarily on real-world simulation or synthetic human motion, but few are tailored to Celluloid(cel) anime character motion: a domain with unique visual and motion characteristics. To bridge this gap and facilitate research in optical flow estimation and downstream tasks such as anime video generation and line drawing colorization, we introduce LinkTo-Anime, the first high-quality dataset specifically designed for cel anime character motion generated with 3D model rendering. LinkTo-Anime provides rich annotations including forward and backward optical flow, occlusion masks, and Mixamo Skeleton. The dataset comprises 395 video sequences, totally 24,230 training frames, 720 validation frames, and 4,320 test frames. Furthermore, a comprehensive benchmark is constructed with various optical flow estimation methods to analyze the shortcomings and limitations across multiple datasets.

cs.CV↗

Arcturus: A Cloud Overlay Network for Global Accelerator with Enhanced Performance and Stability

Global Accelerator (GA) services play a vital role in ensuring low-latency, high-reliability communication for real-time interactive applications. However, existing GA offerings are tightly bound to specific cloud providers, resulting in high costs, rigid deployment, and limited flexibility, especially for large-scale or budget-sensitive deployments. Arcturus is a cloud-native GA framework that revisits the design of GA systems by leveraging low-cost, heterogeneous cloud resources across multiple providers. Rather than relying on fixed, high-end infrastructure, Arcturus dynamically constructs its acceleration network and balances performance, stability, and resource efficiency. To achieve this, Arcturus introduces a two-plane design: a forwarding plane that builds a proxy network with adaptive control, and a scheduling plane that coordinates load and routing through lightweight, quantitative optimization. Evaluations under millions of RPS show that Arcturus outperforms commercial GA services by up to 1.7X in acceleration performance, reduces cost by 71%, and maintains over 80% resource efficiency--demonstrating efficient use of cloud resources at scale.

cs.NI↗

DCD: A Semantic Segmentation Model for Fetal Ultrasound Four-Chamber View

Accurate segmentation of anatomical structures in the apical four-chamber (A4C) view of fetal echocardiography is essential for early diagnosis and prenatal evaluation of congenital heart disease (CHD). However, precise segmentation remains challenging due to ultrasound artifacts, speckle noise, anatomical variability, and boundary ambiguity across different gestational stages. To reduce the workload of sonographers and enhance segmentation accuracy, we propose DCD, an advanced deep learning-based model for automatic segmentation of key anatomical structures in the fetal A4C view. Our model incorporates a Dense Atrous Spatial Pyramid Pooling (Dense ASPP) module, enabling superior multi-scale feature extraction, and a Convolutional Block Attention Module (CBAM) to enhance adaptive feature representation. By effectively capturing both local and global contextual information, DCD achieves precise and robust segmentation, contributing to improved prenatal cardiac assessment.

eess.IV↗

Numerical simulation of wormhole propagation with the mixed hybridized discontinuous Galerkin finite element method

The acid treatment of carbonate reservoirs is a widely employed technique for enhancing the productivity of oil and gas reservoirs. In this paper, we present a novel combined hybridized mixed discontinuous Galerkin (HMDG) finite element method to simulate the dissolution process near the wellbore, commonly referred to as the wormhole phenomenon. The primary contribution of this work lies in the application of hybridization techniques to both the pressure and concentration equations. Additionally, an upwind scheme is utilized to address convection-dominant scenarios, and a ``cut-off" operator is introduced to maintain the boundedness of porosity. Compared to traditional discontinuous Galerkin methods, the proposed approach results in a global system with fewer unknowns and sparser stencils, thereby significantly reducing computational costs. We analyze the existence and uniqueness of the new combined method and derive optimal error estimates using the developed technique. Numerical examples are provided to validate the theoretical analysis.

math.NA↗

LLM-MedQA: Enhancing Medical Question Answering through Case Studies in Large Language Models

Accurate and efficient question-answering systems are essential for delivering high-quality patient care in the medical field. While Large Language Models (LLMs) have made remarkable strides across various domains, they continue to face significant challenges in medical question answering, particularly in understanding domain-specific terminologies and performing complex reasoning. These limitations undermine their effectiveness in critical medical applications. To address these issues, we propose a novel approach incorporating similar case generation within a multi-agent medical question-answering (MedQA) system. Specifically, we leverage the Llama3.1:70B model, a state-of-the-art LLM, in a multi-agent architecture to enhance performance on the MedQA dataset using zero-shot learning. Our method capitalizes on the model's inherent medical knowledge and reasoning capabilities, eliminating the need for additional training data. Experimental results show substantial performance gains over existing benchmark models, with improvements of 7% in both accuracy and F1-score across various medical QA tasks. Furthermore, we examine the model's interpretability and reliability in addressing complex medical queries. This research not only offers a robust solution for medical question answering but also establishes a foundation for broader applications of LLMs in the medical domain.

cs.CL↗

A Self-Learning Multimodal Approach for Fake News Detection

The rapid growth of social media has resulted in an explosion of online news content, leading to a significant increase in the spread of misleading or false information. While machine learning techniques have been widely applied to detect fake news, the scarcity of labeled datasets remains a critical challenge. Misinformation frequently appears as paired text and images, where a news article or headline is accompanied by a related visuals. In this paper, we introduce a self-learning multimodal model for fake news classification. The model leverages contrastive learning, a robust method for feature extraction that operates without requiring labeled data, and integrates the strengths of Large Language Models (LLMs) to jointly analyze both text and image features. LLMs are excel at this task due to their ability to process diverse linguistic data drawn from extensive training corpora. Our experimental results on a public dataset demonstrate that the proposed model outperforms several state-of-the-art classification approaches, achieving over 85% accuracy, precision, recall, and F1-score. These findings highlight the model's effectiveness in tackling the challenges of multimodal fake news detection.

cs.CL↗

Learning from Noisy Labels via Conditional Distributionally Robust Optimization

While crowdsourcing has emerged as a practical solution for labeling large datasets, it presents a significant challenge in learning accurate models due to noisy labels from annotators with varying levels of expertise. Existing methods typically estimate the true label posterior, conditioned on the instance and noisy annotations, to infer true labels or adjust loss functions. These estimates, however, often overlook potential misspecification in the true label posterior, which can degrade model performances, especially in high-noise scenarios. To address this issue, we investigate learning from noisy annotations with an estimated true label posterior through the framework of conditional distributionally robust optimization (CDRO). We propose formulating the problem as minimizing the worst-case risk within a distance-based ambiguity set centered around a reference distribution. By examining the strong duality of the formulation, we derive upper bounds for the worst-case risk and develop an analytical solution for the dual robust risk for each data point. This leads to a novel robust pseudo-labeling algorithm that leverages the likelihood ratio test to construct a pseudo-empirical distribution, providing a robust reference probability distribution in CDRO. Moreover, to devise an efficient algorithm for CDRO, we derive a closed-form expression for the empirical robust risk and the optimal Lagrange multiplier of the dual problem, facilitating a principled balance between robustness and model fitting. Our experimental results on both synthetic and real-world datasets demonstrate the superiority of our method.

cs.LG↗

Understanding Mobile App Reviews to Guide Misuse Audits

Problem: We address the challenge in responsible computing where an exploitable mobile app is misused by one app user (an abuser) against another user or bystander (victim). We introduce the idea of a misuse audit of apps as a way of determining if they are exploitable without access to their implementation. Method: We leverage app reviews to identify exploitable apps and their functionalities that enable misuse. First, we build a computational model to identify alarming reviews (which report misuse). Second, using the model, we identify exploitable apps and their functionalities. Third, we validate them through manual inspection of reviews. Findings: Stories by abusers and victims mostly focus on past misuses, whereas stories by third parties mostly identify stories indicating the potential for misuse. Surprisingly, positive reviews by abusers, which exhibit language with high dominance, also reveal misuses. In total, we confirmed 156 exploitable apps facilitating the misuse. Based on our qualitative analysis, we found exploitable apps exhibiting four types of exploitable functionalities. Implications: Our method can help identify exploitable apps and their functionalities, facilitating misuse audits of a large pool of apps.

cs.CR↗

CrossDF: Improving Cross-Domain Deepfake Detection with Deep Information Decomposition

Deepfake technology poses a significant threat to security and social trust. Although existing detection methods have shown high performance in identifying forgeries within datasets that use the same deepfake techniques for both training and testing, they suffer from sharp performance degradation when faced with cross-dataset scenarios where unseen deepfake techniques are tested. To address this challenge, we propose a Deep Information Decomposition (DID) framework to enhance the performance of Cross-dataset Deepfake Detection (CrossDF). Unlike most existing deepfake detection methods, our framework prioritizes high-level semantic features over specific visual artifacts. Specifically, it adaptively decomposes facial features into deepfake-related and irrelevant information, only using the intrinsic deepfake-related information for real/fake discrimination. Moreover, it optimizes these two kinds of information to be independent with a de-correlation learning module, thereby enhancing the model's robustness against various irrelevant information changes and generalization ability to unseen forgery methods. Our extensive experimental evaluation and comparison with existing state-of-the-art detection methods validate the effectiveness and superiority of the DID framework on cross-dataset deepfake detection.

cs.CV↗

Uniqueness of positive radial solutions of Choquard type equations

In this paper, we consider the following Choquard type equation \begin{equation} \left\{\begin{aligned} &-Δu+λu=γ(Φ_N(|x|)\ast|u|^p)u \ \ \mbox{in $\mathbb{R}^N$}, \\ &\lim\limits_{|x|\to\infty}u(x)=0,\\ \end{aligned}\right. \end{equation} where $N\geq2,λ>0,γ>0, p\in[1,2]$ and $Φ_N(|x|)$ denotes the fundamental solution of the Laplacian $-Δ$ on $\mathbb{R}^N$. This equation does not have a variational frame when $p\neq 2.$ Instead of variational methods, we prove the existence and uniqueness of positive radial solutions of the above equation via the shooting method by establishing some new differential inequalities. The proofs are based on an analysis of the corresponding system of second-order differential equations, and our results extend the existing ones in the literature from $p=2$ to $p\in[1,2]$.

math.AP↗

Visualization of Unconventional Rashba Band and Vortex Zero Mode in Topopogical Superconductor Candidate AuSn$_{4}$

Topological superconductivity (TSC) is a promising platform to host Majorana zero mode (MZM) for topological quantum computing. Recently, the noble metal alloy AuSn$_{4}$ has been identified as an intrinsic surface TSC. However, the atomic visualization of its nontrivial surface states and MZM remains elusive. Here, we report the direct observation of unconventional surface states and vortex zero mode at the gold (Au) terminated surfaces of AuSn$_{4}$, by ultra-low scanning tunneling microscope/spectroscopy. Distinct from the trivial metallic bulk states at tin (Sn) surfaces, the Au terminated surface exhibits pronounced surface states near Fermi level. Our density functional theory calculations indicate that these states arise from unconventional Rashba bands, where two Fermi circles from different bands share identical helical spin textures, chiralities, and group velocities in the same direction. Furthermore, we find that although the superconducting gap, critical temperature, anisotropic in-plane critical field are almost identical on Au and Sn terminated surfaces, the in-gap bound states inside Abrikosov vortex cores show significant differences. The vortex on Sn terminated surfaces exhibits a conventional Caroli-de Gennes-Matricon bound state while the Au surface shows a sharp zero-energy core state with a long non-splitting distance, resembling an MZM in a non-quantum-limit condition. This distinction may result from the dominant contribution of unconventional Rashba bands near Fermi energy from Au terminated surface. Our results provide a new platform for studying unconventional Rashba band and MZM in superconductors.

cond-mat.supr-con↗

Existence and uniqueness of ground state solutions for the planar Schrödinger-Newton equation on the disc

This paper is concerned with the existence and qualitative properties of positive ground state solutions for the planar Schrödinger-Newton equation on the disc. First, we prove the existence and radial symmetry of all the positive ground state solutions by employing the symmetric decreasing rearrangement and Talenti's inequality. Next, we develop Newton's theorem and then use the contraction mapping principle to establish the uniqueness of the positive ground state solution for the Schrödinger-Newton equation on the disc in the two dimensional case. Finally, we show that the unique positive ground state solution converges to the trivial solution as the radius $R$ tending to infinity, which is totally different from the higher dimensional case in \cite{Guo-Wang-Yi}.

math.AP↗

Research on signalized intersection mixed traffic flow platoon control method considering Backward-looking effect

Connected and Autonomous Vehicles (CAVs) technology facilitates the advancement of intelligent transportation. However, intelligent control techniques for mixed traffic flow at signalized intersections involving both CAVs and Human-Driven Vehicles (HDVs) require further investigation into the impact of backward-looking effect. This paper proposes the concept of 1+n+1 mixed platoon considering the backward-looking effect, consisting of one leading CAV, n following HDVs, and one trailing CAV. The leading and trailing CAVs collectively guide the movement of intermediate HDVs at intersections, forming an optimal control framework for platoon-based CAVs at signalized intersections. Initially, a linearized dynamic model for the 1+n+1 mixed platoon is established and compared with a benchmark model focusing solely on controlling the lead vehicle. Subsequently, constraints are formulated for the optimal control framework, aiming to enhance overall intersection traffic efficiency and fuel economy by directly controlling the leading and trailing CAVs in the platoon. Finally, extensive numerical simulations compare vehicle throughput and fuel consumption at signalized intersections under different mixed platoon control methods, validating that considering both front and backward-looking effects in the mixed platoon control method outperforms traditional methods focusing solely on the lead CAV.

physics.app-ph↗

Uncertainty-Aware Explainable Recommendation with Large Language Models

Providing explanations within the recommendation system would boost user satisfaction and foster trust, especially by elaborating on the reasons for selecting recommended items tailored to the user. The predominant approach in this domain revolves around generating text-based explanations, with a notable emphasis on applying large language models (LLMs). However, refining LLMs for explainable recommendations proves impractical due to time constraints and computing resource limitations. As an alternative, the current approach involves training the prompt rather than the LLM. In this study, we developed a model that utilizes the ID vectors of user and item inputs as prompts for GPT-2. We employed a joint training mechanism within a multi-task learning framework to optimize both the recommendation task and explanation task. This strategy enables a more effective exploration of users' interests, improving recommendation effectiveness and user satisfaction. Through the experiments, our method achieving 1.59 DIV, 0.57 USR and 0.41 FCR on the Yelp, TripAdvisor and Amazon dataset respectively, demonstrates superior performance over four SOTA methods in terms of explainability evaluation metric. In addition, we identified that the proposed model is able to ensure stable textual quality on the three public datasets.

cs.IR↗

Object-Driven One-Shot Fine-tuning of Text-to-Image Diffusion with Prototypical Embedding

As large-scale text-to-image generation models have made remarkable progress in the field of text-to-image generation, many fine-tuning methods have been proposed. However, these models often struggle with novel objects, especially with one-shot scenarios. Our proposed method aims to address the challenges of generalizability and fidelity in an object-driven way, using only a single input image and the object-specific regions of interest. To improve generalizability and mitigate overfitting, in our paradigm, a prototypical embedding is initialized based on the object's appearance and its class, before fine-tuning the diffusion model. And during fine-tuning, we propose a class-characterizing regularization to preserve prior knowledge of object classes. To further improve fidelity, we introduce object-specific loss, which can also use to implant multiple objects. Overall, our proposed object-driven method for implanting new objects can integrate seamlessly with existing concepts as well as with high fidelity and generalization. Our method outperforms several existing works. The code will be released.

cs.CV↗

GAN-generated Faces Detection: A Survey and New Perspectives

Generative Adversarial Networks (GAN) have led to the generation of very realistic face images, which have been used in fake social media accounts and other disinformation matters that can generate profound impacts. Therefore, the corresponding GAN-face detection techniques are under active development that can examine and expose such fake faces. In this work, we aim to provide a comprehensive review of recent progress in GAN-face detection. We focus on methods that can detect face images that are generated or synthesized from GAN models. We classify the existing detection works into four categories: (1) deep learning-based, (2) physical-based, (3) physiological-based methods, and (4) evaluation and comparison against human visual performance. For each category, we summarize the key ideas and connect them with method implementations. We also discuss open problems and suggest future research directions.

cs.CV↗