Searcharxiv⌕ Search

arXiv subjects

Yiheng Xiong

Publications and source records attributed to Yiheng Xiong.

15 recordsLinked to original sources

An Axial $U_A(1)_{L_μ-L_τ}$: UV Completion and Experimental Searches

We propose an anomaly-free and renormalizable axial $U_A(1)_{L_μ-L_τ}$ model and study its experimental signatures for $A'$ masses from the MeV scale to the TeV scale. The opposite charges of the left- and right-handed charged leptons forbid the usual muon and tau Yukawa interactions. Their masses are instead generated by a singlet scalar and heavy vector-like leptons through a universal-seesaw mechanism. We focus on heavy vector-like leptons, small light--heavy mixing, and $m_s\gtrsim10~\mathrm{GeV}$. In this limit, the observables considered here depend mainly on $(m_{A'},g_X)$, while the other model parameters are restricted by mixing and perturbativity. We confront this benchmark with current experimental searches. For neutrino trident production, our finite-$m_μ$ calculation shows that $A'$ modifies the axial weak coefficient, rather than the vector coefficient relevant to the usual $L_μ-L_τ$ model. The longitudinal mode enhances muon bremsstrahlung and gives a negative contribution to $(g-2)_μ$; the latter dominates over the scalar contribution in our benchmark. Combining these results with invisible meson decays and four-muon resonance searches, we summarize the phenomenological constraints in the $(m_{A'},g_X)$ plane. For $m_{A'}\gg m_μ$, vector and axial final-state-radiation rates become nearly identical, so the corresponding collider limits can be obtained by rate matching. At a future muon collider, the total rate alone does not fully resolve the interaction structure, whereas angular distributions, especially the forward--backward asymmetry in $μ^+μ^-\toτ^+τ^-$, retain direct sensitivity to chirality.

hep-ph↗

Towards Practical Algorithm Selection for Unsupervised Domain Adaptation in Medical Imaging

Numerous unsupervised domain adaptation (UDA) algorithms exist, but for clinical practice, selecting the best-suited one along with proper hyperparameters often remains unclear, as the unlabeled deployment (target) domain prevents direct evaluation. We propose a label-free criterion that jointly selects the algorithm and hyperparameters for UDA. Given a pool of candidate models from multiple algorithms trained with different hyperparameters, our approach scores each candidate against an agreement reference, and selects the one with the highest score. The agreement reference is constructed in two levels without using target labels. First, we leverage multiple label-free selection signals, using each to nominate a model within every algorithm. Second, the nominated models are aggregated across algorithms to form a reference prediction for each unlabeled target sample. The candidate whose predictions agree most with this reference is then selected for deployment. Experimental results on four brain MRI and four chest X-ray datasets across seven clinically relevant transfer scenarios show that our method achieves better selection performance than other methods and remains effective across different algorithm pools. Our approach takes a step towards practical, label-free algorithm selection for clinical deployment of UDA.

cs.CV↗

How Far from Clinical Deployment? Evaluating the Complete Unsupervised Domain Adaptation Pipeline in Medical Imaging

Deploying unsupervised domain adaptation (UDA) in clinical practice requires choosing which algorithm to use and which of its trained models to ship. However, the deployment (target) domain is unlabeled, so models cannot be evaluated directly on it, leaving it unclear which to select. We address this by evaluating the complete UDA pipeline, considering both adaptation and label-free selection together. Our study covers eleven clinically relevant cross-domain scenarios from nine medical imaging datasets, with ten UDA algorithms and 13 label-free selection methods (validators), evaluating over 80,000 trained models in total. By this, we find that a capable adapted model usually exists, but identifying it without target labels is difficult: the validator-selected models leave a large and structural target performance gap to the best available one, with no evaluated validator consistently reliable. Towards closing it, we explore two strategies, ensembling and a small target-labeling budget; both narrow this gap but do not close it entirely. Overall, deployable UDA depends on the complete pipeline; addressing the less explored selection step could bring much of current UDA closer to clinical use.

cs.CV↗

From Signals to Behaviors: Evidence-Based Android Malware Detection

Android malware remains a persistent threat, and detecting it accurately is a long-standing open problem. Whether an app is malicious depends on what it actually does and the context in which it does it, not on the surface signals it happens to exhibit. Existing detectors instead reason about proxies for behavior, such as learned features or local code slices, and flag whatever deviates from these proxies as malicious. But deviation is not maliciousness: benign apps that merely look unusual are over-flagged, evolving malware that looks ordinary slips through. We argue that detection should be behavior-oriented: recover an app's potentially malicious behaviors and judge which are truly malicious. To realize this, we present Praxis, which structures detection as a hypothesize-confirm-judge pipeline: it hypothesizes candidate behaviors from coarse static signals, confirms each by grounding it in code evidence verified with program analysis, and judges the confirmed behaviors in context: the user's awareness, the app's functional context, and how they compose into an attack. For a malicious app, Praxis returns a verdict and the supported behaviors. We evaluate Praxis against seven baselines across three challenging settings. It achieves the best overall detection performance (87.4% F1), outperforming the baselines by 18.6-34.8 percentage points. On high-permission benign apps, it reduces the false-positive rate to 13.0%, a reduction of 41.1-67.0 percentage points compared with the baselines. Beyond binary detection, Praxis recovers fine-grained malicious behaviors at 87.3% F1, outperforming prior behavior-level approaches by 56.5-73.4 percentage points. Ablation studies show that each stage of the pipeline contributes to the final performance.

cs.CR↗

MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Content

Mobile graphical user interface (GUI) agents driven by vision-language models (VLMs) perceive the screen as rendered pixels and choose actions from what they see, so they cannot reliably separate trusted interface elements from user-generated content. We present MIRAGE (Mobile Injection of Realistic Adversarial GUI Examples), a pipeline that turns benign mobile screenshots into prompt-injection samples by placing attacker-controlled text into ordinary user-generated content regions, without modifying the agent, the application, or the operating system. MIRAGE operates in three stages: a Localizer identifies user-controllable regions on the screenshot, a Generator synthesises context-aware payloads and renders them in the application's native style, and a Curator moderates realism and balances the samples across applications, region types, and attack intents. A key challenge is that an injected screenshot must stay visually indistinguishable from genuine user content while still diverting the agent; we address this by separating the stages that control reach, realism, and distributional balance. On a 1,111-sample benchmark spanning ten applications and eleven attack intents, all five evaluated VLM agents are vulnerable, with attack success rates of 23%-30%, and MIRAGE scores higher on human realism ratings than the strongest prior attack (3.02 versus 2.52 out of 5). We further find that per-sample realism and attack success are uncorrelated, so visual-quality filtering alone cannot reliably defend against this threat.

cs.CR↗

From Exploration to Specification: LLM-Based Property Generation for Mobile App Testing

Mobile apps often suffer from functional bugs that do not cause crashes but instead manifest as incorrect behaviors under specific user interactions. Such bugs are difficult to detect automatically because they often lack explicit test oracles. Property-based testing can effectively expose them by checking intended behavioral properties under diverse interactions. However, its use largely depends on manually written properties, whose construction is difficult and expensive, limiting its practical use for mobile apps. To address this limitation, we propose PropGen, an automated approach for generating properties for Android apps. However, this task is challenging for two reasons: app functionalities are often hard to systematically uncover and execute, and properties are difficult to derive accurately from observed behaviors. To this end, PropGen performs functionality-guided exploration to collect behavioral evidence from app executions, synthesizes properties from the collected evidence, and refines imprecise properties based on testing feedback. We implemented PropGen and evaluated it on 12 real-world Android apps. The results show that PropGen can effectively identify and execute valid app functionalities, generate valid properties, and repair most imprecise ones. Across all apps, PropGen identified 1,210 valid functionalities and correctly executed 977 of them, compared with 491 and 187 for the baseline. It generated 985 properties, 912 of which were valid, and repaired 118 of 127 imprecise ones exposed during testing. With the resulting properties, we found 25 previously unknown functional bugs in the latest versions of the subject apps, many of which were missed by existing functional testing techniques.

cs.SE↗

Improving Random Testing via LLM-powered UI Tarpit Escaping for Mobile Apps

Random GUI testing is a widely-used technique for testing mobile apps. However, its effectiveness is limited by the notorious issue -- UI exploration tarpits, where the exploration is trapped in local UI regions, thus impeding test coverage and bug discovery. In this experience paper, we introduce LLM-powered random GUI Testing, a novel hybrid testing approach to mitigating UI tarpits during random testing. Our approach monitors UI similarity to identify tarpits and query LLMs to suggest promising events for escaping the encountered tarpits. We implement our approach on top of two different automated input generation (AIG) tools for mobile apps: (1) HybridMonkey upon Monkey, a state-of-the-practice tool; and (2) HybridDroidbot upon Droidbot, a state-of-the-art tool. We evaluated them on 12 popular, real-world apps. The results show that HybridMonkey and HybridDroidbot outperform all baselines, achieving average coverage improvements of 54.8% and 44.8%, respectively, and detecting the highest number of unique crashes. In total, we found 75 unique bugs, including 34 previously unknown bugs. To date, 26 bugs have been confirmed and fixed. We also applied HybridMonkey on WeChat, a popular industrial app with billions of monthly active users. HybridMonkey achieved higher activity coverage and found more bugs than random testing.

cs.SE↗

From Natural Language to Executable Properties for Property-based Testing of Mobile Apps

Property-based testing (PBT) is a popular software testing methodology and is effective in validating the functionality of mobile applications (apps for short). However, its adoption in practice remains limited, largely due to the manual effort and technical expertise required to specify executable properties. In this experience paper, we propose a novel structured property synthesis approach that automatically translates property descriptions in natural language into executable properties, and implement it in a tool named iPBT. Our approach decomposes the problem into UI semantic grounding and executable property synthesis. It first builds an enriched widget context via multimodal LLMs to align visual elements with their functional semantics, and then uses an LLM with in-context learning to generate framework-specific executable properties. We evaluate iPBT with a closed-source LLM (GPT-4o) and an open-source LLM (DeepSeek-V3) on 124 diverse property descriptions derived from an existing benchmark dataset. iPBT achieves 95.2% (118/124) accuracy on both LLMs. Notably, an ablation study reveals that the enriched widget context contributes to an absolute improvement of up to 20.2% (from 75.0% to 95.2%). A user study with 10 participants demonstrates that iPBT reduces the time required to write executable properties by 56%, suggesting substantially lower manual effort. Furthermore, evaluations on 1,180 linguistically diverse variations demonstrate iPBT's robustness (87.6% accuracy), indicating its capability to handle varied expressions.

cs.SE↗

FunnyNodules: A Customizable Medical Dataset Tailored for Evaluating Explainable AI

Densely annotated medical image datasets that capture not only diagnostic labels but also the underlying reasoning behind these diagnoses are scarce. Such reasoning-related annotations are essential for developing and evaluating explainable AI (xAI) models that reason similarly to radiologists: making correct predictions for the right reasons. To address this gap, we introduce FunnyNodules, a fully parameterized synthetic dataset designed for systematic analysis of attribute-based reasoning in medical AI models. The dataset generates abstract, lung nodule-like shapes with controllable visual attributes such as roundness, margin sharpness, and spiculation. The target class is derived from a predefined attribute combination, allowing full control over the decision rule that links attributes to the diagnostic class. We demonstrate how FunnyNodules can be used in model-agnostic evaluations to assess whether models learn correct attribute-target relations, to interpret over- or underperformance in attribute prediction, and to analyze attention alignment with attribute-specific regions of interest. The framework is fully customizable, supporting variations in dataset complexity, target definitions, class balance, and beyond. With complete ground truth information, FunnyNodules provides a versatile foundation for developing, benchmarking, and conducting in-depth analyses of explainable AI methods in medical image analysis.

cs.CV↗

Discovery prospects for photophobic axion-like particles at a 100 TeV proton--proton collider

We study heavy photophobic axion-like particles (ALPs) in the limit of an effectively vanishing diphoton coupling, $g_{aγγ}\simeq 0$, for which diphoton production and decay are suppressed and collider phenomenology is driven by electroweak interactions ($aWW$, $aZγ$, $aZZ$). We perform detector-level searches at a future $\sqrt{s}=$ 100 TeV $pp$ collider (SppC/FCC-hh), with an integrated luminosity of $\mathcal{L} =$ 20 ab$^{-1}$. We consider $a\to Zγ$ and $a\to W^+W^-$ decays. For $pp\to jj\,a$ we include both $s$-channel electroweak exchange and vector boson fusion (VBF)-like topologies, while the tri-$W$ signature arises from associated production $pp\to W^\pm a$ (via $s$-channel exchange) followed by $a\to W^+W^-$. We analyze three final states--$Zγjj$ with $Z\to\ell^+\ell^-$, tri-$W$ ($W^\pm W^\pm W^\mp$) with same-sign dimuons plus jets, and $W^+W^-jj$ with opposite-sign, different-flavor dilepton ($e^\pmμ^\mp$) plus jets. Among the two $WW$ final states, the VBF-assisted $jj\,a(\to W^+W^-)$ channel overtakes the purely $s$-channel tri-$W$ mode for $m_a \stackrel{>}{\sim}$ 1 TeV, reflecting 100~TeV signal/background kinematic shifts beyond naive energy/luminosity rescaling. A boosted-decision-tree (BDT) classifier built from kinematic observables provides the final signal--background separation, using detector-level simulations of signal and high-statistics SM backgrounds. At $\sqrt{s}=100$ TeV and $\mathcal{L} =$ 20 ab$^{-1}$, we present discovery sensitivities to the ALP--$W$ coupling $g_{aWW}$ over $m_a\in[100,\,7000]$ GeV. In parallel, we report model-independent discovery thresholds on $σ\times\mathrm{Br}$ for $pp\to jj\,a$ with $a\to Zγ$ and $a\to W^+W^-$, as well as for associated production $pp\to W^\pm a$ with $a\to W^+W^-$....

hep-ph↗

Enhancing Automated Program Repair via Faulty Token Localization and Quality-Aware Patch Refinement

Large language models (LLMs) have recently demonstrated strong potential for automated program repair (APR). However, existing LLM-based techniques primarily rely on coarse-grained external feedback (e.g.,test results) to guide iterative patch generation, while lacking fine-grained internal signals that reveal why a patch fails or which parts of the generated code are likely incorrect. This limitation often leads to inefficient refinement, error propagation, and suboptimal repair performance. In this work, we propose TokenRepair, a novel two-level refinement framework that enhances APR by integrating internal reflection for localizing potentially faulty tokens with external feedback for quality-aware patch refinement. Specifically, TokenRepair first performs internal reflection by analyzing context-aware token-level uncertainty fluctuations to identify suspicious or low-confidence tokens within a patch. It then applies Chain-of-Thought guided rewriting to refine only these localized tokens, enabling targeted and fine-grained correction. To further stabilize the iterative repair loop, TokenRepair incorporates a quality-aware external feedback mechanism that evaluates patch quality and filters out low-quality candidates before refinement. Experimental results show that TokenRepair achieves new state-of-the-art repair performance, correctly fixing 88 bugs on Defects4J 1.2 and 139 bugs on HumanEval-Java, demonstrating substantial improvements ranging from 8.2% to 34.9% across all models on Defects4J 1.2 and from 3.3% to 16.1% on HumanEval-Java.

cs.SE↗

Sensitivities to New Resonance Couplings to $W$-Bosons at the LHC

We propose a search strategy at the HL-LHC for a new neutral particle $X$ that couples to $W$-bosons, using the process $p p \rightarrow W^{\pm} X (\rightarrow W^{+} W^{-})$ with a tri-$W$-boson final state. Focusing on events with two same-sign leptonic $W$-boson decays into muons and a hadronically decaying $W$-boson, our method leverages the enhanced signal-to-background discrimination achieved through a machine-learning-based multivariate analysis. Using the heavy photophobic axion-like particle (ALP) as a benchmark, we evaluate the discovery sensitivities on both production cross section times branching ratio $σ(p p \rightarrow W^{\pm} X) \times \textrm{Br}(X \rightarrow W^{+} W^{-})$ and the coupling $g_{aWW}$ for the particle mass over a wide range of 170-3000 GeV at the HL-LHC with center-of-mass energy $\sqrt{s} = 14$ TeV and integrated luminosity $\mathcal{L} = 3$ $\textrm{ab}^{-1}$. Our results show significant improvements in discovery sensitivity, particularly for masses above 300 GeV, compared to existing limits derived from CMS analyses of Standard Model (SM) tri-$W$-boson production at $\sqrt{s} = 13$ TeV. This study demonstrates the potential of advanced selection techniques in probing the coupling of new particles to $W$-bosons and highlights the HL-LHC's capability to explore the physics beyond the SM.

hep-ph↗

PT43D: A Probabilistic Transformer for Generating 3D Shapes from Single Highly-Ambiguous RGB Images

Generating 3D shapes from single RGB images is essential in various applications such as robotics. Current approaches typically target images containing clear and complete visual descriptions of the object, without considering common realistic cases where observations of objects that are largely occluded or truncated. We thus propose a transformer-based autoregressive model to generate the probabilistic distribution of 3D shapes conditioned on an RGB image containing potentially highly ambiguous observations of the object. To handle realistic scenarios such as occlusion or field-of-view truncation, we create simulated image-to-shape training pairs that enable improved fine-tuning for real-world scenarios. We then adopt cross-attention to effectively identify the most relevant region of interest from the input image for shape generation. This enables inference of sampled shapes with reasonable diversity and strong alignment with the input image. We train and test our model on our synthetic data then fine-tune and test it on real-world data. Experiments demonstrate that our model outperforms state of the art in both scenarios.

cs.CV↗

Prior-RadGraphFormer: A Prior-Knowledge-Enhanced Transformer for Generating Radiology Graphs from X-Rays

The extraction of structured clinical information from free-text radiology reports in the form of radiology graphs has been demonstrated to be a valuable approach for evaluating the clinical correctness of report-generation methods. However, the direct generation of radiology graphs from chest X-ray (CXR) images has not been attempted. To address this gap, we propose a novel approach called Prior-RadGraphFormer that utilizes a transformer model with prior knowledge in the form of a probabilistic knowledge graph (PKG) to generate radiology graphs directly from CXR images. The PKG models the statistical relationship between radiology entities, including anatomical structures and medical observations. This additional contextual information enhances the accuracy of entity and relation extraction. The generated radiology graphs can be applied to various downstream tasks, such as free-text or structured reports generation and multi-label classification of pathologies. Our approach represents a promising method for generating radiology graphs directly from CXR images, and has significant potential for improving medical image analysis and clinical decision-making.

cs.CV↗

Fully Automated Functional Fuzzing of Android Apps for Detecting Non-crashing Logic Bugs

Android apps are GUI-based event-driven software and have become ubiquitous in recent years. Obviously, functional correctness is critical for an app's success. However, in addition to crash bugs, non-crashing functional bugs (in short as "non-crashing bugs" in this work) like inadvertent function failures, silent user data lost and incorrect display information are prevalent, even in popular, well-tested apps. These non-crashing functional bugs are usually caused by program logic errors and manifest themselves on the graphic user interfaces (GUIs). In practice, such bugs pose significant challenges in effectively detecting them because (1) current practices heavily rely on expensive, small-scale manual validation (the lack of automation); and (2) modern fully automated testing has been limited to crash bugs (the lack of test oracles). This paper fills this gap by introducing independent view fuzzing, a novel, fully automated approach for detecting non-crashing functional bugs in Android apps. Inspired by metamorphic testing, our key insight is to leverage the commonly-held independent view property of Android apps to manufacture property-preserving mutant tests from a set of seed tests that validate certain app properties. The mutated tests help exercise the tested apps under additional, adverse conditions. Any property violations indicate likely functional bugs for further manual confirmation. We have realized our approach as an automated, end-to-end functional fuzzing tool, Genie. Given an app, (1) Genie automatically detects non-crashing bugs without requiring human-provided tests and oracles (thus fully automated); and (2) the detected non-crashing bugs are diverse (thus general and not limited to specific functional properties), which set Genie apart from prior work.

cs.SE↗