Searcharxiv⌕ Search

arXiv subjects

Kui Liu

Publications and source records attributed to Kui Liu.

At least 55 records · Page 3Linked to original sources

Practical Program Repair via Preference-based Ensemble Strategy

To date, over 40 Automated Program Repair (APR) tools have been designed with varying bug-fixing strategies, which have been demonstrated to have complementary performance in terms of being effective for different bug classes. Intuitively, it should be feasible to improve the overall bug-fixing performance of APR via assembling existing tools. Unfortunately, simply invoking all available APR tools for a given bug can result in unacceptable costs on APR execution as well as on patch validation (via expensive testing). Therefore, while assembling existing tools is appealing, it requires an efficient strategy to reconcile the need to fix more bugs and the requirements for practicality. In light of this problem, we propose a Preference-based Ensemble Program Repair framework (P-EPR), which seeks to effectively rank APR tools for repairing different bugs. P-EPR is the first non-learning-based APR ensemble method that is novel in its exploitation of repair patterns as a major source of knowledge for ranking APR tools and its reliance on a dynamic update strategy that enables it to immediately exploit and benefit from newly derived repair results. Experimental results show that P-EPR outperforms existing strategies significantly both in flexibility and effectiveness.

cs.SE↗

On the distribution of $k$-free numbers on the view point of random walks

In this paper, we investigate the distribution of $k$-free numbers in a class of $α$-random walks on the integer lattice $\mathbb{Z}$. In these walks, the walker starts from a non-negative integer $r$ and moves to the right by $a$ units with probability $α$, or by $b$ units with probability $1-α$. For $k\geq 3$, we obtain the asymptotic proportion of $k$-free numbers in a path of such $α$-random walks in almost surely sense. This provides a generalization of a classical result on the distribution of $k$-free numbers in arithmetic progressions.

math.NT↗

How are We Detecting Inconsistent Method Names? An Empirical Study from Code Review Perspective

Proper naming of methods can make program code easier to understand, and thus enhance software maintainability. Yet, developers may use inconsistent names due to poor communication or a lack of familiarity with conventions within the software development lifecycle. To address this issue, much research effort has been invested into building automatic tools that can check for method name inconsistency and recommend consistent names. However, existing datasets generally do not provide precise details about why a method name was deemed improper and required to be changed. Such information can give useful hints on how to improve the recommendation of adequate method names. Accordingly, we construct a sample method-naming benchmark, ReName4J, by matching name changes with code reviews. We then present an empirical study on how state-of-the-art techniques perform in detecting or recommending consistent and inconsistent method names based on ReName4J. The main purpose of the study is to reveal a different perspective based on reviewed names rather than proposing a complete benchmark. We find that the existing techniques underperform on our review-driven benchmark, both in inconsistent checking and the recommendation. We further identify potential biases in the evaluation of existing techniques, which future research should consider thoroughly.

cs.SE↗

Towards More Realistic Evaluation for Neural Test Oracle Generation

Effective unit tests can help guard and improve software quality but require a substantial amount of time and effort to write and maintain. A unit test consists of a test prefix and a test oracle. Synthesizing test oracles, especially functional oracles, is a well-known challenging problem. Recent studies proposed to leverage neural models to generate test oracles, i.e., neural test oracle generation (NTOG), and obtained promising results. However, after a systematic inspection, we find there are some inappropriate settings in existing evaluation methods for NTOG. These settings could mislead the understanding of existing NTOG approaches' performance. We summarize them as 1) generating test prefixes from bug-fixed program versions, 2) evaluating with an unrealistic metric, and 3) lacking a straightforward baseline. In this paper, we first investigate the impacts of these settings on evaluating and understanding the performance of NTOG approaches. We find that 1) unrealistically generating test prefixes from bug-fixed program versions inflates the number of bugs found by the state-of-the-art NTOG approach TOGA by 61.8%, 2) FPR (False Positive Rate) is not a realistic evaluation metric and the Precision of TOGA is only 0.38%, and 3) a straightforward baseline NoException, which simply expects no exception should be raised, can find 61% of the bugs found by TOGA with twice the Precision. Furthermore, we introduce an additional ranking step to existing evaluation methods and propose an evaluation metric named Found@K to better measure the cost-effectiveness of NTOG approaches. We propose a novel unsupervised ranking method to instantiate this ranking step, significantly improving the cost-effectiveness of TOGA. Eventually, we propose a more realistic evaluation method TEval+ for NTOG and summarize seven rules of thumb to boost NTOG approaches into their practical usages.

cs.SE↗

Multisensor fusion-based digital twin in additive manufacturing for in-situ quality monitoring and defect correction

Early detection and correction of defects are critical in additive manufacturing (AM) to avoid build failures. In this paper, we present a multisensor fusion-based digital twin for in-situ quality monitoring and defect correction in a robotic laser direct energy deposition process. Multisensor fusion sources consist of an acoustic sensor, an infrared thermal camera, a coaxial vision camera, and a laser line scanner. The key novelty and contribution of this work are to develop a spatiotemporal data fusion method that synchronizes and registers the multisensor features within the part's 3D volume. The fused dataset can be used to predict location-specific quality using machine learning. On-the-fly identification of regions requiring material addition or removal is feasible. Robot toolpath and auto-tuned process parameters are generated for defecting correction. In contrast to traditional single-sensor-based monitoring, multisensor fusion allows for a more in-depth understanding of underlying process physics, such as pore formation and laser-material interactions. The proposed methods pave the way for self-adaptation AM with higher efficiency, less waste, and cleaner production.

eess.IV↗

Generation of large-scale continuous-variable cluster states multiplexed both in time and frequency domains

Large-scale continuous variable (CV) cluster state is necessary in quantum information processing based on measurement-based quantum computing (MBQC). Specially, generating large-scale CV cluster state multiplexed in time domain is easier to implement and has strong scalability in experiment. Here one-dimensional (1D) large-scale dual-rail CV cluster states multiplexed both in time and frequency domains are parallelly generated, which can be further extended to three-dimensional (3D) CV cluster state by combining two time-delay NOPA systems with beamsplitters. It is shown that the number of parallel arrays depends on the corresponding frequency comb lines and the partite number of each array can be very large (million), and scale of the 3D cluster state can be ultra-large. This scheme provides some special-structured CV cluster states, which will be valuable for quantum computing of hybrid domains.

quant-ph↗

App Review Driven Collaborative Bug Finding

Software development teams generally welcome any effort to expose bugs in their code base. In this work, we build on the hypothesis that mobile apps from the same category (e.g., two web browser apps) may be affected by similar bugs in their evolution process. It is therefore possible to transfer the experience of one historical app to quickly find bugs in its new counterparts. This has been referred to as collaborative bug finding in the literature. Our novelty is that we guide the bug finding process by considering that existing bugs have been hinted within app reviews. Concretely, we design the BugRMSys approach to recommend bug reports for a target app by matching historical bug reports from apps in the same category with user app reviews of the target app. We experimentally show that this approach enables us to quickly expose and report dozens of bugs for targeted apps such as Brave (web browser app). BugRMSys's implementation relies on DistilBERT to produce natural language text embeddings. Our pipeline considers similarities between bug reports and app reviews to identify relevant bugs. We then focus on the app review as well as potential reproduction steps in the historical bug report (from a same-category app) to reproduce the bugs. Overall, after applying BugRMSys to six popular apps, we were able to identify, reproduce and report 20 new bugs: among these, 9 reports have been already triaged, 6 were confirmed, and 4 have been fixed by official development teams, respectively.

cs.SE↗

The Best of Both Worlds: Combining Learned Embeddings with Engineered Features for Accurate Prediction of Correct Patches

A large body of the literature on automated program repair develops approaches where patches are automatically generated to be validated against an oracle (e.g., a test suite). Because such an oracle can be imperfect, the generated patches, although validated by the oracle, may actually be incorrect. Our empirical work investigates different representation learning approaches for code changes to derive embeddings that are amenable to similarity computations of patch correctness identification, and assess the possibility of accurate classification of correct patch by combining learned embeddings with engineered features. Experimental results demonstrate the potential of learned embeddings to empower Leopard (a patch correctness predicting framework implemented in this work) with learning algorithms in reasoning about patch correctness: a machine learning predictor with BERT transformer-based learned embeddings associated with XGBoost achieves an AUC value of about 0.803 in the prediction of patch correctness on a new dataset of 2,147 labeled patches that we collected for the experiments. Our investigations show that deep learned embeddings can lead to complementary/better performance when comparing against the state-of-the-art, PATCH-SIM, which relies on dynamic information. By combining deep learned embeddings and engineered features, Panther (the upgraded version of Leopard implemented in this work) outperforms Leopard with higher scores in terms of AUC, +Recall and -Recall, and can accurately identify more (in)correct patches that cannot be predicted by the classifiers only with learned embeddings or engineered features. Finally, we use an explainable ML technique, SHAP, to empirically interpret how the learned embeddings and engineered features are contributed to the patch correctness prediction.

cs.SE↗

Realization of higher Precision in Interferometric-weak-value-based Small-tilt Measurement

We experimentally realize a great precision enhancement in the small tilt measurement by using a Sagnac interferometer and balanced homodyne detection (BHD) of high-order optical modes, together with the weak value amplification (WVA) technique. Smaller minimum measurable tilt (MMT) and higher signal-to-noise ratio (SNR) can be obtained by using BHD, compared with the split detection (SD). The precision of 3.8 nrad can be obtained under our present experimental condition. It is shown that combining WVA technique and BHD can strengthen each other's advantages and can behave better for some special application scenarios, such as extremely weak output, wider measurement bandwidth, etc. Moreover, the precision can be further enhanced by experimental parameter optimization.

quant-ph↗

Visible lattice points in higher dimensional random walks and biases among them

For any integers $k\geq 2$, $q\geq 1$ and any finite set $\mathcal{A}=\{{\boldsymbolα}_1,\cdots,{\boldsymbolα}_q\}$, where ${ \boldsymbolα_t}=(α_{t,1},\cdots,α_{t,k})~(1\leq t\leq q)$ with $0<α_{t,1},\cdots,α_{t,k}<1$ and $α_{t,1}+\cdots+α_{t,k}=1$, this paper concerns the visibility of lattice points in the type-$\mathcal{A}$ random walk on the lattice $\mathbb{Z}^k$. We show that the proportion of visible lattice points on a random path of the walk is almost surely $1/ζ(k)$, where $ζ(s)$ is the Riemann zeta-function, and we also consider consecutive visibility of lattice points in the type-$\mathcal{A}$ random walk and give the proportion of the corresponding visible steps. Moreover, we find a new phenomenon that visible steps in both of the above cases are not evenly distributed. Our proof relies on tools from probability theory and analytic number theory.

math.NT↗

Normal-mode splitting in the optomechanical system with an optical parametric amplifier and coherent feedback

Strong coupling in optomechanical systems is the basic condition for observing many quantum phenomena such as optomechanical squeezing and entanglement. Normal-mode splitting (NMS) is the most evident signature of strong coupling systems. Here we show the NMS in the spectra of the movable mirror and the output field in an optomechanical system can be flexibly engineered by a combination of optical parametric amplifier (OPA) and coherent feedback (CF). Moreover, the NMS could be enhanced by optimizing the parameters such as input optical power, OPA gain and phase, CF strength in terms of amplitude reflectivity of beam splitter.

quant-ph↗

Is this Change the Answer to that Problem? Correlating Descriptions of Bug and Code Changes for Evaluating Patch Correctness

In this work, we propose a novel perspective to the problem of patch correctness assessment: a correct patch implements changes that "answer" to a problem posed by buggy behaviour. Concretely, we turn the patch correctness assessment into a Question Answering problem. To tackle this problem, our intuition is that natural language processing can provide the necessary representations and models for assessing the semantic correlation between a bug (question) and a patch (answer). Specifically, we consider as inputs the bug reports as well as the natural language description of the generated patches. Our approach, Quatrain, first considers state of the art commit message generation models to produce the relevant inputs associated to each generated patch. Then we leverage a neural network architecture to learn the semantic correlation between bug reports and commit messages. Experiments on a large dataset of 9135 patches generated for three bug datasets (Defects4j, Bugs.jar and Bears) show that Quatrain can achieve an AUC of 0.886 on predicting patch correctness, and recalling 93% correct patches while filtering out 62% incorrect patches. Our experimental results further demonstrate the influence of inputs quality on prediction performance. We further perform experiments to highlight that the model indeed learns the relationship between bug reports and code change descriptions for the prediction. Finally, we compare against prior work and discuss the benefits of our approach.

cs.SE↗

Predicting Patch Correctness Based on the Similarity of Failing Test Cases

Towards predicting patch correctness in APR, we propose a simple, but novel hypothesis on how the link between the patch behaviour and failing test specifications can be drawn: similar failing test cases should require similar patches. We then propose BATS, an unsupervised learning-based system to predict patch correctness by checking patch Behaviour Against failing Test Specification. BATS exploits deep representation learning models for code and patches: for a given failing test case, the yielded embedding is used to compute similarity metrics in the search for historical similar test cases in order to identify the associated applied patches, which are then used as a proxy for assessing generated patch correctness. Experimentally, we first validate our hypothesis by assessing whether ground-truth developer patches cluster together in the same way that their associated failing test cases are clustered. Then, after collecting a large dataset of 1278 plausible patches (written by developers or generated by some 32 APR tools), we use BATS to predict correctness: BATS achieves an AUC between 0.557 to 0.718 and a recall between 0.562 and 0.854 in identifying correct patches. Compared against previous work, we demonstrate that our approach outperforms state-of-the-art performance in patch correctness prediction, without the need for large labeled patch datasets in contrast with prior machine learning-based approaches. While BATS is constrained by the availability of similar test cases, we show that it can still be complementary to existing approaches: used in conjunction with a recent approach implementing supervised learning, BATS improves the overall recall in detecting correct patches. We finally show that BATS can be complementary to the state-of-the-art PATCH-SIM dynamic approach of identifying the correct patches for APR tools.

cs.SE↗

$k$-free lattice points in random walks

Let $\mathbb{Z}^2$ be the two-dimensional integer lattice. For an integer $k\geq 1$, a non-zero lattice point is $k$-free if the greatest common divisor of its coordinates is a $k$-free number. We consider the proportions of $k$-free and twin $k$-free lattice points on a path of an $α$-random walker in $\mathbb{Z}^2$. Using the second-moment method and tools from analytic number theory, we prove that these two proportions are $1/ζ(2k)$ and $\prod_{p}(1-2p^{-2k})$, respectively, where $ζ$ is the Riemann zeta function and the infinite product takes over all primes.

math.NT↗

Characterizing Sensor Leaks in Android Apps

While extremely valuable to achieve advanced functions, mobile phone sensors can be abused by attackers to implement malicious activities in Android apps, as experimentally demonstrated by many state-of-the-art studies. There is hence a strong need to regulate the usage of mobile sensors so as to keep them from being exploited by malicious attackers. However, despite the fact that various efforts have been put in achieving this, i.e., detecting privacy leaks in Android apps, we have not yet found approaches to automatically detect sensor leaks in Android apps. To fill the gap, we designed and implemented a novel prototype tool, SEEKER, that extends the famous FlowDroid tool to detect sensor-based data leaks in Android apps. SEEKER conducts sensor-focused static taint analyses directly on the Android apps' bytecode and reports not only sensor-triggered privacy leaks but also the sensor types involved in the leaks. Experimental results using over 40,000 real-world Android apps show that SEEKER is effective in detecting sensor leaks in Android apps, and malicious apps are more interested in leaking sensor data than benign apps.

cs.CR↗

Beep: Fine-grained Fix Localization by Learning to Predict Buggy Code Elements

Software Fault Localization refers to the activity of finding code elements (e.g., statements) that are related to a software failure. The state-of-the-art fault localization techniques, however, produce coarse-grained results that can deter manual debugging or mislead automated repair tools. In this work, we focus specifically on the fine-grained identification of code elements (i.e., tokens) that must be changed to fix a buggy program: we refer to it as fix localization. This paper introduces a neural network architecture (named Beep) that builds on AST paths to predict the buggy code element as well as the change action that must be applied to repair a program. Leveraging massive data of bugs and patches within the CoCoNut dataset, we trained a model that was (1) effective in localizing the buggy tokens with the Mean First Rank significantly higher than a statistics based baseline and a machine learning-based baseline, and (2) effective in predicting the repair operators (with the associated buggy code elements) with a Recall@1= 30-45% and the Mean First Rank=7-12 (evaluated by CoCoNut, ManySStuBs4J, and Defects4J datasets). To showcase how fine-grained fix localization can help program repair, we employ it in two repair pipelines where we use either a code completion engine to predict the correct token or a set of heuristics to search for the suitable donor code. A key strength of accurate fix localization for program repair is that it reduces the chance of patch overfitting, a challenge in generate-and-validate automated program repair: both two repair pipelines achieve a correctness ratio of 100%, i.e., all generated patches are found to be correct. Moreover, accurate fix localization helps enhance the efficiency of program repair.

cs.SE↗

On some sums involving the integral part function

Denote by $τ$ k (n), $ω$(n) and $μ$ 2 (n) the number of representations of n as product of k natural numbers, the number of distinct prime factors of n and the characteristic function of the square-free integers, respectively. Let [t] be the integral part of real number t. For f = $ω$, 2 $ω$ , $μ$ 2 , $τ$ k , we prove that n x f x n = x d 1 f (d) d(d + 1) + O $ε$ (x $θ$ f +$ε$) for x $\rightarrow$ $\infty$, where $θ$ $ω$ = 53 110 , $θ$ 2 $ω$ = 9 19 , $θ$ $μ$2 = 2 5 , $θ$ $τ$ k = 5k--1 10k--1 and $ε$ > 0 is an arbitrarily small positive number. These improve the corresponding results of Bordell{è}s.

math.NT↗

A variant of the prime number theorem

Let $Λ(n)$ be the von Mangoldt function, and let $[t]$ be the integral part of real number $t$. In this note, we prove that for any $\varepsilon>0$ the asymptotic formula $$ \sum_{n\le x} Λ\Big(\Big[\frac{x}{n}\Big]\Big) = x\sum_{d\ge 1} \frac{Λ(d)}{d(d+1)} + O_{\varepsilon}\big(x^{9/19+\varepsilon}\big) \qquad (x\to\infty)$$ holds. This improves a recent result of Bordellès, which requires $\frac{97}{203}$ in place of $\frac{9}{19}$.

math.NT↗