SearcharxivSearch

arXiv subjects

Xiaoting Zhang

Publications and source records attributed to Xiaoting Zhang.

At least 19 recordsLinked to original sources

SmellCC: A Tool for Automated Code Smells Remediation

Code smells significantly threaten software maintainability by accumulating technical debt, yet developers often lack the resources to manually address these flaws under tight release schedules. While static analysis tools like SonarQube provide precise detection, they function largely as passive alert systems, leaving the burden of refactoring on developers. To bridge this gap, we present a novel cleaning tool, namely SmellCC, a Visual Studio Code extension that augments SonarQube with an LLM-based pipeline to automatically detect and refactor Python code smells. By employing Chain-of-Thought (CoT) and few-shot learning, SmellCC provides in-place, one-click remediation for the top-10 most frequent smells, effectively preventing the accumulation of technical debt during development. Our quantitative evaluation demonstrates that our SmellCC is promising in helping developers effectively eliminate code smells (96.8\% cleaning rate) with high accuracy (i.e., 91.3\%), ensuring that the refactored code remains syntactically correct and behavior-preserving, thereby significantly improving long-term software maintainability.

cs.SE

On the Shoulders of Giants: Empowering Automated Smart Contract Auditing via the GiAnt Corpus

High-quality smart contract auditing datasets are crucial for evaluating security tools and advancing smart contract security research. Two major limitations of existing datasets are the manual-induced scalability bottleneck and the deficiency in data granularity and diversity. To address these limitations, we propose GiANT, an automated framework designed to curate smart contract auditing datasets by distilling vulnerability insights from real-world auditing reports. GiANT employs a divide-and-conquer strategy coupled with the Chain-of-Thought technique to extract structured vulnerability information from Code4rena reports, followed by an LLM-as-a-judge mechanism to perform rigorous quality assurance. To evaluate GiANT's effectiveness, we run it on 388 real-world audit reports and generate the GiAnt Corpus comprising 7,711 vulnerability findings across five severity levels. Manual assessment of the dataset demonstrates exceptional reliability in information extraction, achieving a mean quality score of $4.76\pm0.37$ (out of 5) with inter-rater agreement $\kappa$ of 0.88. We further validate the practicality of our dataset by benchmarking 4 state-of-the-art LLMs on vulnerability detection, code summarization, mitigation recommendation, and automated gas optimization tasks, to establish performance baselines, thereby providing a valuable data foundation for future research in automated smart contract auditing.

cs.CR

Clean Code, Better Models: Enhancing LLM Performance with Smell-Cleaned Dataset

The Large Language Models (LLMs) have demonstrated great potential in code-related tasks. However, most research focuses on improving the output quality of LLMs (e.g., correctness), and less attention has been paid to the LLM input (e.g., the training code quality). Given that code smells are widely existed in practice and can negatively impact software maintainability and readability, this study takes the first systematic research to assess and improve dataset quality in terms of code smells. In this work, we first conduct a preliminary study to explore the presence of code smells in a popular benchmark dataset (i.e., CodeSearchNet-Python}) and evaluate the output of several popular LLMs (i.e., DeepSeek-Coder, CodeLlama, and MagiCoder), revealing that code smell issues extensively exist in LLM's input (e.g., benchmark dataset) and output (e.g., generated code). We then conduct our systematic research by taking three main steps: Firstly, we propose an LLM-based code smell cleaning tool, named SmellCC, which automatically refactors and removes code smells. To evaluate the correctness of the code refactoring, we construct a test set of 50 repositories sourced from the CodeSearchNet-Python benchmark for functional testing. Then we apply our curated smell-cleaned dataset to fine-tune two LLMs (i.e., DeepSeek-V2 and Qwen-Coder) to explore their potential for generating high-quality code. Thirdly, we investigate the impact of code smells on two downstream tasks: code completion and code search. Lastly, we derive several actionable implications for software engineering researchers and industry practitioners from our findings.

cs.SE

Contractibility and total semi-stability conditions of Euclidean quivers

We study the bounded derived category $\mathcal{D}$ of an Euclidean quiver, or equivalently, that of coherent sheaves on a tame weighted projective line. We give a description of the moduli space $\mathrm{ToSS}$ of the total semi-stability conditions on $\mathcal{D}$, which implies that $\mathrm{ToSS}$ can linearly contract to any chosen non-concentrated stability condition in it. For type $\widetilde{A_{p,q}}$, this gives an alternative proof of the contractibility of the whole space of stability conditions.

math.RT

SemDP: Semantic-level Differential Privacy Protection for Face Datasets

While large-scale face datasets have advanced deep learning-based face analysis, they also raise privacy concerns due to the sensitive personal information they contain. Recent schemes have implemented differential privacy to protect face datasets. However, these schemes generally treat each image as a separate database, which does not fully meet the core requirements of differential privacy. In this paper, we propose a semantic-level differential privacy protection scheme that applies to the entire face dataset. Unlike pixel-level differential privacy approaches, our scheme guarantees that semantic privacy in faces is not compromised. The key idea is to convert unstructured data into structured data to enable the application of differential privacy. Specifically, we first extract semantic information from the face dataset to build an attribute database, then apply differential perturbations to obscure this attribute data, and finally use an image synthesis model to generate a protected face dataset. Extensive experimental results show that our scheme can maintain visual naturalness and balance the privacy-utility trade-off compared to the mainstream schemes.

cs.CV

Simplifying Textured Triangle Meshes in the Wild

This paper introduces a method for simplifying textured surface triangle meshes in the wild while maintaining high visual quality. While previous methods achieve excellent results on manifold meshes by using the quadric error metric, they struggle to produce high-quality outputs for meshes in the wild, which typically contain non-manifold elements and multiple connected components. In this work, we propose a method for simplifying these wild textured triangle meshes. We formulate mesh simplification as a problem of decimating simplicial 2-complexes to handle multiple non-manifold mesh components collectively. Building on the success of quadric error simplification, we iteratively collapse 1-simplices (vertex pairs). Our approach employs a modified quadric error that converges to the original quadric error metric for watertight manifold meshes, while significantly improving the results on wild meshes. For textures, instead of following existing strategies to preserve UVs, we adopt a novel perspective which focuses on computing mesh correspondences throughout the decimation, independent of the UV layout. This combination yields a textured mesh simplification system that is capable of handling arbitrary triangle meshes, achieving to high-quality results on wild inputs without sacrificing the excellent performance on clean inputs. Our method guarantees to avoid common problems in textured mesh simplification, including the prevalent problem of texture bleeding. We extensively evaluate our method on multiple datasets, showing improvements over prior techniques through qualitative, quantitative, and user study evaluations.

cs.GR

Infrared spectroscopy study of kagome material CsTi$_3$Bi$_5$

The kagome material CsTi$_3$Bi$_5$, which is isostructural to the extensively studied charge density wave (CDW) compound CsV$_3$Sb$_5$, exhibits intriguing electronic features within its two-dimensional kagome lattices of titanium atoms. Here, we perform optical spectroscopic measurements together with the first-principles calculations on single-crystalline CsTi$_3$Bi$_5$ to investigate its electronic properties comprehensively. It is found that the overall optical spectra are very similar to those of CsV$_3$Sb$_5$, but the existence of CDW instability is ruled out in CsTi$_3$Bi$_5$. Via careful comparison to the optical responses of CsV$_3$Sb$_5$, we attribute this difference to a significant reduction in the itinerant carrier density of CsTi$_3$Bi$_5$, which is associated with the absence of van Hove singularity near the Fermi level at $M$ point. This result supports the scenario that the CDW in CsV$_3$Sb$_5$ is driven by the nesting of van Hove singularity. Additionally, we unveil some exotic low-lying absorption features, which provide clear evidence of flat bands in CsTi$_3$Bi$_5$. Our findings contribute to a deeper understanding of exotic phenomena in CsTi$_3$Bi$_5$ and provide valuable insights into the role of van Hove singularity in CsV$_3$Sb$_5$.

cond-mat.str-el

MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis

We present a Multi-Instance Generation (MIG) task, simultaneously generating multiple instances with diverse controls in one image. Given a set of predefined coordinates and their corresponding descriptions, the task is to ensure that generated instances are accurately at the designated locations and that all instances' attributes adhere to their corresponding description. This broadens the scope of current research on Single-instance generation, elevating it to a more versatile and practical dimension. Inspired by the idea of divide and conquer, we introduce an innovative approach named Multi-Instance Generation Controller (MIGC) to address the challenges of the MIG task. Initially, we break down the MIG task into several subtasks, each involving the shading of a single instance. To ensure precise shading for each instance, we introduce an instance enhancement attention mechanism. Lastly, we aggregate all the shaded instances to provide the necessary information for accurately generating multiple instances in stable diffusion (SD). To evaluate how well generation models perform on the MIG task, we provide a COCO-MIG benchmark along with an evaluation pipeline. Extensive experiments were conducted on the proposed COCO-MIG benchmark, as well as on various commonly used benchmarks. The evaluation results illustrate the exceptional control capabilities of our model in terms of quantity, position, attribute, and interaction. Code and demos will be released at https://migcproject.github.io/.

cs.CV

Fusion-stable structures on triangulated categories

Let $\mathcal{G}$ be a fusion category acting on a triangulated category $\mathcal{D}$, in the sense that $\mathcal{D}$ is a $\mathcal{G}$-module category. Our motivation example is fusion-weighted species, which is essentially Heng's construction. We study $\mathcal{G}$-stable tilting, cluster and stability structures on $\mathcal{D}$. In particular, we prove the deformation theorem for $\mathcal{G}$-stable stability conditions. A first application is that Duffield-Tumarkin's categorification of cluster exchange graphs of finite Coxeter-Dynkin type can be naturally realized as fusion-stable cluster exchange graphs. Another application is that the universal cover of the hyperplane arrangements of any finite Coxeter-Dynkin type can be realized as the space of fusion-stable stability conditions for certain ADE Dynkin quiver. This provides an alternative uniform proof of $K(\pi,1)$-conjecture in the finite Coxeter-Dynkin case.

math.RT

Geometric classification of total stability spaces

We construct a geometric model for the root category $\mathcal{D}^b(Q)/[2]$ of any Dynkin diagram $Q$, which is an $h_Q$-gon $\mathbf{V}_Q$ with cores, where $h_Q$ is the Coxeter number and $\mathcal{D}^b(Q)$ is the bounded derived category associated to $Q$. As an application, we classify all spaces $\mathrm{ToSt}\mathcal{D}$ of total stability conditions on triangulated categories $\mathcal{D}$, where $\mathcal{D}$ must be of the form $\mathcal{D}^b(Q)$. More precisely, we prove that $\mathrm{ToSt}\mathcal{D}^b(Q)/[2]$ is isomorphic to a suitable moduli space of stable $h_Q$-gons of type $Q$. In particular, an $h_Q$-gon $\mathbf{V}$ of type $D_n$ is a (centrally) symmetric doubly punctured $2(n-1)$-gon. $\mathbf{V}$ is stable if it is convex and the punctures are inside the level-$(n-2)$ diagonal-gon. Another interesting case is $E_6$, where the (stable) $h_Q$-gon (dodecagon) can be realized as a pair of planar tiling pattern.

math.RT

Finitary birepresentations of finitary bicategories

In this paper, we discuss the generalization of finitary $2$-representation theory of finitary $2$-categories to finitary birepresentation theory of finitary bicategories. In previous papers on the subject, the classification of simple transitive $2$-representations of a given $2$-category was reduced to that for certain subquotients. These reduction results were all formulated as bijections between equivalence classes of $2$-representations. In this paper, we generalize them to biequivalences between certain $2$-categories of birepresentations. Furthermore, we prove an analog of the double centralizer theorem in finitary birepresentation theory.

math.RT

Adjunction in the absence of identity

We develop a bicategorical setup in which one can speak about adjoint 1-morphisms even in the absence of genuine identity 1-morphisms. We also investigate which part of 2-representation theory of 2-categories extends to this new setup.

math.CT

Furnishing Your Room by What You See: An End-to-End Furniture Set Retrieval Framework with Rich Annotated Benchmark Dataset

Understanding interior scenes has attracted enormous interest in computer vision community. However, few works focus on the understanding of furniture within the scenes and a large-scale dataset is also lacked to advance the field. In this paper, we first fill the gap by presenting DeepFurniture, a richly annotated large indoor scene dataset, including 24k indoor images, 170k furniture instances and 20k unique furniture identities. On the dataset, we introduce a new benchmark, named furniture set retrieval. Given an indoor photo as input, the task requires to detect all the furniture instances and search a matched set of furniture identities. To address this challenging task, we propose a feature and context embedding based framework. It contains 3 major contributions: (1) An improved Mask-RCNN model with an additional mask-based classifier is introduced for better utilizing the mask information to relieve the occlusion problems in furniture detection context. (2) A multi-task style Siamese network is proposed to train the feature embedding model for retrieval, which is composed of a classification subnet supervised by self-clustered pseudo attributes and a verification subnet to estimate whether the input pair is matched. (3) In order to model the relationship of the furniture entities in an interior design, a context embedding model is employed to re-rank the retrieval results. Extensive experiments demonstrate the effectiveness of each module and the overall system.

cs.CV

Quantum (dual) Grassmann superalgebra as $\mathcal U_q(\mathfrak{gl}(m|n))$-module algebra and beyond

We introduce and define the quantum affine $(m|n)$-superspace (or say quantum Manin superspace) $A_q^{m|n}$ and its dual object, the quantum Grassmann superalgebra $Ω_q(m|n)$. Correspondingly, a quantum Weyl algebra $\mathcal W_q(2(m|n))$ of $(m|n)$-type is introduced as the quantum differential operators (QDO for short) algebra $\textrm{Diff}_q(Ω_q)$ defined over $Ω_q(m|n)$, which is a smash product of the quantum differential Hopf algebra $\mathfrak D_q(m|n)$ (isomorphic to the bosonization of the quantum Manin superspace) and the quantum Grassmann superalgebra $Ω_q(m|n)$. An interested point of this approach here is that even though $\mathcal W_q(2(m|n))$ itself is in general no longer a Hopf algebra, so are some interesting sub-quotients existed inside. This point of view gives us one of main expected results, that is, the quantum (restricted) Grassmann superalgebra $Ω_q$ is made into the $\mathcal U_q(\mathfrak g)$-module (super)algebra structure,$Ω_q=Ω_q(m|n)$ for $q$ generic, or $Ω_q(m|n, \bold 1)$ for $q$ root of unity, and $\mathfrak g=\mathfrak{gl}(m|n)$ or $\mathfrak {sl}(m|n)$, the general or special linear Lie superalgebra. This QDO approach provides us with explicit realization models for some simple $\mathcal U_q(\mathfrak g)$-modules, together with the concrete information on their dimensions. Similar results hold for the quantum dual Grassmann superalgebra $Ω_q^!$ as $\mathcal U_q(\mathfrak g)$-module algebra.In the paper some examples of pointed Hopf algebras can arise from the QDOs, whose idea is an expansion of the spirit noted by Manin in \cite{Ma}, \& \cite{Ma1}.

math.QA

2-categories of symmetric bimodules and their 2-representations

In this article we analyze the structure of $2$-categories of symmetric projective bimodules over a finite dimensional algebra with respect to the action of a finite abelian group. We determine under which condition the resulting $2$-category is fiat (in the sense of \cite{MM1}) and classify simple transitive $2$-representations of this $2$-category (under some mild technical assumption). We also study several classes of examples in detail.

math.RT

Extreme representations of semirings

This is a write-up of the discussions during the meetings of the study group on representation theory of semirings which was organized at the Department of Mathematics, Uppsala University, during the academic year 2017-2018. The main emphasis is on classification of various classes of "irreducible" representations for various concrete semirings.

math.RT