Searcharxiv⌕ Search

arXiv subjects

Han Xu

Publications and source records attributed to Han Xu.

At least 91 records · Page 5Linked to original sources

Exploring Memorization in Fine-tuned Language Models

Large language models (LLMs) have shown great capabilities in various tasks but also exhibited memorization of training data, raising tremendous privacy and copyright concerns. While prior works have studied memorization during pre-training, the exploration of memorization during fine-tuning is rather limited. Compared to pre-training, fine-tuning typically involves more sensitive data and diverse objectives, thus may bring distinct privacy risks and unique memorization behaviors. In this work, we conduct the first comprehensive analysis to explore language models' (LMs) memorization during fine-tuning across tasks. Our studies with open-sourced and our own fine-tuned LMs across various tasks indicate that memorization presents a strong disparity among different fine-tuning tasks. We provide an intuitive explanation of this task disparity via sparse coding theory and unveil a strong correlation between memorization and attention score distribution.

cs.AI↗

Tame quivers and affine bases II: nonsimply-laced cases

In [Tame_quivers_and_affine_bases_I], we give a Ringel-Hall algebra approach to the canonical bases in the symmetric affine cases. In this paper, we extend the results to general symmetrizable affine cases by using Ringel-Hall algebras of representations of a valued quiver. We obtain a bar-invariant basis $\mathbf{B}'=\{C(\mathbf{c},t_λ)|(\mathbf{c},t_λ)\in\mathcal{G}^a\}$ in the generic composition algebra $\mathcal{C}^*$ and prove that $\mathcal{B}'=\mathbf{B}'\sqcup(-\mathbf{B}')$ coincides with Lusztig's signed canonical basis $\mathcal{B}$. Moreover, in type $\tilde{B}_n,\tilde{C}_n$, $\mathbf{B}'$ is the canonical basis $\mathbf{B}$.

math.RT↗

A Scalable Network-Aware Multi-Agent Reinforcement Learning Framework for Decentralized Inverter-based Voltage Control

This paper addresses the challenges associated with decentralized voltage control in power grids due to an increase in distributed generations (DGs). Traditional model-based voltage control methods struggle with the rapid energy fluctuations and uncertainties of these DGs. While multi-agent reinforcement learning (MARL) has shown potential for decentralized secondary control, scalability issues arise when dealing with a large number of DGs. This problem lies in the dominant centralized training and decentralized execution (CTDE) framework, where the critics take global observations and actions. To overcome these challenges, we propose a scalable network-aware (SNA) framework that leverages network structure to truncate the input to the critic's Q-function, thereby improving scalability and reducing communication costs during training. Further, the SNA framework is theoretically grounded with provable approximation guarantee, and it can seamlessly integrate with multiple multi-agent actor-critic algorithms. The proposed SNA framework is successfully demonstrated in a system with 114 DGs, providing a promising solution for decentralized voltage control in increasingly complex power grid systems.

math.OC↗

Accurate Time-segmented Loss Model for SiC MOSFETs in Electro-thermal Multi-Rate Simulation

Compared with silicon (Si) power devices, Silicon carbide (SiC) devices have the advantages of fast switching speed and low on-resistance. However, the effects of non-ideal characteristics of SiC MOSFETs and stray parameters (especially parasitic inductance) on switching losses need to be further evaluated. In this paper, a transient loss model based on SiC MOSFET and SiC Schottky barrier diode (SBD) switching pairs is proposed. The transient process analysis is simplified by time segmentation of the transient process of power switching devices. The electro-thermal simulation calculates the junction temperature and updates the temperature-related parameters with the proposed loss model and the thermal network model. A multi-rate data exchange strategy is proposed to solve the problem of disparity in timescales between circuit simulation and thermal network simulation. The CREE CMF20120D SiC MOSFET device is used for the experimental verification. The experimental results verify the accuracy of the model which provides guidance for the circuit design of SiC MOSFETs. All the parameters of the loss model can be extracted from the datasheet, which is practical in power electronics design.

eess.SY↗

An Event-Based Synchronization Framework for Controller Hardware-in-the-loop Simulation of Electric Railway Power Electronics Systems

The Controller Hardware_in_the_loop (CHIL) simulation is gaining popularity as a cost_effective, efficient, and reliable tool in the design and development process of fast_growing electrified transportation power converters. However, it is challenging to implement the conventional CHIL simulations on the railway power converters with complex topologies and high switching frequencies due to strict real_time constraints. Therefore, this paper proposes an event-based synchronization CHIL (ES_CHIL) framework for high_fidelity simulation of these electrified railway power converters. Different from conventional CHIL simulations synchronized through the time axis, the ES_CHIL framework is synchronized through the event axis. Therefore, it can ease the real_time constraint and broaden the upper bound on the system size and switching frequency. Besides, models and algorithms with higher accuracy, such as the diode model with natural commutation processes, can be used in the ES-CHIL framework. The proposed framework is validated for a 350 kW wireless power transformer system containing 24 fully controlled devices and 36 diodes by comparing it with Simulink and physical experiments. This research improves the fidelity and application range of the power converters CHIL simulation. Thus, it helps to accelerate the prototype design and performance evaluation process for electrified railways and other applications with such complex converters.

eess.SY↗

Probabilistic Categorical Adversarial Attack & Adversarial Training

The existence of adversarial examples brings huge concern for people to apply Deep Neural Networks (DNNs) in safety-critical tasks. However, how to generate adversarial examples with categorical data is an important problem but lack of extensive exploration. Previously established methods leverage greedy search method, which can be very time-consuming to conduct successful attack. This also limits the development of adversarial training and potential defenses for categorical data. To tackle this problem, we propose Probabilistic Categorical Adversarial Attack (PCAA), which transfers the discrete optimization problem to a continuous problem that can be solved efficiently by Projected Gradient Descent. In our paper, we theoretically analyze its optimality and time complexity to demonstrate its significant advantage over current greedy based attacks. Moreover, based on our attack, we propose an efficient adversarial training framework. Through a comprehensive empirical study, we justify the effectiveness of our proposed attack and defense algorithms.

cs.LG↗

FPGA-Based Implicit-Explicit Real-time Simulation Solver for Railway Wireless Power Transfer with Nonlinear Magnetic Coupling Components

Railway Wireless Power Transfer (WPT) is a promising non-contact power supply solution, but constructing prototypes for controller testing can be both costly and unsafe. Real-time hardware-in-the-loop simulation is an effective and secure testing tool, but simulating the dynamic charging process of railway WPT systems is challenging due to the continuous changes in the nonlinear magnetic coupling components. To address this challenge, we propose an FPGA-based half-step implicit-explicit (IMEX) simulation solver. The proposed solver adopts an IMEX algorithm to solve the piecewise linear and nonlinear parts of the system separately, which enables FPGAs to solve nonlinear components while achieving high numerical stability. Additionally, we divide a complete integration step into two half-steps to reduce computational time delays. Our proposed method offers a promising solution for the real-time simulation of railway WPT systems. The novelty of our approach lies in the use of the IMEX algorithm and the half-step integration method, which significantly improves the accuracy and efficiency of the simulation. Our simulations and experiments demonstrate the effectiveness and accuracy of the proposed solver, which provides a new approach for simulating and optimizing railway WPT systems with nonlinear magnetic coupling components.

eess.SY↗

Numerical Derivative-based Flexible Integration Algorithm for Power Electronic Systems Simulation Considering Nonlinear Components

Simulation is an efficient tool in the design and control of power electronic systems. However, quick and accurate simulation of them is still challenging, especially when the system contains a large number of switches and state variables. Conventional general-purpose integration algorithms assume nonlinearity within systems but face inefficiency in handling the piecewise characteristics of power electronic switches. While some specialized algorithms can adapt to the piecewise characteristics, most of these methods require systems to be piecewise linear. In this article, a numerical derivative-based flexible integration algorithm is proposed. This algorithm can adapt to the piecewise characteristic caused by switches and have no difficulty when nonlinear non-switching components are present in the circuit. This algorithm consists of a recursive numerical scheme that obtains high-order time derivatives of nonlinear components and a decoupling strategy that further increases computational efficiency. The proposed method is applied to solve a motor derive system and a large-scale power conversion system (PCS) to verify its accuracy and efficiency by comparing experimental waveforms and simulated results given by commercial software. Our proposed method demonstrates several-fold acceleration compared to multiple commonly used algorithms in Simulink.

eess.SY↗

Quantum Monte Carlo simulations of thermodynamic properties of attractive SU($3$) Dirac fermions

We employ the determinant quantum Monte Carlo method to study the finite-temperature properties of the half-filled attractive SU($3$) Hubbard model on a honeycomb lattice. We calculate the phase diagram in which the phase boundary separates the disordered phase and the charge-density-wave (CDW) phase and the transition temperature $T_{\text{tr}}(|U|)$ varies non-monotonically with attractive Hubbard interaction $|U|$. As the Hubbard $|U|$ increases at constant temperature $T<\text{max}(T_{\text{tr}}(|U|))$, the system first undergoes a transition from thermal Dirac semimetal phase to CDW phase, and eventually the CDW state is thermally melted at a strong Hubbard $|U|$ where the system enters a trion liquid phase. In between the two transition points the non-monotonic $|U|$ dependence of CDW order strength is strikingly different from the zero-temperature monotonic behavior. In the trion CDW state where off-site trions arise from quantum fluctuations (a fermion inside an on-site trion hops to a nearest-neighbor site), the simulated triple occupancy at constant Hubbard $|U|$ surprisingly increases with temperature, implying that the formation of off-site trions is suppressed by the thermal delocalization of on-site trions. We have also calculated the entropy-temperature relations for various attractive Hubbrad interactions, which exhibit the prominent characteristic of the Pomeranchuk effect. Our work has revealed that the formation of on-site and off-site trions has significant consequences for thermodynamic properties of SU(3) Dirac fermions.

cond-mat.quant-gas↗

On the Generalization of Training-based ChatGPT Detection Methods

ChatGPT is one of the most popular language models which achieve amazing performance on various natural language tasks. Consequently, there is also an urgent need to detect the texts generated ChatGPT from human written. One of the extensively studied methods trains classification models to distinguish both. However, existing studies also demonstrate that the trained models may suffer from distribution shifts (during test), i.e., they are ineffective to predict the generated texts from unseen language tasks or topics. In this work, we aim to have a comprehensive investigation on these methods' generalization behaviors under distribution shift caused by a wide range of factors, including prompts, text lengths, topics, and language tasks. To achieve this goal, we first collect a new dataset with human and ChatGPT texts, and then we conduct extensive studies on the collected dataset. Our studies unveil insightful findings which provide guidance for developing future methodologies or data collection strategies for ChatGPT detection.

cs.CL↗

Trion states and quantum criticality of attractive SU(3) Dirac fermions

We perform the projector quantum Monte Carlo (QMC) simulation to study the trion formation and quantum phase transition in the half-filled attractive SU(3) Hubbard model on a honeycomb lattice. With increasing attractive Hubbard interaction, our simulations demonstrate a continuous quantum phase transition from the semimetal to charge density wave (CDW) at the critical coupling $U_c/t=-1.52(2)$. The critical exponents $ν=0.82(3)$ and $η=0.58(4)$ determined by the QMC simulation remarkably disagree with those of the $N=3$ chiral Ising universality class suggested by the effective Gross-Neveu-Yukawa (GNY) theory, but coincide with the $N=1$ chiral Ising universality class. In the CDW phase, we show that on-site and off-site trions coexist and the off-site trion forms a local bond state. Our work not only illustrates the formation of off-site trions in two-dimensional Hubbard model, but also raises doubts about the extent of applicability of GNY model on the attractive SU(3) Dirac fermions.

cond-mat.quant-gas↗

Tame quivers and affine bases I: a Hall algebra approach to the canonical bases

For quantum group of affine type, Lusztig gave an explicit construction of the affine canonical basis by simple perverse sheaves. In this paper, we construct a bar-invariant basis by using a PBW basis arising from representations of the corresponding tame quiver. We prove that this bar-invariant basis coincides with Lusztig's canonical basis and obtain a concrete bijection between the elements in theses two bases. The index set of these bases is listed orderly by modules in preprojective, regular non-homogeneous, preinjective components and irreducible characters of symmetric groups. Our results are based on the work of Lin-Xiao-Zhang and closely related with the work of Beck-Nakajima. A crucial method in our construction is a generalization of that by Deng-Du-Xiao.

math.RT↗

Diff-Retinex: Rethinking Low-light Image Enhancement with A Generative Diffusion Model

In this paper, we rethink the low-light image enhancement task and propose a physically explainable and generative diffusion model for low-light image enhancement, termed as Diff-Retinex. We aim to integrate the advantages of the physical model and the generative network. Furthermore, we hope to supplement and even deduce the information missing in the low-light image through the generative network. Therefore, Diff-Retinex formulates the low-light image enhancement problem into Retinex decomposition and conditional image generation. In the Retinex decomposition, we integrate the superiority of attention in Transformer and meticulously design a Retinex Transformer decomposition network (TDN) to decompose the image into illumination and reflectance maps. Then, we design multi-path generative diffusion networks to reconstruct the normal-light Retinex probability distribution and solve the various degradations in these components respectively, including dark illumination, noise, color deviation, loss of scene contents, etc. Owing to generative diffusion model, Diff-Retinex puts the restoration of low-light subtle detail into practice. Extensive experiments conducted on real-world low-light datasets qualitatively and quantitatively demonstrate the effectiveness, superiority, and generalization of the proposed method.

cs.CV↗

PCDF: A Parallel-Computing Distributed Framework for Sponsored Search Advertising Serving

Traditional online advertising systems for sponsored search follow a cascade paradigm with retrieval, pre-ranking,ranking, respectively. Constrained by strict requirements on online inference efficiency, it tend to be difficult to deploy useful but computationally intensive modules in the ranking stage. Moreover, ranking models currently used in the industry assume the user click only relies on the advertisements itself, which results in the ranking stage overlooking the impact of organic search results on the predicted advertisements (ads). In this work, we propose a novel framework PCDF(Parallel-Computing Distributed Framework), allowing to split the computation cost into three parts and to deploy them in the pre-module in parallel with the retrieval stage, the middle-module for ranking ads, and the post-module for re-ranking ads with external items. Our PCDF effectively reduces the overall inference latency compared with the classic framework. The whole module is end-to-end offline training and adapt for the online learning paradigm. To our knowledge, we are the first to propose an end-to-end solution for online training and deployment on complex CTR models from the system framework side.

cs.IR↗

Nonzero angular momentum density wave phases in SU($N$) fermions with singlet-bond and triplet-current interactions

We employ the sign-problem-free projector determinant quantum Monte Carlo method to study a microscopic model of SU($N$) fermions with singlet-bond and triplet-current interactions on the square lattice. We find the gapped singlet $p_x$ and gapless triplet $d_{x^2-y^2}$ density wave states in the half-filled $N=4$ model. Specifically, the triplet $d_{x^2-y^2}$ density wave order is observed in the weak triplet-current interaction regime. As the triplet-current interaction strength is further increased, our simulations demonstrate a transition to the singlet $p_x$ density wave state, accompanied by a gapped mixed-ordered area where the two orders coexist. With increasing singlet-bond interaction strength, the triplet $d_{x^2-y^2}$-wave order persists up to a critical point after which the singlet $p_x$ density wave state is stabilized, while the ground state is disordered in between the two ordered phases. The analytical continuation is then performed to derive the single-particle spectrum. In the spectra of triplet $d_{x^2-y^2}$ and singlet $p_x$ density waves, the anisotropic Dirac cone and the parabolic shape around the Dirac point are observed, respectively. As for the mixed-ordered area, a single-particle gap opens and the velocities remain anisotropic at the Dirac point.

cond-mat.str-el↗

Jointly Attacking Graph Neural Network and its Explanations

Graph Neural Networks (GNNs) have boosted the performance for many graph-related tasks. Despite the great success, recent studies have shown that GNNs are highly vulnerable to adversarial attacks, where adversaries can mislead the GNNs' prediction by modifying graphs. On the other hand, the explanation of GNNs (GNNExplainer) provides a better understanding of a trained GNN model by generating a small subgraph and features that are most influential for its prediction. In this paper, we first perform empirical studies to validate that GNNExplainer can act as an inspection tool and have the potential to detect the adversarial perturbations for graphs. This finding motivates us to further initiate a new problem investigation: Whether a graph neural network and its explanations can be jointly attacked by modifying graphs with malicious desires? It is challenging to answer this question since the goals of adversarial attacks and bypassing the GNNExplainer essentially contradict each other. In this work, we give a confirmative answer to this question by proposing a novel attack framework (GEAttack), which can attack both a GNN model and its explanations by simultaneously exploiting their vulnerabilities. Extensive experiments on two explainers (GNNExplainer and PGExplainer) under various real-world datasets demonstrate the effectiveness of the proposed method.

cs.LG↗

VRKG4Rec: Virtual Relational Knowledge Graphs for Recommendation

Incorporating knowledge graph as side information has become a new trend in recommendation systems. Recent studies regard items as entities of a knowledge graph and leverage graph neural networks to assist item encoding, yet by considering each relation type individually. However, relation types are often too many and sometimes one relation type involves too few entities. We argue that it is not efficient nor effective to use every relation type for item encoding. In this paper, we propose a VRKG4Rec model (Virtual Relational Knowledge Graphs for Recommendation), which explicitly distinguish the influence of different relations for item representation learning. We first construct virtual relational graphs (VRKGs) by an unsupervised learning scheme. We also design a local weighted smoothing (LWS) mechanism for encoding nodes, which iteratively updates a node embedding only depending on the embedding of its own and its neighbors, but involve no additional training parameters. We also employ the LWS mechanism on a user-item bipartite graph for user representation learning, which utilizes encodings of items with relational knowledge to help training representations of users. Experiment results on two public datasets validate that our VRKG4Rec model outperforms the state-of-the-art methods. The implementations are available at https://github.com/lulu0913/VRKG4Rec.

cs.IR↗