Searcharxiv⌕ Search

arXiv subjects

Guo Li

Publications and source records attributed to Guo Li.

36 records · Page 2Linked to original sources

FedHL: Federated Learning for Heterogeneous Low-Rank Adaptation via Unbiased Aggregation

Federated Learning (FL) facilitates the fine-tuning of Foundation Models (FMs) using distributed data sources, with Low-Rank Adaptation (LoRA) gaining popularity due to its low communication costs and strong performance. While recent work acknowledges the benefits of heterogeneous LoRA in FL and introduces flexible algorithms to support its implementation, our theoretical analysis reveals a critical gap: existing methods lack formal convergence guarantees due to parameter truncation and biased gradient updates. Specifically, adapting client-specific LoRA ranks necessitates truncating global parameters, which introduces inherent truncation errors and leads to subsequent inaccurate gradient updates that accumulate over training rounds, ultimately degrading performance. To address the above issues, we propose \textbf{FedHL}, a simple yet effective \textbf{Fed}erated Learning framework tailored for \textbf{H}eterogeneous \textbf{L}oRA. By leveraging the full-rank global model as a calibrated aggregation basis, FedHL eliminates the direct truncation bias from initial alignment with client-specific ranks. Furthermore, we derive the theoretically optimal aggregation weights by minimizing the gradient drift term in the convergence upper bound. Our analysis shows that FedHL guarantees $\mathcal{O}(1/\sqrt{T})$ convergence rate, and experiments on multiple real-world datasets demonstrate a 1-3\% improvement over several state-of-the-art methods.

cs.LG↗

A Method to Decipher "Genome" from Interatomic Cohesion in the Exploration for a "Central Dogma" Replacement in Material Science

In the ball-stick model, interatomic cohesions are considered "sticks". But enormous details and features of the "sticks" are usually oversimplified as indexed quantities or equivocated as geometry characteristics. These indexed quantities or geometry characteristics not only limit the explanatory capability to a few chemical/physical aspects but also eliminate generativity for expected resemblance. And these limitations can be related to the information loss during the conversion. Herein, inspired by the central dogma, a framework is introduced to compact interatomic cohesions into a detailed residue-by-residue "genome" with matched encoding/decoding tools. The framework fuses the quantum mechanical aspects, auto feature extraction, nanostructures and/or simulations, and generative models. As a proof of concept, the realization introduced in this work adopted bosonic/fermionic features, an autoencoder with image recognition processes, Density Functional Theory simulations, and a thiolate-protected gold nanocluster dataset. After repetitive modeling, validating, and analysis based on 26,528 simulated interatomic images, the interatomic cohesion can be almost losslessly encoded into an 8-value-genome, and the genome encoder-decoder pair is also obtained. The model is then automatically extended into a generative model which converts any arbitrary 8-value-genome to a bond image.

cond-mat.mes-hall↗

Tenplex: Dynamic Parallelism for Deep Learning using Parallelizable Tensor Collections

Deep learning (DL) jobs use multi-dimensional parallelism, i.e. combining data, model, and pipeline parallelism, to use large GPU clusters efficiently. Long-running jobs may experience changes to their GPU allocation: (i) resource elasticity during training adds or removes GPUs; (ii) hardware maintenance may require redeployment on different GPUs; and (iii) GPU failures force jobs to run with fewer devices. Current DL frameworks tie jobs to a set of GPUs and thus lack support for these scenarios. In particular, they cannot change the multi-dimensional parallelism of an already-running job in an efficient and model-independent way. We describe Scalai, a state management library for DL systems that enables jobs to change their parallelism dynamically after the GPU allocation is updated at runtime. Scalai achieves this through a new abstraction, a parallelizable tensor collection (PTC), that externalizes the job state during training. After a GPU change, Scalai uses the PTC to transform the job state: the PTC repartitions the dataset state under data parallelism and exposes it to DL workers through a virtual file system; and the PTC obtains the model state as partitioned checkpoints and transforms them to reflect the new parallelization configuration. For efficiency, Scalai executes PTC transformations in parallel with minimum data movement between workers. Our experiments show that Scalai enables DL jobs to support dynamic parallelization with low overhead.

cs.DC↗

Functional Group Induced Transformations in Stacking and Electron Structure in Mo2CTx/NiS Heterostructures

The two-dimensional transition metal carbide/nitride family (MXenes) has garnered significant attention due to their highly customizable surface functional groups. Leveraging modern material science techniques, the customizability of MXenes can be enhanced further through the construction of associated heterostructures. As indicated by recent research, the Mo2CTx/NiS heterostructure has emerged as a promising candidate exhibiting superior physical and chemical application potential. The geometrical structure of Mo2CTx/NiS heterostructure is modeled and 6 possible configurations are validated by Density Functional Theory simulations. The variation in functional groups leads to structural changes in Mo2CTx/NiS interfaces, primarily attributed to the competition between van der Waals and covalent interactions. The presence of different functional groups results in significant band fluctuations near the Fermi level for Ni and Mo atoms, influencing the role of atoms and electron's ability to escape near the interface. This, in turn, modulates the strength of covalent interactions at the MXenes/NiS interface and alters the ease of dissociation of the MXenes/NiS complex. Notably, the Mo2CO2/NiS(P6_3/mmc) heterostructure exhibits polymorphism, signifying that two atomic arrangements can stabilize the structure. The transition process between these polymorphs is also simulated, further indicating the modulation of the electronic level of properties by a sliding operation.

cond-mat.mes-hall↗

Quiver: Supporting GPUs for Low-Latency, High-Throughput GNN Serving with Workload Awareness

Systems for serving inference requests on graph neural networks (GNN) must combine low latency with high throughout, but they face irregular computation due to skew in the number of sampled graph nodes and aggregated GNN features. This makes it challenging to exploit GPUs effectively: using GPUs to sample only a few graph nodes yields lower performance than CPU-based sampling; and aggregating many features exhibits high data movement costs between GPUs and CPUs. Therefore, current GNN serving systems use CPUs for graph sampling and feature aggregation, limiting throughput. We describe Quiver, a distributed GPU-based GNN serving system with low-latency and high-throughput. Quiver's key idea is to exploit workload metrics for predicting the irregular computation of GNN requests, and governing the use of GPUs for graph sampling and feature aggregation: (1) for graph sampling, Quiver calculates the probabilistic sampled graph size, a metric that predicts the degree of parallelism in graph sampling. Quiver uses this metric to assign sampling tasks to GPUs only when the performance gains surpass CPU-based sampling; and (2) for feature aggregation, Quiver relies on the feature access probability to decide which features to partition and replicate across a distributed GPU NUMA topology. We show that Quiver achieves up to 35 times lower latency with an 8 times higher throughput compared to state-of-the-art GNN approaches (DGL and PyG).

cs.DC↗

Spin Cooperated Catalytic Activities in Mn-N4 based Single-atom Nanozyme: Mechanisms and a Brief Charge-spin Model

Although developing artificial enzymes has made great progress, there is still a gap between artificial enzymes and natural enzymes in catalytic performance. Designing and constructing efficient artificial biocatalysts is extremely desirable because of their high stability, low cost and easy storage. Here, we report a synthesized amino-functionalized graphene quantum dots-based manganese single atom catalyst (SAC) Mn-N4, which exhibits POD-, CAT, SOD-like activities, especially the superior SOD-like activity. Recent studies have reported Mn-based SAzymes, however, the multi-enzyme mimicking catalytic mechanisms for Mn-N4 are not comprehensive and in-depth enough. Therefore, we combine density functional theory (DFT) calculations and machine learning (ML) to validate the performance of the multi-enzyme mimicking activities. The DFT simulations show that Mn-N4 owns a highly effective SOD in the "one-side adsorption" with a very low energy barrier of 0.077 eV, which can be attributed to variation of the preferred spin states of Mn-O2.- system and its "spin flip-collection lock" in the SOD-like catalytic procedure. Furthermore, spin related charge distributions on Mn-N4 configurations by machine learning (ML) analysis suggest that the pattern of spin and natural charge/valence electron distribution will exhibit similarity in the structures of multiple intermediate steps of multi-enzyme mimicking activities. This work not only puts forward the catalytic mechanisms of Mn-N4 SAzymes, but also provides essential guidance for future design of highly performance artificial enzymes.

cond-mat.mes-hall↗

For Intelligent and Higher Spectrum Efficiency: A Variable Packing Ratio Transmission System Based on Faster-than-Nyquist and Deep Learning

With the rapid development of various services in wireless communications, spectrum resource has become increasingly valuable. Faster than Nyquist (FTN) signaling, proposed in the 1970s, is a promising paradigm for improving spectrum utilization. This paper proposes intelligent variable-packing-ratio (VPR)-based transmissions for high spectrum efficiency (SE) and security, respectively. Aided by deep learning (DL)-based estimation, the proposed scheme for high SE can achieve a higher capacity with negligible modification to existing communication paradigms (e.g., spectrum allocation or frame structure). Also, for VPR-based secure transmission, a dynamic generation scheme is proposed to produce randomly distributed positions to switch the packing ratio, which can effectively avoid detections and attacks. In addition, we propose a simplified DL-based packing ratio estimation for both of these two scenarios so that the receiver can estimate the packing ratio without any in-band or out-band control messages. Simulation results show that the proposed simplified estimation achieves nearly the same accuracy and convergence speed as the original multi-branch fully-connected structure with a complexity reduction of 20 folds. Finally, we derive the closed-form SE of the proposed VPR transmission under different channels. The numerical results validate the correctness of the derivation and demonstrate the SE gains of the VPR scheme beyond conventional Nyquist transmission.

eess.SP↗

Fast and Flexible Human Pose Estimation with HyperPose

Estimating human pose is an important yet challenging task in multimedia applications. Existing pose estimation libraries target reproducing standard pose estimation algorithms. When it comes to customising these algorithms for real-world applications, none of the existing libraries can offer both the flexibility of developing custom pose estimation algorithms and the high-performance of executing these algorithms on commodity devices. In this paper, we introduce Hyperpose, a novel flexible and high-performance pose estimation library. Hyperpose provides expressive Python APIs that enable developers to easily customise pose estimation algorithms for their applications. It further provides a model inference engine highly optimised for real-time pose estimation. This engine can dynamically dispatch carefully designed pose estimation tasks to CPUs and GPUs, thus automatically achieving high utilisation of hardware resources irrespective of deployment environments. Extensive evaluation results show that Hyperpose can achieve up to 3.1x~7.3x higher pose estimation throughput compared to state-of-the-art pose estimation libraries without compromising estimation accuracy. By 2021, Hyperpose has received over 1000 stars on GitHub and attracted users from both industry and academy.

cs.CV↗

Efficient Reinforcement Learning Development with RLzoo

Many researchers and developers are exploring for adopting Deep Reinforcement Learning (DRL) techniques in their applications. They however often find such an adoption challenging. Existing DRL libraries provide poor support for prototyping DRL agents (i.e., models), customising the agents, and comparing the performance of DRL agents. As a result, the developers often report low efficiency in developing DRL agents. In this paper, we introduce RLzoo, a new DRL library that aims to make the development of DRL agents efficient. RLzoo provides developers with (i) high-level yet flexible APIs for prototyping DRL agents, and further customising the agents for best performance, (ii) a model zoo where users can import a wide range of DRL agents and easily compare their performance, and (iii) an algorithm that can automatically construct DRL agents with custom components (which are critical to improve agent's performance in custom applications). Evaluation results show that RLzoo can effectively reduce the development cost of DRL agents, while achieving comparable performance with existing DRL libraries.

cs.AI↗

Learning Generalized Spatial-Temporal Deep Feature Representation for No-Reference Video Quality Assessment

In this work, we propose a no-reference video quality assessment method, aiming to achieve high-generalization capability in cross-content, -resolution and -frame rate quality prediction. In particular, we evaluate the quality of a video by learning effective feature representations in spatial-temporal domain. In the spatial domain, to tackle the resolution and content variations, we impose the Gaussian distribution constraints on the quality features. The unified distribution can significantly reduce the domain gap between different video samples, resulting in a more generalized quality feature representation. Along the temporal dimension, inspired by the mechanism of visual perception, we propose a pyramid temporal aggregation module by involving the short-term and long-term memory to aggregate the frame-level quality. Experiments show that our method outperforms the state-of-the-art methods on cross-dataset settings, and achieves comparable performance on intra-dataset configurations, demonstrating the high-generalization capability of the proposed method.

eess.IV↗

Synthetic high-order PT symmetry in a single coil resonator

The exploration of non-Hermitian systems with parity-time (PT) symmetry has witnessed immense research interest both fundamentally and technologically in a wide range of subject areas in physics and engineering. One significant example of the principal emerging fields in this context is the PT symmetric wireless applications using multiple coils that are spatially separated but mutually coupled with position-dependent coupling strength. Such a spatial PT configuration limits the flexibility and miniaturization of the PT symmetric designs. As far as this is concerned, inspired by scattering induced two opposite whispering-gallery (WG) modes in an optical resonator, analogously here we experimentally demonstrate a specially constructed second-order (2-nd order) PT symmetry in a single coil resonator, whose currents with two different directions are excited by internal bypass capacitor. Our proposed structure has the following peculiar feature: First, the bypass capacitor induces coupling in spectral resonances allow us to observe a 2-nd order phase transition between symmetry regimes, without the need of a second coil in the spatial PT case. Under this circumstance, this specially constructed PT symmetry can be regarded as synthetic PT symmetry, which is enabled by coupling modes with different directions. Second, by introducing two or more internal bypass capacitors, the synthetic high-order PT symmetric system bearing such as third-order exceptional point in a single coil resonator can be realized. These results will provide a new paradigm to realize higher-order PT symmetry towards the investigation of non-Hermitian physics in a synthetic perspective, which can be extended to other physical platforms such as optics and acoustics.

physics.app-ph↗

Receiver Design for Faster-than-Nyquist Signaling: Deep-learning-based Architectures

Faster-than-Nyquist (FTN) is a promising paradigm to improve bandwidth utilization at the expense of additional intersymbol interference (ISI). In this paper, we apply state-of-the-art deep learning (DL) technology into receiver design for FTN signaling and propose two DL-based new architectures. Firstly, we propose an FTN signal detection based on DL and connect it with the successive interference cancellation (SIC) to replace traditional detection algorithms. Simulation results show that this architecture can achieve near-optimal performance in both uncoded and coded scenarios. Additionally, we propose a DL-based joint signal detection and decoding for FTN signaling to replace the complete baseband part in traditional FTN receivers. The performance of this new architecture has also been illustrated by simulation results. Finally, both the proposed DL-based receiver architecture has the robustness to signal to noise ratio (SNR). In a nutshell, DL has been proved to be a powerful tool for the FTN receiver design.

eess.SP↗

Seamless maps of major elements of the Moon: Results from high-resolution geostationary satellite

Major elements such as Fe, Ti, Mg, Al, Ca, and Si play very important roles in understanding the origin and evolution of the Moon. Previous maps of these major elements derived from orbital data are based on mosaic images or low-resolution Gamma ray data. The hue variations and gaps among orbital boundaries in the mosaic images are not conducive to geological studies. This paper aims to produce seamless and homogenous distribution maps of major elements using the single-exposure image of the whole lunar disk obtained by China's high-resolution geostationary satellite, Gaofen-4, with a spatial resolution of ~500 m. The elemental contents of soil samples returned by Apollo and Luna missions were used as ground truth, and were correlated with the reflectance of the sampling sites extracted from Gaofen-4 data. The final distribution maps of these major oxides are generated with the statistical regression model. With these products the average contents and proportions of the major elements for maria and highlands were estimated and compared. The results showed that SiO2 and TiO2 have the highest and lowest fractions in mare and highland areas, respectively. Besides, the relative concentrations of these elements could serve as indicators of geologic processes, e.g., the obviously asymmetric distributions of Al2O3, CaO, and SiO2 around Tycho crater may suggest that Tycho crater was formed by an oblique impact from the southwest direction.

astro-ph.EP↗

Unveiling the secrets of the mid-infrared Moon

The Moon's optical characteristics in visible and long-wavelength infrared (LWIR) have long been observed with our eyes or with instruments. What the mid-infrared (MIR) Moon looks like is still a mystery. For the first time we present detailed appearance of the MIR Moon observed by a high-resolution geostationary satellite and reveal the essence behind its appearance. The appearance of the MIR Moon is opposite to its normal visible appearance. In addition the MIR Moon shows limb darkening. Both the absolute and the relative brightness distribution of the MIR lunar disk changes with the solar incidence angle. The signatures of the MIR Moon are controlled by both the reflection and emission of the lunar surface. We also show first-ever brightness temperature maps of the lunar disk without needing a mosaic, which better show the temperature variation across the lunar disk. They reveal that the relationship between brightness temperature and solar incidence angle i is cos1/bi, and the power parameter is smaller than the Lambertian temperature model of cos1/4i observed for lunar orbit-based measurements. The slower decrease of the brightness temperature when moving away from the sub-solar point than the Lambertian model is due to topographic effects. The brightness temperature is dominated by albedo and the solar incidence angle and influenced by the topography. Our results indicate that the Moon in the MIR exhibits many interesting phenomena which were previously unknown, and contains abundant information about lunar reflection and thermal emission for future study.

astro-ph.EP↗

TailorGAN: Making User-Defined Fashion Designs

Attribute editing has become an important and emerging topic of computer vision. In this paper, we consider a task: given a reference garment image A and another image B with target attribute (collar/sleeve), generate a photo-realistic image which combines the texture from reference A and the new attribute from reference B. The highly convoluted attributes and the lack of paired data are the main challenges to the task. To overcome those limitations, we propose a novel self-supervised model to synthesize garment images with disentangled attributes (e.g., collar and sleeves) without paired data. Our method consists of a reconstruction learning step and an adversarial learning step. The model learns texture and location information through reconstruction learning. And, the model's capability is generalized to achieve single-attribute manipulation by adversarial learning. Meanwhile, we compose a new dataset, named GarmentSet, with annotation of landmarks of collars and sleeves on clean garment images. Extensive experiments on this dataset and real-world samples demonstrate that our method can synthesize much better results than the state-of-the-art methods in both quantitative and qualitative comparisons.

cs.CV↗

Active collaboration in relative observation for Multi-agent visual SLAM based on Deep Q Network

This paper proposes a unique active relative localization mechanism for multi-agent Simultaneous Localization and Mapping(SLAM),in which a agent to be observed are considered as a task, which is performed by others assisting that agent by relative observation. A task allocation algorithm based on deep reinforcement learning are proposed for this mechanism. Each agent can choose whether to localize other agents or to continue independent SLAM on it own initiative. By this way, the process of each agent SLAM will be interacted by the collaboration. Firstly, based on the characteristics of ORBSLAM, a unique observation function which models the whole MAS is obtained. Secondly, a novel type of Deep Q network(DQN) called MAS-DQN is deployed to learn correspondence between Q Value and state-action pair,abstract representation of agents in MAS are learned in the process of collaboration among agents. Finally, each agent must act with a certain degree of freedom according to MAS-DQN. The simulation results of comparative experiments prove that this mechanism improves the efficiency of cooperation in the process of multi-agent SLAM.

cs.AI↗

Blind Estimation Algorithms for I/Q Imbalance in Direct Down-conversion Receivers

As known, receivers with in-phase and quadrature phase (I/Q) down conversion, especially direct-conversion architectures, always suffer from I/Q imbalance. I/Q imbalance is caused by amplitude and phase mismatch between I/Q paths. The performance degradation resulting from I/Q imbalance can not be mitigated with simply higher signal to noise ratio (SNR). Thus, I/Q imbalance compensation in the digital domain is critical. There are two main contributions in this paper. Firstly, we proposed a blind estimation algorithm for I/Q imbalance parameters based on joint first and second order statistics (FSS) which has lower complexity than conventional Gaussian maximum likelihood estimation (GMLE). This can be used for further processing such as equalization in the presence of receiver IQ imbalance. In addition, we find out the reason of the error floor in conventional I/Q imbalance compensation method based on the conjugate signal model (CSM). The proposed joint first order statistics and conjugate signal model (FSCSM) compensation algorithm can reach the ideal bit error rate (BER) performance.

eess.SP↗

Molecular Adsorption on Metal Surfaces with a van der Waals Density Functional

The adsorption of 1,4-benzenediamine (BDA) on the Au(111) surface and azobenzene on the Ag(111) surface is investigated using density functional theory (DFT) with a non-local density functional (vdW-DF) and a semi-local Perdew-Burke-Ernzerhof (PBE) functional. For BDA on Au(111), the inclusion of London dispersion interactions not only dramatically enhances the molecule-substrate binding, resulting in adsorption energies consistent with experimental results, but also significantly alters the BDA binding geometry. For azobenzene on Ag(111), the vdW-DF produces superior adsorption energies compared to those obtained with other dispersion corrected DFT approaches. These results provide evidence for the applicability of the vdW-DF method and serves as a practical benchmark for the investigation of molecules adsorbed on noble metal surfaces.

cond-mat.mtrl-sci↗