SearcharxivSearch

arXiv subjects

Dongming Li

Publications and source records attributed to Dongming Li.

13 recordsLinked to original sources

LDA-1/2 for Molecular Systems: A Real-Space Finite-Element Benchmark on the GW100 Set

The LDA-1/2 method provides an efficient correction to semilocal density functional theory for improving ionization energies and band gaps, yet its application to molecular systems has remained limited. In this work, we present an all-electron finite-element implementation of LDA-1/2 within the NESSIE electronic-structure framework and apply it to the GW100 molecular benchmark set with systematically controllable numerical accuracy. The self-energy correction is constructed explicitly from neutral and half-ionized calculations for each molecule, avoiding the use of precomputed atomic correction potentials. The real-space finite-element formulation enables systematic convergence with respect to the discretization and provides a controlled assessment of LDA-1/2 performance. For the GW100 set, the present implementation yields a mean absolute error of 0.472 eV and a root-mean-square error of 0.645 eV relative to CCSD(T) reference ionization energies, substantially improving upon conventional LDA and the previously reported LAPW implementation of LDA-1/2, while achieving accuracy comparable to G0W0@PBE. Convergence tests show that third-order finite elements are sufficient to reach or approach chemical accuracy relative to higher-order calculations for the representative systems considered. The resulting corrected Hamiltonian also improves several lower lying valence states relative to LDA, although the improvement becomes less systematic away from the HOMO. This work provides accurate LDA-1/2 benchmark data for molecular systems and establishes a rigorous finite-element foundation for future molecular GW calculations.

cond-mat.mtrl-sci

Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism

LLM serving is increasingly multi-tenant: the same deployment must handle latency-critical interactive requests and more relaxed background workloads under a fixed GPU budget. This creates a tiered-SLO setting where maximizing overall goodput (requests that satisfy both TTFT and TPOT targets) is challenging because workload mix, request lengths, and load intensity vary over time. Existing systems mainly optimize request-level controls (e.g., queuing and batching) while keeping execution configuration largely static, which limits adaptation under multi-tier contention. We present Nitsum, a distributed LLM serving system that treats tensor parallelism (TP) as a first-class runtime control surface rather than a static deployment choice. Nitsum jointly optimizes TP level, prefill/decode GPU split, and request scheduling. To make frequent TP adaptation practical, Nitsum introduces TP-aware weight reuse and fast KV migration. Experiments on real traces and targeted microbenchmarks show that Nitsum improves SLO-compliant goodput over SoTA by up to 5.3 times.

cs.DC

QAssemble: A Pure Python Package for Quantum Many-Body Theory

QAssemble is a pure-Python package for the quantum many-body problem. It implements various functional approaches, such as tight-binding, Hartree-Fock, and GW approximations within a unified object-oriented architecture. Each physical concept--crystal structure, Hamiltonian, Green's function, self-energy, polarizability, screened Coulomb interaction--is represented as a distinct class. The modular design prioritizes code clarity and extensibility, leveraging NumPy, SciPy, and libdlr for numerical operations. Performance-critical kernels, including the polarizability bubble, Dyson equation inversion, and lattice Fourier transforms, are systematically vectorized and combined with the discrete Lehmann representation to achieve practical efficiency within a pure-Python environment. We validate QAssemble on the electronic structure of graphene with local and non-local interactions. Furthermore, benchmarks on a five-orbital extended Hund-Hubbard model demonstrate that this strategy delivers up to a 60x speedup over traditional loop-based Matsubara implementations. QAssemble supports both batch execution for production calculations and interactive workflows for method development.

cond-mat.str-el

On the Secrecy Performance of Continuous-Aperture Arrays Over Fading Channels

The secrecy performance of continuous-aperture array (CAPA)-based wiretap channels in terms of secrecy rate and secrecy outage probability (SOP) is analyzed. First, the system models of CAPA systems with maximum-ratio transmission under a Rayleigh fading channel are established, and approximate probability density functions for the legitimate user Bob's signal-to-noise ratio (SNR) and the eavesdropper Eve's SNR are derived using Mercer's theorem and Landau's eigenvalue theorem. Three scenarios are considered, including a single Eve, multiple independent Eves, and multiple collaborative Eves. Next, the expressions of the secrecy rate and SOP under these three scenarios are derived, and the high-SNR slope, high-SNR power offset, diversity order, and array gain in Bob's high-SNR region are obtained. It is then theoretically proven that, in all three scenarios, the CAPA system achieves the same high-SNR slope and the same diversity order, with the latter being equal to the spatial degrees of freedom. Moreover, the CAPA system with a single Eve has the smallest high-SNR offset and the highest array gain, whereas the CAPA system with multiple collaborative Eves exhibits the largest high-SNR offset and the lowest array gain. Finally, the theoretical analyses of secrecy rate, SOP, high-SNR performance are validated by the simulation results, and a higher secrecy rate and a lower SOP are achieved by the CAPA systems compared to the spatially-discrete array systems with half-wavelength antenna spacing.

cs.IT

Silhouette Score Efficient Radio Frequency Fingerprint Feature Extraction

Radio frequency fingerprint (RFF) identification technology, which exploits relatively stable hardware imperfections, is highly susceptible to constantly changing channel effects. Although various channel-robust RFF feature extraction methods have been proposed, they predominantly rely on experimental comparisons rather than theoretical analyses. This limitation hinders the progress of channel-robust RFF feature extraction and impedes the establishment of theoretical guidance for its design. In this paper, we establish a unified theoretical performance analysis framework for different RFF feature extraction methods using the silhouette score as an evaluation metric, and propose a precoding-based channel-robust RFF feature extraction method that enhances the silhouette score without requiring channel estimation. First, we employ the silhouette score as an evaluation metric and obtain the theoretical performance of various RFF feature extraction methods using the Taylor series expansion. Next, we mitigate channel effects by computing the reciprocal of the received signal in the frequency domain at the device under authentication. We then compare these methods across three different scenarios: the deterministic channel scenario, the independent and identically distributed (i.i.d.) stochastic channel scenario, and the non-i.i.d. stochastic channel scenario. Finally, simulation and experimental results demonstrate that the silhouette score is an efficient metric to evaluate classification accuracy. Furthermore, the results indicate that the proposed precoding-based channel-robust RFF feature extraction method achieves the highest silhouette score and classification accuracy under channel variations.

eess.SP

Division-based Receiver-agnostic RFF Identification in WiFi Systems

In physical-layer security schemes, radio frequency fingerprint (RFF) identification of WiFi devices is susceptible to receiver differences, which can significantly degrade classification performance when a model is trained on one receiver but tested on another. In this paper, we propose a division-based receiver-agnostic RFF extraction method for WiFi systems, which removes the receivers' effects by dividing different preambles in the frequency domain. The proposed method requires only a single receiver for training and does not rely on additional calibration or stacking processes. First, for flat fading channel scenarios, the legacy short training field (L-STF) and legacy long training field (L-LTF) of the unknown device are divided by those of the reference device in the frequency domain. The receiver-dependent effects can be eliminated with the requirement of only a single receiver for training, and the higher-dimensional RFF features can be extracted. Second, for frequency-selective fading channel scenarios, the high-throughput long training field (HT-LTF) is divided by the L-LTF in the frequency domain. Only a single receiver is required for training and the higher-dimensional RFF features that are both channel-invariant and receiver-agnostic are extracted. Finally, simulation and experimental results demonstrate that the proposed method effectively mitigate the impacts of channel variations and receiver differences. The classification results show that, even when training on a single receiver and testing on a different one, the proposed method achieves classification accuracy improvements of 15.5% and 28.45% over the state-of-the-art approach in flat fading and frequency-selective fading channel scenarios, respectively.

eess.SP

The magnetar model's energy crisis for a prolific repeating fast radio burst source

Fast radio bursts (FRBs) are widely considered to originate from magnetars that power the explosion through releasing magnetic energy. Active repeating FRBs have been seen to produce hundreds of bursts per hour and can stay active for months, thus may provide stringent constraints on the energy budget of FRBs' central engine. Within a time span of 214 days, we detected 11,553 bursts from the hyper-active FRB 20240114A that reached a peak burst rate of 729 hr$^{-1}$. This is the largest burst sample from any single FRB source, exceeding the cumulative total of all published bursts from all known FRBs to date. Assuming typical values of radio efficiency and beaming factor, the estimated total isotropic burst energy of this source exceeds 86% of the dipolar magnetic energy of a typical magnetar. The total released energy from this source exceeds that of other known repeaters by about one and a half orders of magnitude, yielding the most stringent lower limit of $4.7\times10^{32}$ G cm$^3$ for the magnetar's magnetic moment. The source remained active at the end of this observation campaign. Our findings thus require either the FRB's central magnetar engine's possessing exceptionally high emission efficiency or a more powerful compact object than a typical magnetar.

astro-ph.HE

A Reference Architecture for Autonomous Networks: An Agent-Based Approach

The vision of autonomous systems is becoming increasingly important in many application areas, where the aim is to replace humans with agents. These include autonomous vehicles and other agents' applications in business processes and problem-solving. For networks, the increasing scale and operation and management (O&M) complexity drive the need for autonomous networks (AN). The technical objective of AN is to ensure trustworthy O&M without human intervention for higher efficiency and lower operating costs. However, realizing AN seems more difficult than autonomous vehicles. It encounters challenges of networks' structural and functional complexity, which operate as distributed dynamic systems governed by various technical and economic constraints. A key problem lies in formulating a rigorous development methodology that facilitates a seamless transition from traditional networks to AN. Central to this methodology is the definition of a reference architecture for network agents, which specifies the required functionalities for their realization, regardless of implementation choices. This article proposes a reference architecture characterizing main functional features, illustrating its application with network use cases. It shows how artificial intelligence components can be used to implement the required functionality and its coordination. The latter is achieved through the management and generation of shared domain-specific knowledge stored in long-term memory, ensuring the overall consistency of decisions and their execution. The article concludes with a discussion of architecture specialization for building network layer agents. It also identifies the main technical challenges ahead, such as satisfying essential requirements at development or runtime, as well as the issue of coordinating agents to achieve collective intelligence in meeting overall network goals.

cs.NI

FEAST nonlinear eigenvalue algorithm for $GW$ quasiparticle equations

The use of Green's function in quantum many-body theory often leads to nonlinear eigenvalue problems, as Green's function needs to be defined in energy domain. The $GW$ approximation method is one of the typical examples. In this article, we introduce a method based on the FEAST eigenvalue algorithm for accurately solving the nonlinear eigenvalue $G_0W_0$ quasiparticle equation, eliminating the need for the Kohn-Sham wavefunction approximation. Based on the contour integral method for nonlinear eigenvalue problem, the energy (eigenvalue) domain is extended to complex plane. Hypercomplex number is introduced to the contour deformation calculation of $GW$ self-energy to carry imaginary parts of both Green's functions and FEAST quadrature nodes. Calculation results for various molecules are presented and compared with a more conventional graphical solution approximation method. It is confirmed that the Highest Occupied Molecular Orbital (HOMO) from the Kohn-Sham equation is very close to that of $GW$, while the Least Unoccupied Molecular Orbital (LUMO) shows noticeable differences.

physics.comp-ph

Methodological Explainability Evaluation of an Interpretable Deep Learning Model for Post-Hepatectomy Liver Failure Prediction Incorporating Counterfactual Explanations and Layerwise Relevance Propagation: A Prospective In Silico Trial

Artificial intelligence (AI)-based decision support systems have demonstrated value in predicting post-hepatectomy liver failure (PHLF) in hepatocellular carcinoma (HCC). However, they often lack transparency, and the impact of model explanations on clinicians' decisions has not been thoroughly evaluated. Building on prior research, we developed a variational autoencoder-multilayer perceptron (VAE-MLP) model for preoperative PHLF prediction. This model integrated counterfactuals and layerwise relevance propagation (LRP) to provide insights into its decision-making mechanism. Additionally, we proposed a methodological framework for evaluating the explainability of AI systems. This framework includes qualitative and quantitative assessments of explanations against recognized biomarkers, usability evaluations, and an in silico clinical trial. Our evaluations demonstrated that the model's explanation correlated with established biomarkers and exhibited high usability at both the case and system levels. Furthermore, results from the three-track in silico clinical trial showed that clinicians' prediction accuracy and confidence increased when AI explanations were provided.

cs.CV

Preble: Efficient Distributed Prompt Scheduling for LLM Serving

Prompts to large language models (LLMs) have evolved beyond simple user questions. For LLMs to solve complex problems, today's practices are to include domain-specific instructions, illustration of tool usages, and/or long context such as textbook chapters in prompts. As such, many parts of prompts are repetitive across requests. Recent works propose to cache and reuse KV state of prompts. However, they are all confined to a single-GPU optimization, while production LLM serving systems are distributed by nature. This paper proposes Preble, the first distributed LLM serving platform that targets and optimizes for prompt sharing. We designed a distributed scheduling system that co-optimizes KV state reuse and computation load-balancing with a new scheduling algorithm and a hierarchical scheduling mechanism. Our evaluation of Preble with real workloads and request arrival patterns on two open-source LLMs shows that Preble outperforms the SOTA serving systems by 1.5X to 14.5X on average latency and 2X to 10X on p99 latency.

cs.DC

A method of calculating bandstructure in real-space with application to all-electron and full potential

We introduce a practical and efficient approach for calculating the all-electron full potential bandstructure in real space, employing a finite element basis. As an alternative to the k-space method, the method involves the self-consistent solution of the Kohn-Sham equation within a larger finite system that encloses the unit-cell. It is based on the fact that the net potential of the unit-cell converges at a certain radius point. Bandstructure results are then obtained by performing non-self-consistent calculations in the Brillouin zone. Numerous numerical experiments demonstrate that the obtained valence and conduction bands are in excellent agreement with the pseudopotential k-space method. Moreover, we successfully observe the band bending of core electrons.

cond-mat.mtrl-sci

Insight from NLP Analysis: COVID-19 Vaccines Sentiments on Social Media

Social media is an appropriate source for analyzing public attitudes towards the COVID-19 vaccine and various brands. Nevertheless, there are few relevant studies. In the research, we collected tweet posts by the UK and US residents from the Twitter API during the pandemic and designed experiments to answer three main questions concerning vaccination. To get the dominant sentiment of the civics, we performed sentiment analysis by VADER and proposed a new method that can count the individual's influence. This allows us to go a step further in sentiment analysis and explain some of the fluctuations in the data changing. The results indicated that celebrities could lead the opinion shift on social media in vaccination progress. Moreover, at the peak, nearly 40\% of the population in both countries have a negative attitude towards COVID-19 vaccines. Besides, we investigated how people's opinions toward different vaccine brands are. We found that the Pfizer vaccine enjoys the most popular among people. By applying the sentiment analysis tool, we discovered most people hold positive views toward the COVID-19 vaccine manufactured by most brands. In the end, we carried out topic modelling by using the LDA model. We found residents in the two countries are willing to share their views and feelings concerning the vaccine. Several death cases have occurred after vaccination. Due to these negative events, US residents are more worried about the side effects and safety of the vaccine.

cs.CL