SearcharxivSearch

arXiv subjects

Shan Zhou

Publications and source records attributed to Shan Zhou.

At least 19 recordsLinked to original sources

Liquid structure adjacent to solid surfaces follows the superposition principle

Liquid structure at solid-liquid interfaces is critical for many natural and engineered processes ranging from biological signal transduction to electrochemical energy conversion. Advanced experimental and computational methods have provided insights into the structure of liquids adjacent to planar substrates at the nanoscale. However, realistic solid-liquid interfaces are inevitably inhomogeneous across multiple length scales, presenting a complexity that surpasses the capabilities of existing approaches. Here we bridge the complexity gap by discovering and utilizing a hitherto hidden principle of interfacial liquid--superposition. Experimentally, we use 3D atomic force microscopy (3D-AFM) to image the interfacial structure of a wide range of organic and aqueous solvents and electrolytes, uncovering universal liquid density oscillations and emergent liquid layer reconfigurations at heterogeneous substrate sites. We further develop an analytical model, coined solid-liquid superposition (SLS), which solves the interfacial liquid density distribution based on a key descriptor: the effective total correlation function (ETCF) between a liquid molecule and nearby solid atoms. SLS not only explains all the experimentally observed interfacial liquid distribution profiles from the angstrom to near-micron scale, but also predicts more precise atomic-scale interference patterns which are further corroborated by molecular dynamics (MD) simulations. This study unveils a key structural descriptor of interfacial liquids, and establishes a theoretical framework for rapidly and accurately predicting liquid structures adjacent to solid surfaces with arbitrary morphology and size scale.

physics.chem-ph

Valence-free open nanoparticle superlattices

A cornerstone of advanced materials design is establishing a framework for assembling nanoparticle superstructures with tailored symmetries. A longstanding challenge has been assembling diamond-like superstructures for photonic devices. Traditionally, such open superstructures require functionalized nanoparticles with directional or anisotropic interactions, reminiscent of valence bonding in a diamond. Here, we present a robust strategy for assembling valence-free nanoparticles into a broad array of cubic superstructures. By grafting nanoparticles with oppositely charged, end-functionalized water-soluble polymers of adjustable molecular weight, we gain control over electrostatic interactions and conformational constraints. This unified approach yields lattices analogous to rock salt, CsCl, zinc-blende, diamond, and the rare simple cubic phase, with tunable lattice constants. Theoretical models and simulations elucidate the underlying interactions, providing a framework for engineering valence-free nanoparticle superlattices.

cond-mat.mtrl-sci

Correlative angstrom-scale microscopy and spectroscopy of graphite-water interfaces

Water at solid surfaces is key for many processes ranging from biological signal transduction to membrane separation and renewable energy conversion. However, under realistic conditions, which often include environmental and surface charge variations, the interfacial water structure remains elusive. Here we overcome this limit by combining three-dimensional atomic force microscopy and interface-sensitive Raman spectroscopy to characterize the graphite-water interfacial structure in situ. Through correlative analysis of the spatial liquid density maps and vibrational peaks within ~2 nm of the graphite surface, we find the existence of two interfacial configurations at open circuit potential, a transient state where pristine water exhibits strong hydrogen bond (HB) breaking effects, and a steady state with hydrocarbons dominating the interface and weak HB breaking in the surrounding water. At sufficiently negative potentials, both states transition into a stable structure featuring pristine water with a broader distribution of HB configurations. Our three-state model resolves many long-standing controversies on interfacial water structure.

physics.chem-ph

LLMCO2: Advancing Accurate Carbon Footprint Prediction for LLM Inferences

Throughout its lifecycle, a large language model (LLM) generates a substantially larger carbon footprint during inference than training. LLM inference requests vary in batch size, prompt length, and token generation number, while cloud providers employ different GPU types and quantities to meet diverse service-level objectives for accuracy and latency. It is crucial for both users and cloud providers to have a tool that quickly and accurately estimates the carbon impact of LLM inferences based on a combination of inference request and hardware configurations before execution. Estimating the carbon footprint of LLM inferences is more complex than training due to lower and highly variable model FLOPS utilization, rendering previous equation-based models inaccurate. Additionally, existing machine learning (ML) prediction methods either lack accuracy or demand extensive training data, as they inadequately handle the distinct prefill and decode phases, overlook hardware-specific features, and inefficiently sample uncommon inference configurations. We introduce \coo, a graph neural network (GNN)-based model that greatly improves the accuracy of LLM inference carbon footprint predictions compared to previous methods.

cs.LG

Bending, breaking, and reconnecting of the electrical double layers at heterogeneous electrodes

In electrochemical systems, the structure of electrical double layers (EDLs) near electrode surfaces is crucial for energy conversion and storage functions. While the electrodes in real-world systems are usually heterogeneous, to date the investigation of EDLs is mainly limited to flat model solid surfaces. To bridge this gap, here we image the EDL structure of an ionic liquid-based electrolyte at a heterogeneous graphite electrode using our recently developed electrochemical 3D atomic force microscopy. These interfaces feature the formation of thin, nanoscale adlayer/cluster domains that closely mimic the early-stage solid-electrolyte interphases in many battery systems. We observe multiple discrete layers in the EDL near the flat electrode, which restructures at the heterogeneous interphase sites. Depending on the local size of the interphase clusters, the EDLs exhibit bending, breaking, and/or reconnecting behaviors, likely due to the combined steric and long-range interaction effects. These results shed light on the fundamental structure and reconfiguration mechanism of EDLs at heterogeneous interfaces.

physics.chem-ph

Inference Performance Optimization for Large Language Models on CPUs

Large language models (LLMs) have shown exceptional performance and vast potential across diverse tasks. However, the deployment of LLMs with high performance in low-resource environments has garnered significant attention in the industry. When GPU hardware resources are limited, we can explore alternative options on CPUs. To mitigate the financial burden and alleviate constraints imposed by hardware resources, optimizing inference performance is necessary. In this paper, we introduce an easily deployable inference performance optimization solution aimed at accelerating LLMs on CPUs. In this solution, we implement an effective way to reduce the KV cache size while ensuring precision. We propose a distributed inference optimization approach and implement it based on oneAPI Collective Communications Library. Furthermore, we propose optimization approaches for LLMs on CPU, and conduct tailored optimizations for the most commonly used models. The code is open-sourced at https://github.com/intel/xFasterTransformer.

cs.AI

Distributed Inference Performance Optimization for LLMs on CPUs

Large language models (LLMs) hold tremendous potential for addressing numerous real-world challenges, yet they typically demand significant computational resources and memory. Deploying LLMs onto a resource-limited hardware device with restricted memory capacity presents considerable challenges. Distributed computing emerges as a prevalent strategy to mitigate single-node memory constraints and expedite LLM inference performance. To reduce the hardware limitation burden, we proposed an efficient distributed inference optimization solution for LLMs on CPUs. We conduct experiments with the proposed solution on 5th Gen Intel Xeon Scalable Processors, and the result shows the time per output token for the LLM with 72B parameter is 140 ms/token, much faster than the average human reading speed about 200ms per token.

cs.DC

Metastable Cation-Disordered Niobium Tungsten Oxides as Li-ion Battery Anode Materials

Metastable cation-disordered compounds have greatly expanded the synthesizable compositions of solid-state materials and drawn sharp attention among battery electrochemists. While such a strategy has been very successful in a few well-known structures, such as rock salts, metastable cation-disordered materials for other structural types, especially for non-close packed structures, are peculiarly underexplored. In this work, we develop a new series of fully cation-disordered metastable niobium tungsten oxides with a simple structure, and name this new structural type anti-Li3N. Furthermore, we find that metastable anti-Li3N NWOs transform to a cation-disordered cubic structure when applied as a Li-ion battery anode, highlighting an intriguing non-close packed to close packed conversion between two cation-disordered phases, as evidenced in various physicochemical characterizations, in terms of diffraction, electronic, and vibrational structures. This work enriches the structural and compositional space of niobium tungsten oxide families, cation-disordered solid-state materials, and the working mechanisms of Li-ion battery anodes.

cond-mat.mtrl-sci

Preventing Discriminatory Decision-making in Evolving Data Streams

Bias in machine learning has rightly received significant attention over the last decade. However, most fair machine learning (fair-ML) work to address bias in decision-making systems has focused solely on the offline setting. Despite the wide prevalence of online systems in the real world, work on identifying and correcting bias in the online setting is severely lacking. The unique challenges of the online environment make addressing bias more difficult than in the offline setting. First, Streaming Machine Learning (SML) algorithms must deal with the constantly evolving real-time data stream. Second, they need to adapt to changing data distributions (concept drift) to make accurate predictions on new incoming data. Adding fairness constraints to this already complicated task is not straightforward. In this work, we focus on the challenges of achieving fairness in biased data streams while accounting for the presence of concept drift, accessing one sample at a time. We present Fair Sampling over Stream ($FS^2$), a novel fair rebalancing approach capable of being integrated with SML classification algorithms. Furthermore, we devise the first unified performance-fairness metric, Fairness Bonded Utility (FBU), to evaluate and compare the trade-off between performance and fairness of different bias mitigation methods efficiently. FBU simplifies the comparison of fairness-performance trade-offs of multiple techniques through one unified and intuitive evaluation, allowing model designers to easily choose a technique. Overall, extensive evaluations show our measures surpass those of other fair online techniques previously reported in the literature.

cs.LG

Chiral Assemblies of Pinwheel Superlattices on Substrates

The unique topology and physics of chiral superlattices make their self-assembly from nanoparticles a holy grail for (meta)materials. Here we show that tetrahedral gold nanoparticles can spontaneously transform from a perovskite-like low-density phase with corner-to-corner connections into pinwheel assemblies with corner-to-edge connections and denser packing. While the corner-sharing assemblies are achiral, pinwheel superlattices become strongly mirror-asymmetric on solid substrates as demonstrated by chirality measures. Liquid-phase transmission electron microscopy and computational models show that van der Waals and electrostatic interactions between nanoparticles control thermodynamic equilibrium. Variable corner-to-edge connections among tetrahedra enable fine-tuning of chirality. The domains of the bilayer superlattices display strong chiroptical activity identified by photon-induced near-field electron microscopy and finite-difference time-domain simulations. The simplicity and versatility of the substrate-supported chiral superlattices facilitate manufacturing of metastructured coatings with unusual optical, mechanical and electronic characteristics.

cond-mat.mtrl-sci

The center of monoidal 2-categories in 3+1D Dijkgraaf-Witten Theory

In this work, for a finite group $G$ and a 4-cocycle $ω\in Z^4(G, \mathbf{k}^\times)$, we compute explicitly the center of the monoidal 2-category $\operatorname{2Vec}_G^ω$ of $ω$-twisted $G$-graded 1-categories of finite dimensional $\mathbf{k}$-vector spaces. This center gives a precise mathematical description of the topological defects in the associated 3+1D Dijkgraaf-Witten TQFT. We prove that this center is a braided monoidal 2-category with a trivial sylleptic center.

math.QA

Subleading Microstate Counting in the Dual to Massive Type IIA

We study the topologically twisted index of a certain Chern-Simons matter theory with $SU(N)$ level $k$ gauge group on a genus $g$ Riemann surface times a circle. For this theory it is known that the logarithm of the topologically twisted index grows as $N^{5/3}$ and that it matches the Bekenstein-Hawking entropy of certain magnetically charged asymptotically $AdS_4\times S^6$ black holes in massive type IIA supergravity. Through a combination of numerical and analytical techniques we study the subleading in $N$ structure. We demonstrate precise analytic cancellation of terms of orders $N\log\,N$ and $N^{1/3}\log N$ and show numerical cancellation for terms of order $N$. As a result, the first subleading correction is of order $N^{2/3}$. Furthermore, we provide evidence for the presence of a term of the form $(g-1)(7/18) \log \,N$ which constitutes a microscopic prediction for the one-loop contribution coming from the massless gravitational degrees of freedom in the massive IIA black hole.

hep-th

Deploying Deep Ranking Models for Search Verticals

In this paper, we present an architecture executing a complex machine learning model such as a neural network capturing semantic similarity between a query and a document; and deploy to a real-world production system serving 500M+users. We present the challenges that arise in a real-world system and how we solve them. We demonstrate that our architecture provides competitive modeling capability without any significant performance impact to the system in terms of latency. Our modular solution and insights can be used by other real-world search systems to realize and productionize recent gains in neural networks.

cs.IR

Comments on Higher Rank Wilson Loops in ${\cal N}=2^*$

For ${\cal N}=2^*$ theory with $U(N)$ gauge group we evaluate expectation values of Wilson loops in representations described by a rectangular Young tableau with $n$ rows and $k$ columns. The evaluation reduces to a two-matrix model and we explain, using a combination of numerical and analytical techniques, the general properties of the eigenvalue distributions in various regimes of parameters $(N,λ,n,k)$ where $λ$ is the 't Hooft coupling. In the large $N$ limit we present analytic results for the leading and sub-leading contributions. In the particular cases of only one row or one column we reproduce previously known results for the totally symmetry and totally antisymmetric representations. We also extensively discusss the ${\cal N}=4$ limit of the ${\cal N}=2^*$ theory. While establishing these connections we clarify aspects of various orders of limits and how to relax them; we also find it useful to explicitly address details of the genus expansion. As a result, for the totally symmetric Wilson loop we find new contributions that improve the comparison with the dual holographic computation at one loop order in the appropriate regime.

hep-th

Reliable and robust entanglement witness

Entanglement, a critical resource for quantum information processing, needs to be witnessed in many practical scenarios. Theoretically, witnessing entanglement is by measuring a special Hermitian observable, called entanglement witness (EW), which has non-negative expected outcomes for all separable states but can have negative expectations for certain entangled states. In practice, an EW implementation may suffer from two problems. The first one is \emph{reliability}. Due to unreliable realization devices, a separable state could be falsely identified as an entangled one. The second problem relates to \emph{robustness}. A witness may not to optimal for a target state and fail to identify its entanglement. To overcome the reliability problem, we employ a recently proposed measurement-device-independent entanglement witness, in which the correctness of the conclusion is independent of the implemented measurement devices. In order to overcome the robustness problem, we optimize the EW to draw a better conclusion given certain experimental data. With the proposed EW scheme, where only data postprocessing needs to be modified comparing to the original measurement-device-independent scheme, one can efficiently take advantage of the measurement results to maximally draw reliable conclusions.

quant-ph

Distributed Power Control and Coding-Modulation Adaptation in Wireless Networks using Annealed Gibbs Sampling

In wireless networks, the transmission rate of a link is determined by received signal strength, interference from simultaneous transmissions, and available coding-modulation schemes. Rate allocation is a key problem in wireless network design, but a very challenging problem because: (i) wireless interference is global, i.e., a transmission interferes all other simultaneous transmissions, and (ii) the rate-power relation is non-convex and non-continuous, where the discontinuity is due to limited number of coding-modulation choices in practical systems. In this paper, we propose a distributed power control and coding-modulation adaptation algorithm using annealed Gibbs sampling, which achieves throughput optimality in an arbitrary network topology. We consider a realistic Signal-to-Interference-and-Noise-Ratio (SINR) based interference model, and assume continuous power space and finite rate options (coding-modulation choices). Our algorithm first decomposes network-wide interference to local interference by properly choosing a "neighborhood" for each transmitter and bounding the interference from non-neighbor nodes. The power update policy is then carefully designed to emulate a Gibbs sampler over a Markov chain with a continuous state space. We further exploit the technique of simulated annealing to speed up the convergence of the algorithm to the optimal power and coding-modulation configuration. Finally, simulation results demonstrate the superior performance of the proposed algorithm.

cs.NI

On Delay Constrained Multicast Capacity of Large-Scale Mobile Ad-Hoc Networks

This paper studies the delay constrained multicast capacity of large scale mobile ad hoc networks (MANETs). We consider a MANET consists of $n_s$ multicast sessions. Each multicast session has one source and $p$ destinations. The wireless mobiles move according to a two-dimensional i.i.d. mobility model. Each source sends identical information to the $p$ destinations in its multicast session, and the information is required to be delivered to all the $p$ destinations within $D$ time-slots. Given the delay constraint $D,$ we first prove that the capacity per multicast session is $O(\min\{1, (\log p)(\log (n_sp)) \sqrt{\frac{D}{n_s}}\}).$ Given non-negative functions $f(n)$ and $g(n)$: $f(n)=O(g(n))$ means there exist positive constants $c$ and $m$ such that $f(n) \leq cg(n)$ for all $ n\geq m;$ $f(n)=Ω(g(n))$ means there exist positive constants $c$ and $m$ such that $f(n)\geq cg(n)$ for all $n\geq m;$ $f(n)=Θ(g(n))$ means that both $f(n)=Ω(g(n))$ and $f(n)=O(g(n))$ hold; $f(n)=o(g(n))$ means that $\lim_{n\to \infty} f(n)/g(n)=0;$ and $f(n)=ω(g(n))$ means that $\lim_{n\to \infty} g(n)/f(n)=0.$ We then propose a joint coding/scheduling algorithm achieving a throughput of $Θ(\min\{1,\sqrt{\frac{D}{n_s}}\}).$ Our simulations show that the joint coding/scheduling algorithm achieves a throughput of the same order ($Θ(\min\{1, \sqrt{\frac{D}{n_s}}\})$) under random walk model and random waypoint model.

cs.NI

On the nature of the $π_2(1880)

The strong decays of the $π_2(1880)$ as the $2 ^1D_2$ quark-antiquark state are investigated in the $^3P_0$ model and the flux-tube model, respectively. The results are similar in the two models. It is found that the decay patterns of the conventional $2 ^1D_2$ meson and the $2^{-+}$ light hybrid are very different, and the experimental evidence for the $π_2(1880)$ is consistent with it being the conventional $2 ^1D_2$ meson rather than the $2^{-+}$ light hybrid. The possibility of the $π_2(1880)$ being a mixture of the conventional $q\bar{q}$ and the hybrid is discussed.

hep-ph