SearcharxivSearch

arXiv subjects

Zhenlin Wang

Publications and source records attributed to Zhenlin Wang.

At least 19 recordsLinked to original sources

Optically locked low-noise photonic microwave oscillator

The next-generation sensing and communication applications rely on high-frequency microwave generation with low-noise. The microwave photonic technology is promising by the practical application is limited by its complex architecture so far. Here, we demonstrate an optically locked low-noise photonic microwave oscillator, so that all the optical components are packaged within a small module of 166 mL, and low noise microwave generation is achieved at 10.4 GHz with single-sideband phase noise of -54 dBc/Hz at 10 Hz, -141 dBc/Hz at 10 kHz, and -162 dBc/Hz at 10 MHz offset. Above performance arises from a dual-laser self-injection-locking scheme to a single Fabry-Perot cavity with high Q exceeding 10^8, with over 20 dB common-mode noise suppression. The low-noise nature of such reference is coherently transferred to the X-band through a high-performance TFLN electro-optic comb chip, thereby overcoming long-standing barriers in photonic microwave integration to enable truly field-deployable low-noise microwave generation.

physics.optics

Scalable Generalized Meta-Spanners Enabling Parallel Multitasking Optical Manipulation

Optical manipulation techniques offer exceptional contactless control but are fundamentally limited in their ability to perform parallel multitasking. To achieve high-density, versatile manipulation with subwavelength photonic devices, it is essential to sculpt light fields in multiple dimensions. Here, we overcome this challenge by introducing generalized optical meta-spanners (GOMSs) based on metasurfaces. Relying on complex-amplitude modulation, this platform generates lens-free, customizable optical fields that suppress diffractive losses. As a result, several advanced functionalities are simultaneously achieved, including longitudinally varying manipulation and in-plane spanner arrays, which outperforms the same operations realized by conventional donut-shaped orbital flows. Furthermore, the particle dynamics is reconfigurable simply by switching the input and output polarizations, facilitating robust multi-channel control. We experimentally validate the proposed approach by demonstrating single-particle dynamics and the parallel manipulation of particle ensembles, revealing exceptional stability for multitasking operations. These results demonstrate an ultracompact platform scalable to a much larger number of optical spanners, advancing metadevices from wavefront sculptors to particle manipulators. We envision that the GOMS will catalyze innovations in cross-disciplinary fields such as targeted drug delivery and cell-level biomechanics.

physics.optics

XBOF: A Cost-Efficient CXL JBOF with Inter-SSD Compute Resource Sharing

Enterprise SSDs integrate numerous computing resources (e.g., ARM processor and onboard DRAM) to satisfy the ever-increasing performance requirements of I/O bursts. While these resources substantially elevate the monetary costs of SSDs, the sporadic nature of I/O bursts causes severe SSD resource underutilization in just a bunch of flash (JBOF) level. Tackling this challenge, we propose XBOF, a cost-efficient JBOF design, which only reserves moderate computing resources in SSDs at low monetary cost, while achieving demanded I/O performance through efficient inter-SSD resource sharing. Specifically, XBOF first disaggregates SSD architecture into multiple disjoint parts based on their functionality, enabling fine-grained SSD internal resource management. XBOF then employs a decentralized scheme to manage these disaggregated resources and harvests the computing resources of idle SSDs to assist busy SSDs in handling I/O bursts. This idea is facilitated by the cache-coherent capability of Compute eXpress Link (CXL), with which the busy SSDs can directly utilize the harvested computing resources to accelerate metadata processing. The evaluation results show that XBOF improves SSD resource utilization by 50.4% and saves 19.0% monetary costs with a negligible performance loss, compared to existing JBOF designs.

cs.OS

High-gain optical parametric amplification with a continuous-wave pump using a domain-engineered thin-film lithium niobate waveguide

While thin film lithium niobate (TFLN) is known for efficient signal generation, on-chip signal amplification remains challenging from fully integrated optical communication circuits. Here we demonstrate the continuous-wave-pump optical parametric amplification (OPA) using an x-cut domain-engineered TFLN waveguide, with high gain over the telecom band up to 13.9 dB, and test it for high signal-to-noise ratio signal amplification using a commercial optical communication module pair. Fabricated in wafer scale using common process as devices including modulators, this OPA device marks an important step in TFLN photonic integration.

physics.optics

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models

Fine-tuning plays a crucial role in adapting models to downstream tasks with minimal training efforts. However, the rapidly increasing size of foundation models poses a daunting challenge for accommodating foundation model fine-tuning in most commercial devices, which often have limited memory bandwidth. Techniques like model sharding and tensor parallelism address this issue by distributing computation across multiple devices to meet memory requirements. Nevertheless, these methods do not fully leverage their foundation nature in facilitating the fine-tuning process, resulting in high computational costs and imbalanced workloads. We introduce a novel Distributed Dynamic Fine-Tuning (D2FT) framework that strategically orchestrates operations across attention modules based on our observation that not all attention modules are necessary for forward and backward propagation in fine-tuning foundation models. Through three innovative selection strategies, D2FT significantly reduces the computational workload required for fine-tuning foundation models. Furthermore, D2FT addresses workload imbalances in distributed computing environments by optimizing these selection strategies via multiple knapsack optimization. Our experimental results demonstrate that the proposed D2FT framework reduces the training computational costs by 40% and training communication costs by 50% with only 1% to 2% accuracy drops on the CIFAR-10, CIFAR-100, and Stanford Cars datasets. Moreover, the results show that D2FT can be effectively extended to recent LoRA, a state-of-the-art parameter-efficient fine-tuning technique. By reducing 40% computational cost or 50% communication cost, D2FT LoRA top-1 accuracy only drops 4% to 6% on Stanford Cars dataset.

cs.LG

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees

Offloading large language models (LLMs) state to host memory during inference promises to reduce operational costs by supporting larger models, longer inputs, and larger batch sizes. However, the design of existing memory offloading mechanisms does not take latency service-level objectives (SLOs) into consideration. As a result, they either lead to frequent SLO violations or underutilize host memory, thereby incurring economic loss and thus defeating the purpose of memory offloading. This paper presents Select-N, a latency-SLO-aware memory offloading system for LLM serving. A key challenge in designing Select-N is to reconcile the tension between meeting SLOs and maximizing host memory usage. Select-N overcomes it by exploiting a unique characteristic of modern LLMs: during serving, the computation time of each decoder layer is deterministic. Leveraging this, Select-N introduces offloading interval, an internal tunable knob that captures the tradeoff between SLOs and host memory usage, thereby reducing the aforementioned challenge to pick an optimal offloading interval. With that, Select-N proposes a two-stage approach to automatically pick the offloading interval. The first stage is offline that generates the range of optimal offloading interval, while the second stage adjusts offloading interval at the granularity of inference iteration based on runtime hardware status. Our evaluation shows that Select-N consistently meets SLOs and improves the serving throughput over existing mechanisms by 1.85X due to maximizing the use of host memory.

cs.DC

Inference of weak-form partial differential equations describing migration and proliferation mechanisms in wound healing experiments on cancer cells

Targeting signaling pathways that drive cancer cell migration or proliferation is a common therapeutic approach. A popular experimental technique, the scratch assay, measures the migration and proliferation-driven cell closure of a defect in a confluent cell monolayer. These assays do not measure dynamic effects. To improve analysis of scratch assays, we combine high-throughput scratch assays, video microscopy, and system identification to infer partial differential equation (PDE) models of cell migration and proliferation. We capture the evolution of cell density fields over time using live cell microscopy and automated image processing. We employ weak form-based system identification techniques for cell density dynamics modeled with first-order kinetics of advection-diffusion-reaction systems. We present a comparison of our methods to results obtained using traditional inference approaches on previously analyzed 1-dimensional scratch assay data. We demonstrate the application of this pipeline on high throughput 2-dimensional scratch assays and find that low levels of trametinib inhibit wound closure primarily by decreasing random cell migration by approximately 20%. Our integrated experimental and computational pipeline can be adapted for quantitatively inferring the effect of biological perturbations on cell migration and proliferation in various cell lines.

q-bio.CB

Spectral-isolated photonic topological corner mode with a tunable mode area and stable frequency

Emergent collective modes in lattices give birth to many intriguing physical phenomena in condensed matter physics. Among these collective modes, large-area modes typically feature small-level spacings, while a mode with stable frequency tends to be spatially tightly confined. Here, we theoretically propose and experimentally demonstrate a spectral-isolated photonic topological corner mode with a tunable mode area and stable frequency in a two-dimensional photonic crystal. This mode emerges from hybridizing the large-area homogeneous mode and in-gap topological corner modes. Remarkably, this large-area homogeneous mode possesses unique chirality and has a tunable mode area under the change of the mass term of the inner topological non-trivial lattice. We experimentally observe such topological large-area corner modes(TLCM) in a 2D photonic system and demonstrate the robustness by introducing disorders in the structure. Our findings have propelled the forefront of higher-order topology research, transitioning it from single-lattice systems to multi-lattice systems. They may support promising potential applications, particularly in vertical-cavity surface-emitting lasers.

cond-mat.mes-hall

Characterizing Out-of-Distribution Error via Optimal Transport

Out-of-distribution (OOD) data poses serious challenges in deployed machine learning models, so methods of predicting a model's performance on OOD data without labels are important for machine learning safety. While a number of methods have been proposed by prior work, they often underestimate the actual error, sometimes by a large margin, which greatly impacts their applicability to real tasks. In this work, we identify pseudo-label shift, or the difference between the predicted and true OOD label distributions, as a key indicator to this underestimation. Based on this observation, we introduce a novel method for estimating model performance by leveraging optimal transport theory, Confidence Optimal Transport (COT), and show that it provably provides more robust error estimates in the presence of pseudo-label shift. Additionally, we introduce an empirically-motivated variant of COT, Confidence Optimal Transport with Thresholding (COTT), which applies thresholding to the individual transport costs and further improves the accuracy of COT's error estimates. We evaluate COT and COTT on a variety of standard benchmarks that induce various types of distribution shift -- synthetic, novel subpopulation, and natural -- and show that our approaches significantly outperform existing state-of-the-art methods with an up to 3x lower prediction error.

cs.LG

A treatment of particle-electrolyte sharp interface fracture in solid-state batteries with multi-field discontinuities

In this work, we present a computational framework for coupled electro-chemo-(nonlinear) mechanics at the particle scale for solid-state batteries. The framework accounts for interfacial fracture between the active particles and solid electrolyte due to intercalation stresses. We extend discontinuous finite element methods for a sharp interface treatment of discontinuities in concentrations, fluxes, electric fields and in displacements, the latter arising from active particle-solid electrolyte interface fracture. We model the degradation in the charge transfer process that results from the loss of contact due to fracture at the electrolyte-active particle interfaces. Additionally, we account for the stress-dependent kinetics that can influence the charge transfer reactions and solid state diffusion. The discontinuous finite element approach does not require a conformal mesh. This offers the flexibility to construct arbitrary particle shapes and geometries that are based on design, or are obtained from microscopy images. The finite element mesh, however, can remain Cartesian, and independent of the particle geometries. We demonstrate this computational framework on micro-structures that are representative of solid-sate batteries with single and multiple anode and cathode particles.

cond-mat.mtrl-sci

FHPM: Fine-grained Huge Page Management For Virtualization

As more data-intensive tasks with large footprints are deployed in virtual machines (VMs), huge pages are widely used to eliminate the increasing address translation overhead. However, once the huge page mapping is established, all the base page regions in the huge page share a single extended page table (EPT) entry, so that the hypervisor loses awareness of accesses to base page regions. None of the state-of-the-art solutions can obtain access information at base page granularity for huge pages. We observe that this can lead to incorrect decisions by the hypervisor, such as incorrect data placement in a tiered memory system and unshared base page regions when sharing pages. This paper proposes FHPM, a fine-grained huge page management for virtualization without hardware and guest OS modification. FHPM can identify access information at base page granularity, and dynamically promote and demote pages. A key insight of FHPM is to redirect the EPT huge page directory entries (PDEs) to new companion pages so that the MMU can track access information within huge pages. Then, FHPM can promote and demote pages according to the current hot page pressure to balance address translation overhead and memory usage. At the same time, FHPM proposes a VM-friendly page splitting and collapsing mechanism to avoid extra VM-exits. In combination, FHPM minimizes the monitoring and management overhead and ensures that the hypervisor gets fine-grained VM memory accesses to make the proper decision. We apply FHPM to improve tiered memory management (FHPM-TMM) and to promote page sharing (FHPM-Share). FHPM-TMM achieves a performance improvement of up to 33% and 61% over the pure huge page and base page management. FHPM-Share can save 41% more memory than Ingens, a state-of-the-art page sharing solution, with comparable performance.

cs.OS

Disentangled higher-orbital bands and chiral symmetric topology in confined Mie resonance photonic crystals

Topological phases based on tight-binding models have been extensively studied in recent decades. By mimicking the linear combination of atomic orbitals in tight-binding models based on the evanescent couplings between resonators in classical waves, numerous experimental demonstrations of topological phases have been successfully conducted. However, in dielectric photonic crystals, the Mie resonances' states decay too slowly as $1/r$ when $r$ $\to$ $\infty$, leading to intrinsically different physical properties between tight-binding models and dielectric photonic crystals. Here, we propose a confined Mie resonance photonic crystal by embedding perfect electric conductors in between dielectric rods, leading to a perfectly matched band structure as the tight-binding models with nearest-neighbour couplings. As a consequence, disentangled band structure spanned by higher atomic orbitals is observed. Moreover, we also achieve a three-dimensional photonic crystal with a complete photonic bandgap and third-order topology based on our design. Our implementation provides a versatile platform for studying exotic higher-orbital bands and achieving tight-binding-like 3D topological photonic crystals.

cond-mat.mes-hall

Predicting Out-of-Distribution Error with Confidence Optimal Transport

Out-of-distribution (OOD) data poses serious challenges in deployed machine learning models as even subtle changes could incur significant performance drops. Being able to estimate a model's performance on test data is important in practice as it indicates when to trust to model's decisions. We present a simple yet effective method to predict a model's performance on an unknown distribution without any addition annotation. Our approach is rooted in the Optimal Transport theory, viewing test samples' output softmax scores from deep neural networks as empirical samples from an unknown distribution. We show that our method, Confidence Optimal Transport (COT), provides robust estimates of a model's performance on a target domain. Despite its simplicity, our method achieves state-of-the-art results on three benchmark datasets and outperforms existing methods by a large margin.

cs.LG

Cascadable in-memory computing based on symmetric writing and read out

The building block of in-memory computing with spintronic devices is mainly based on the magnetic tunnel junction with perpendicular interfacial anisotropy (p-MTJ). The resulting asymmetric write and read-out operations impose challenges in downscaling and direct cascadability of p-MTJ devices. Here, we propose that a new symmetric write and read-out mechanism can be realized in perpendicular-anisotropy spin-orbit (PASO) quantum materials based on Fe3GeTe2 and WTe2. We demonstrate that field-free and deterministic reversal of the perpendicular magnetization can be achieved by employing unconventional charge to z-spin conversion. The resulting magnetic state can be readily probed with its intrinsic inverse process, i.e., z-spin to charge conversion. Using the PASO quantum material as a fundamental building block, we implement the functionally complete set of logic-in-memory operations and a more complex nonvolatile half-adder logic function. Our work highlights the potential of PASO quantum materials for the development of scalable energy-efficient and ultrafast spintronic computing.

cond-mat.mes-hall

HMM-V: Heterogeneous Memory Management for Virtualization

The memory demand of virtual machines (VMs) is increasing, while DRAM has limited capacity and high power consumption. Non-volatile memory (NVM) is an alternative to DRAM, but it has high latency and low bandwidth. We observe that the VM with heterogeneous memory may incur up to a $1.5\times$ slowdown compared to a DRAM VM, if not managed well. However, none of the state-of-the-art heterogeneous memory management designs are customized for virtualization on a real system. In this paper, we propose HMM-V, a Heterogeneous Memory Management system for Virtualization. HMM-V automatically determines page hotness and migrates pages between DRAM and NVM to achieve performance close to the DRAM system. First, HMM-V tracks memory accesses through page table manipulation, but reduces the cost by leveraging Intel page-modification logging (PML) and a multi-level queue. Second, HMM-V quantifies the ``temperature'' of page and determines the hot set with bucket-sorting. HMM-V then efficiently migrates pages with minimal access pause and handles dirty pages with the assistance of PML. Finally, HMM-V provides pooling management to balance precious DRAM across multiple VMs to maximize utilization and overall performance. HMM-V is implemented on a real system with Intel Optane DC persistent memory. The four-VM co-running results show that HMM-V outperforms NUMA balancing and hardware management (Intel Optane memory mode) by $51\%$ and $31\%$, respectively.

cs.OS

Bulk-LDOS Correspondence in Topological Insulators

Seeking the criterion for diagnosing topological phases in real materials has been one of the major tasks in topological physics. Currently, bulk-boundary correspondence based on spectral measurements of in gap topological boundary states and the fractional corner anomaly derived from the measurement of the fractional spectral charge are two main approaches to characterize topologically insulating phases. However, these two methods require a complete band-gap with either in-gap states or strict spatial symmetry of the overall sample which significantly limits their applications to more generalized cases. Here we propose and demonstrate an approach to link the non-trivial hierarchical bulk topology to the multidimensional partition of local-density of states (LDOS) respectively, denoted as the bulk-LDOS correspondence. Specifically, in a finite-size topologically nontrivial photonic crystal, we observe that the distribution of LDOS is divided into three partitioned regions of the sample - the two-dimensional interior bulk area (avoiding edge and corner areas), one-dimensional edge region (avoiding the corner area), and zero-dimensional corner sites. In contrast, the LDOS is distributed across the entire two-dimensional bulk area across the whole spectrum for the topologically trivial cases. Moreover, we present the universality of this criterion by validating this correspondence in both a higher-order topological insulator without a complete band gap and with disorders. Our findings provide a general way to distinguish topological insulators and unveil the unexplored features of topological directional band-gap materials without in-gap states.

cond-mat.mes-hall

Max-Min Grouped Bandits

In this paper, we introduce a multi-armed bandit problem termed max-min grouped bandits, in which the arms are arranged in possibly-overlapping groups, and the goal is to find the group whose worst arm has the highest mean reward. This problem is of interest in applications such as recommendation systems and resource allocation, and is also closely related to widely-studied robust optimization problems. We present two algorithms based successive elimination and robust optimization, and derive upper bounds on the number of samples to guarantee finding a max-min optimal or near-optimal group, as well as an algorithm-independent lower bound. We discuss the degree of tightness of our bounds in various cases of interest, and the difficulties in deriving uniformly tight bounds.

stat.ML

Best Arm Identification with Safety Constraints

The best arm identification problem in the multi-armed bandit setting is an excellent model of many real-world decision-making problems, yet it fails to capture the fact that in the real-world, safety constraints often must be met while learning. In this work we study the question of best-arm identification in safety-critical settings, where the goal of the agent is to find the best safe option out of many, while exploring in a way that guarantees certain, initially unknown safety constraints are met. We first analyze this problem in the setting where the reward and safety constraint takes a linear structure, and show nearly matching upper and lower bounds. We then analyze a much more general version of the problem where we only assume the reward and safety constraint can be modeled by monotonic functions, and propose an algorithm in this setting which is guaranteed to learn safely. We conclude with experimental results demonstrating the effectiveness of our approaches in scenarios such as safely identifying the best drug out of many in order to treat an illness.

cs.LG