SearcharxivSearch

arXiv subjects

Jinhong Li

Publications and source records attributed to Jinhong Li.

11 recordsLinked to original sources

Sufficient Dimension Reduction via Inverse Conditional Mean or Variance Independence

This paper presents a unified framework for sufficient dimension reduction (SDR) that generalizes several existing SDR techniques and offers new insights into the connection between inverse conditional moment independence and dimension reduction. The framework is built on two forms of inverse independence between the response vector and predictors: inverse conditional mean independence (ICMI) and inverse conditional variance independence (ICVI). For each form, we develop two general classes of matrices capable of recovering the central subspace, based on projection and kernel techniques respectively. This yields four distinct estimators: projection- and kernel-based variants under both ICMI and ICVI frameworks. Under standard regularity conditions, we establish the theoretical properties of these estimators and derive their convergence rates in high-dimensional settings. The proposed methods exhibit robustness to outliers in the response variable while maintaining computational competitiveness. Simulation studies and real-data analyses demonstrate the practical effectiveness of the proposed methods.

stat.ME

The Design and Implementation of a High-Performance Log-Structured RAID System for ZNS SSDs

Zoned Namespace (ZNS) defines a new abstraction for host software to flexibly manage storage in flash-based SSDs as append-only zones. It also provides a Zone Append primitive to further boost the write performance of ZNS SSDs by exploiting intra-zone parallelism. However, making Zone Append effective for reliable and scalable storage, in the form of a RAID array of multiple ZNS SSDs, is non-trivial, since Zone Append offloads address management to ZNS SSDs and requires hosts to specifically manage RAID stripes across multiple drives. We propose ZapRAID, a high-performance log-structured RAID system for ZNS SSDs by carefully exploiting Zone Append to achieve high write parallelism and lightweight stripe management. ZapRAID adopts a group-based data layout with a coarse-grained ordering across multiple groups of stripes, such that it can use small-size metadata for stripe management on a per-group basis under Zone Append. It further adopts hybrid data management to simultaneously achieve intra-zone and inter-zone parallelism through a careful combination of both Zone Write and Zone Append primitives. We implement ZapRAID as a user-space block device, and evaluate ZapRAID using microbenchmarks, trace-driven experiments, and real-application experiments. Our evaluation results show that ZapRAID achieves high write throughput and maintains high performance in normal reads, degraded reads, crash recovery, and full-drive recovery.

cs.DC

SID: A Novel Class of Nonparametric Tests of Independence for Censored Outcomes

We propose a new class of metrics, called the survival independence divergence (SID), to test dependence between a right-censored outcome and covariates. A key technique for deriving the SIDs is to use a counting process strategy, which equivalently transforms the intractable independence test due to the presence of censoring into a test problem for complete observations. The SIDs are equal to zero if and only if the right-censored response and covariates are independent, and they are capable of detecting various types of nonlinear dependence. We propose empirical estimates of the SIDs and establish their asymptotic properties. We further develop a wild bootstrap method to estimate the critical values and show the consistency of the bootstrap tests. The numerical studies demonstrate that our SID-based tests are highly competitive with existing methods in a wide range of settings.

stat.ME

An In-Depth Comparative Analysis of Cloud Block Storage Workloads: Findings and Implications

Cloud block storage systems support diverse types of applications in modern cloud services. Characterizing their I/O activities is critical for guiding better system designs and optimizations. In this paper, we present an in-depth comparative analysis of production cloud block storage workloads through the block-level I/O traces of billions of I/O requests collected from two production systems, Alibaba Cloud and Tencent Cloud Block Storage. We study their characteristics of load intensities, spatial patterns, and temporal patterns. We also compare the cloud block storage workloads with the notable public block-level I/O workloads from the enterprise data centers at Microsoft Research Cambridge, and identify the commonalities and differences of the three sources of traces. To this end, we provide 6 findings through the high-level analysis and 16 findings through the detailed analysis on load intensity, spatial patterns, and temporal patterns. We discuss the implications of our findings on load balancing, cache efficiency, and storage cluster management in cloud block storage systems.

cs.DC

Efficient LSM-Tree Key-Value Data Management on Hybrid SSD/HDD Zoned Storage

Zoned storage devices, such as zoned namespace (ZNS) solid-state drives (SSDs) and host-managed shingled magnetic recording (HM-SMR) hard-disk drives (HDDs), expose interfaces for host-level applications to support fine-grained, high-performance storage management. Combining ZNS SSDs and HM-SMR HDDs into a unified hybrid storage system is a natural direction to scale zoned storage at low cost, yet how to effectively incorporate zoned storage awareness into hybrid storage is a non-trivial issue. We make a case for key-value (KV) stores based on log-structured merge trees (LSM-trees) as host-level applications, and present HHZS, a middleware system that bridges an LSM-tree KV store with hybrid zoned storage devices based on hints. HHZS leverages hints issued by the flushing, compaction, and caching operations of the LSM-tree KV store to manage KV objects in placement, migration, and caching in hybrid ZNS SSD and HM-SMR HDD zoned storage. Experiments show that our HHZS prototype, when running on real ZNS SSD and HM-SMR HDD devices, achieves the highest throughput compared with all baselines under various settings.

cs.PF

Separating Data via Block Invalidation Time Inference for Write Amplification Reduction in Log-Structured Storage

Log-structured storage has been widely deployed in various domains of storage systems, yet its garbage collection incurs write amplification (WA) due to the rewrites of live data. We show that there exists an optimal data placement scheme that minimizes WA using the future knowledge of block invalidation time (BIT) of each written block, yet it is infeasible to realize in practice. We propose a novel data placement algorithm for reducing WA, SepBIT, that aims to infer the BITs of written blocks from storage workloads and separately place the blocks into groups with similar estimated BITs. We show via both mathematical and production trace analyses that SepBIT effectively infers the BITs by leveraging the write skewness property in practical storage workloads. Trace analysis and prototype experiments show that SepBIT reduces WA and improves I/O throughput, respectively, compared with state-of-the-art data placement schemes. SepBIT is currently deployed to support the log-structured block storage management at Alibaba Cloud.

cs.DC

Repair Pipelining for Erasure-Coded Storage: Algorithms and Evaluation

We propose repair pipelining, a technique that speeds up the repair performance in general erasure-coded storage. By carefully scheduling the repair of failed data in small-size units across storage nodes in a pipelined manner, repair pipelining reduces the single-block repair time to approximately the same as the normal read time for a single block in homogeneous environments. We further design different extensions of repair pipelining algorithms for heterogeneous environments and multi-block repair operations. We implement a repair pipelining prototype, called ECPipe, and integrate it as a middleware system into two versions of Hadoop Distributed File System (HDFS) (namely HDFS-RAID and HDFS-3) as well as Quantcast File System (QFS). Experiments on a local testbed and Amazon EC2 show that repair pipelining significantly improves the performance of degraded reads and full-node recovery over existing repair techniques.

cs.DC

On identifying magnetized anomalies using geomagnetic monitoring II. A Magnetohydrodynamic Model

This paper is a continuation and an extension of our recent work [13] on the identification of magnetized anomalies using geomagnetic monitoring, which aims to establish a rigorous mathematical theory for the geomagnetic detection technology. Suppose a collection of magnetized anomalies is presented in the shell of the Earth. By monitoring the variation of the magnetic field of the Earth due to the presence of the anomalies, we establish sufficient conditions for the unique recovery of those unknown anomalies. In [13], the geomagnetic model was described by a linear Maxwell system. In this paper, we consider a much more sophisticated and complicated magnetohydrodynamic model, which stems from the widely accepted dynamo theory of geomagnetics.

math.AP

An inverse scattering approach for geometric body generation: a machine learning perspective

In this paper, we are concerned with the 2D and 3D geometric shape generation by prescribing a set of characteristic values of a specific geometric body. One of the major motivations of our study is the 3D human body generation in various applications. We develop a novel method that can generate the desired body with customized characteristic values. The proposed method follows a machine-learning flavour that generates the inferred geometric body with the input characteristic parameters from a training dataset. One of the critical ingredients and novelties of our method is the borrowing of inverse scattering techniques in the theory of wave propagation to the body generation. This is done by establishing a delicate one-to-one correspondence between a geometric body and the far-field pattern of a source scattering problem governed by the Helmholtz system. It in turn enables us to establish a one-to-one correspondence between the geometric body space and the function space defined by the far-field patterns. Hence, the far-field patterns can act as the shape generators. The shape generation with prescribed characteristic parameters is achieved by first manipulating the shape generators and then reconstructing the corresponding geometric body from the obtained shape generator by a stable multiple-frequency Fourier method. Our method is easy to implement and produces more efficient and stable body generations. We provide both theoretical analysis and extensive numerical experiments for the proposed method. The study is the first attempt to introduce inverse scattering approaches in combination with machine learning to the geometric body generation and it opens up many opportunities for further developments.

cs.GR

On identifying magnetized anomalies using geomagnetic monitoring

We propose and investigate the inverse problem of identifying magnetized anomalies beneath the Earth using the geomagnetic monitoring. Suppose a collection of magnetized anomalies presented in the shell of the Earth. The presence of the anomalies interrupts the magnetic field of the Earth, monitored above the Earth. Using the difference of the magnetic fields before and after the presence of the magnetized anomalies, we show that one can uniquely recover the locations as well as their material parameters of the anomalies. Our study provides a rigorous mathematical theory to the geomagnetic detection technology that has been used in practice.

math.AP

On novel elastic structures inducing plasmonic resonances with finite frequencies and cloaking due to anomalous localized resonances

This paper is concerned with the theoretical study of plasmonic resonances for linear elasticity governed by the Lamé system in $\mathbb{R}^3$, and their application for cloaking due to anomalous localized resonances. We derive a very general and novel class of elastic structures that can induce plasmonic resonances. It is shown that if either one of the two convexity conditions on the Lamé parameters is broken, then we can construct certain plasmon structures that induce resonances. This significantly extends the relevant existing studies in the literature where the violation of both convexity conditions is required. Indeed, the existing plasmonic structures are a particular case of the general structures constructed in our study. Furthermore, we consider the plasmonic resonances within the finite frequency regime, and rigorously verify the quasi-static approximation for diametrically small plasmonic inclusions. Finally, as an application of the newly found structures, we construct a plasmonic device of the core-shell-matrix form that can induce cloaking due to anomalous localized resonance in the quasi-static regime, which also includes the existing study as a special case.

math.AP