SearcharxivSearch

arXiv subjects

Zhengyu Wang

Publications and source records attributed to Zhengyu Wang.

At least 19 recordsLinked to original sources

A Large-Dimensional Analysis of ESPRIT DoA Estimation: Inconsistency and a Correction via RMT

In this paper, we perform asymptotic analyses of the widely used ESPRIT direction-of-arrival (DoA) estimator for large arrays, where the array size $N$ and the number of snapshots $T$ grow to infinity at the same pace. In this large-dimensional regime, the sample covariance matrix (SCM) is known to be a poor eigenspectral estimator of the population covariance. We show that the classical ESPRIT algorithm, that relies on the SCM, and as a consequence of the large-dimensional inconsistency of the SCM, produces inconsistent DoA estimates as $N,T \to \infty$ with $N/T \to c \in (0,\infty)$, for both widely-~and~closely-spaced DoAs. Leveraging tools from random matrix theory (RMT), we propose an improved G-ESPRIT method and prove its consistency in the same large-dimensional setting. From a technical perspective, we derive a novel bound on the eigenvalue differences between two potentially non-Hermitian matrices, which may be of independent interest. Numerical simulations are provided to corroborate our theoretical findings.

eess.SP

Not Search, But Scan: Benchmarking MLLMs on Scan-Oriented Academic Paper Reasoning

With the rapid progress of multimodal large language models (MLLMs), AI already performs well at literature retrieval and certain reasoning tasks, serving as a capable assistant to human researchers, yet it remains far from autonomous research. The fundamental reason is that current work on academic paper reasoning is largely confined to a search-oriented paradigm centered on pre-specified targets, with reasoning grounded in relevance retrieval, which struggles to support researcher-style full-document understanding, reasoning, and verification. To bridge this gap, we propose \textbf{ScholScan}, a new benchmark for academic paper reasoning. ScholScan introduces a scan-oriented task setting that asks models to read and cross-check entire papers like human researchers, scanning the document to identify consistency issues. The benchmark comprises 1,800 carefully annotated questions drawn from nine error categories across 13 natural-science domains and 715 papers, and provides detailed annotations for evidence localization and reasoning traces, together with a unified evaluation protocol. We assessed 15 models across 24 input configurations and conducted a fine-grained analysis of MLLM capabilities for all error categories. Across the board, retrieval-augmented generation (RAG) methods yield no significant improvements, revealing systematic deficiencies of current MLLMs on scan-oriented tasks and underscoring the challenge posed by ScholScan. We expect ScholScan to be the leading and representative work of the scan-oriented task paradigm.

cs.AI

Metasurface-Enabled Extremely Large-Scale Antenna Systems: Transceiver Architecture, Physical Modeling, and Channel Estimation

Extremely large-scale antenna arrays (ELAAs) have emerged as a pivotal technology for addressing the unprecedented performance demands of next-generation wireless communication systems. To enhance their practicality, we propose metasurface-enabled extremely large-scale antenna (MELA) systems -- novel transceiver architectures that employ reconfigurable transmissive metasurfaces to facilitate efficient over-the-air RF-to-antenna coupling and phase control. This architecture eliminates the need for bulky switch matrices and costly phase-shifter networks typically required in conventional solutions. Physically grounded models are developed to characterize electromagnetic field propagation through individual transmissive unit cells, capturing the fundamental physics of wave transformation and transmission. Additionally, distance-dependent approximate models are introduced, exhibiting structural properties conducive to efficient parameter estimation and signal processing. Based on the channel model, a two stage channel estimation framework is proposed for the scenarios comprising users in the hybrid near- and far-fields. In the first stage, a dictionary-driven beamspace filtering technique enables rapid angular-domain scanning. In the refinement stage, the rotational symmetry of subarrays is exploited to design super-resolution estimators that jointly recover angular and range parameters. An analytical expression for the half-power beamwidth of MELA is derived, revealing its near-optimal spatial resolution relative to conventional ELAA architectures. Numerical experiments further validate the high-resolution of the proposed channel estimation algorithm and the fidelity of the electromagnetic model, positioning the MELA architecture as a highly competitive and forward-looking solution for practical ELAA deployment.

eess.SP

FinMMDocR: Benchmarking Financial Multimodal Reasoning with Scenario Awareness, Document Understanding, and Multi-Step Computation

We introduce FinMMDocR, a novel bilingual multimodal benchmark for evaluating multimodal large language models (MLLMs) on real-world financial numerical reasoning. Compared to existing benchmarks, our work delivers three major advancements. (1) Scenario Awareness: 57.9% of 1,200 expert-annotated problems incorporate 12 types of implicit financial scenarios (e.g., Portfolio Management), challenging models to perform expert-level reasoning based on assumptions; (2) Document Understanding: 837 Chinese/English documents spanning 9 types (e.g., Company Research) average 50.8 pages with rich visual elements, significantly surpassing existing benchmarks in both breadth and depth of financial documents; (3) Multi-Step Computation: Problems demand 11-step reasoning on average (5.3 extraction + 5.7 calculation steps), with 65.0% requiring cross-page evidence (2.4 pages average). The best-performing MLLM achieves only 58.0% accuracy, and different retrieval-augmented generation (RAG) methods show significant performance variations on this task. We expect FinMMDocR to drive improvements in MLLMs and reasoning-enhanced methods on complex multimodal reasoning tasks in real-world scenarios.

cs.CV

Source Localization and Power Estimation through RISs: Performance Analysis and Prototype Validations

This paper investigates the capabilities and effectiveness of backward localization centered on reconfigurable intelligent surfaces (RISs). In the backward sensing paradigm, the region of interest (RoI) is illuminated using a set of diverse radiation patterns. These patterns encode spatial information into a sequence of measurements, which are subsequently processed to reconstruct the RoI. We show that a single RIS can estimate the direction of arrival of incident waves by leveraging configurational diversity, and that the spatial diversity provided by multiple RISs further improves the accuracy of source localization and power estimation. The underlying structure of the sensing operator in the multi-snapshot measurement process is clarified. For single-RIS localization, the sensing operator is decomposed into a product of structured matrices, each corresponding to a specific physical process: wave propagation to and from the RIS, the relative phase offsets of elements with respect to the reference point, and the applied phase configuration of each element. A unified framework for identifying key performance indicators is established by analyzing the conditioning of the sensing operators. In the multi-RIS setting, we derive--via rank analysis--the governing law among the RoI size, the number of elements, and the number of measurements. Upper bounds on the relative error of the least squares reconstruction algorithm are derived. These bounds clarify how key performance indicators affect estimation error and provide valuable guidance for system-level optimization. Numerical experiments confirm that the trend of the relative error is consistent with the theoretical bounds.

eess.SP

Toward Wireless Localization Using Multiple Reconfigurable Intelligent Surfaces

This paper investigates the capabilities and effectiveness of backward sensing centered on reconfigurable intelligent surfaces (RISs). We demonstrate that the direction of arrival (DoA) estimation of incident waves in the far-field regime can be accomplished using a single RIS by leveraging configurational diversity. Furthermore, we identify that the spatial diversity achieved through deploying multiple RISs enables accurate localization of multiple power sources. Physically accurate and mathematically concise models are introduced to characterize forward signal aggregations via RISs. By employing linearized approximations inherent in the far-field region, the measurement process for various configurations can be expressed as a system of linear equations. The mathematical essence of backward sensing lies in solving this system. A theoretical framework for determining key performance indicators is established through condition number analysis of the sensing operators. In the context of localization using multiple RISs, we examine relationships among the rank of sensing operators, the size of the region of interest (RoI), and the number of elements and measurements. For DoA estimations, we provide an upper bound for the relative error of the least squares reconstruction algorithm. These quantitative analyses offer essential insights for system design and optimization. Numerical experiments validate our findings. To demonstrate the practicality of our proposed RIS-centric sensing approach, we develop a proof-of-concept prototype using universal software radio peripherals (USRP) and employ a magnitude-only reconstruction algorithm tailored for this system. To our knowledge, this represents the first trial of its kind.

eess.SP

Asymptotic CRB Analysis of Random RIS-Assisted Large-Scale Localization Systems

This paper studies the performance of a randomly RIS-assisted multi-target localization system, in which the configurations of the RIS are randomly set to avoid high-complexity optimization. We first focus on the scenario where the number of RIS elements is significantly large, and then obtain the scaling law of Cramér-Rao bound (CRB) under certain conditions, which shows that CRB decreases in the third or fourth order as the RIS dimension increases. Second, we extend our analysis to large systems where both the number of targets and sensors is substantial. Under this setting, we explore two common RIS models: the constant module model and the discrete amplitude model, and illustrate how the random RIS configuration impacts the value of CRB. Numerical results demonstrate that asymptotic formulas provide a good approximation to the exact CRB in the proposed randomly configured RIS systems.

cs.IT

Wireless Regional Imaging through Reconfigurable Intelligent Surfaces: Passive Mode

In this paper, we propose a multi-RIS-aided wireless imaging framework in 3D facing the distributed placement of multi-sensor networks. The system creates a randomized reflection pattern by adjusting the RIS phase shift, enabling the receiver to capture signals within the designated space of interest (SoI). Firstly, a multi-RIS-aided linear imaging channel modeling is proposed. We introduce a theoretical framework of computational imaging to recover the signal strength distribution of the SOI. For the RIS-aided imaging system, the impact of multiple parameters on the performance of the imaging system is analyzed. The simulation results verify the correctness of the proposal. Furthermore, we propose an amplitude-only imaging algorithm for the RIS-aided imaging system to mitigate the problem of phase unpredictability. Finally, the performance verification of the imaging algorithm is carried out by proof of concept experiments under reasonable parameter settings.

eess.IV

Design of Reconfigurable Intelligent Surfaces for Wireless Communication: A Review

This paper addresses the hardware structure of Reconfigurable Intelligent Surfaces (RIS) and presents a comprehensive overview of RIS design, considering both unit design and prototype systems. It commences by tracing the evolutionary trajectory of RIS, originating from static cell-structured hypersurfaces. The article conducts a meticulous examination from the standpoint of adaptability, elucidating the diverse array of unit structures and design philosophies that underlie existing RIS frameworks. Following this, the study systematically categorizes and synthesizes channel modeling research for RIS-facilitated wireless communication, leveraging both physical insights and statistical data. Additionally, the article provides a detailed exposition of current RIS experimental setups and their corresponding empirical findings, delving into the attributes of prototype design and system functionalities. Moreover, this work introduces an in-house developed RIS prototype. The prototype undergoes rigorous empirical evaluation, encompassing multi-hop RIS signal amplification, image reconstruction, and real-world indoor signal coverage experiments. The empirical results robustly affirm the efficacy of RIS in effectively mitigating signal coverage blind spots and enabling radio wave imaging. With RIS-enhanced augmentation, the average indoor signal gain surpasses 8 dB.

eess.SY

Multi-RIS-aided Wireless Communications in Real-world: Prototyping and Field Trials

The performance of multiple reconfigurable intelligent surfaces (RISs) receives limited attention in previous studies. This article fills this research gap by investigating the capabilities of multiple RISs in real-world networks. We propose a simplified yet highly scalable sandwich architecture for implementing one-bit unit cells, with the flexibility to accommodate multi-bit unit cells. To effectively control multiple RISs, we present a cost-effective remote-controlling scheme and develop a cloud-based RIS management system. Through a series of four field trials, we demonstrate the effectiveness of multi-hop routing schemes in establishing reliable links. Our experiments reveal significant improvements in signal strength and data transmission in multi-RIS-aided Wi-Fi and commercial 5G networks. Furthermore, we investigate the power scaling law of RIS-aided beamforming and provide insights into the roles of the later nodes in multi-hop relay chains.

eess.SY

Joint Beamforming Design and 3D DoA Estimation for RIS-aided Communication System

In this paper, we consider a reconfigurable intelligent surface (RIS)-assisted 3D direction-of-arrival (DoA) estimation system, in which a uniform planar array (UPA) RIS is deployed to provide virtual line-of-sight (LOS) links and reflect the uplink pilot signal to sensors. To overcome the mutually coupled problem between the beamforming design at the RIS and DoA estimation, we explore the separable sparse representation structure and propose an alternating optimization algorithm. The grid-based DoA estimation is modeled as a joint-sparse recovery problem considering the grid bias, and the Joint-2D-OMP method is used to estimate both on-grid and off-grid parts. The corresponding Cramér-Rao lower bound (CRLB) is derived to evaluate the estimation. Then, the beampattern at the RIS is optimized to maximize the signal-to-noise (SNR) at sensors according to the estimated angles. Numerical results show that the proposed alternating optimization algorithm can achieve lower estimation error compared to benchmarks of random beamforming design.

eess.SP

Point Cloud Quality Assessment using 3D Saliency Maps

Point cloud quality assessment (PCQA) has become an appealing research field in recent days. Considering the importance of saliency detection in quality assessment, we propose an effective full-reference PCQA metric which makes the first attempt to utilize the saliency information to facilitate quality prediction, called point cloud quality assessment using 3D saliency maps (PQSM). Specifically, we first propose a projection-based point cloud saliency map generation method, in which depth information is introduced to better reflect the geometric characteristics of point clouds. Then, we construct point cloud local neighborhoods to derive three structural descriptors to indicate the geometry, color and saliency discrepancies. Finally, a saliency-based pooling strategy is proposed to generate the final quality score. Extensive experiments are performed on four independent PCQA databases. The results demonstrate that the proposed PQSM shows competitive performances compared to multiple state-of-the-art PCQA metrics.

cs.CV

Towards Analytical Electromagnetic Models for Reconfigurable Intelligent Surfaces

Physically accurate and mathematically tractable models are presented to characterize scattering and reflection properties of reconfigurable intelligent surfaces (RISs). We take continuous and discrete strategies to model a single patch and patch array and their interactions with multiple incident electromagnetic (EM) waves. The proposed models consider the effect of the incident and scattered angles, polarization features, and the topology and geometry of RISs. Particularly, a simple system of linear equations can describe the multiple-input multiple-output (MIMO) behaviors of RISs under reasonable assumptions. It can serve as a fundamental model for analyzing and optimizing the performance of RIS-aided systems in the far-field regime. The proposed models are employed to identify the advantages and limitations of three typical configurations. One important finding is that complicated beam reshaping functionality can not be endowed by popular phase compensation configurations. A possible solution is the simultaneous configurations of collecting area and phase shifting. Numerical simulations verify the effectiveness of the proposed configuration schemes.

eess.SY

(Nearly) Sample-Optimal Sparse Fourier Transform in Any Dimension; RIPless and Filterless

In this paper, we consider the extensively studied problem of computing a $k$-sparse approximation to the $d$-dimensional Fourier transform of a length $n$ signal. Our algorithm uses $O(k \log k \log n)$ samples, is dimension-free, operates for any universe size, and achieves the strongest $\ell_\infty/\ell_2$ guarantee, while running in a time comparable to the Fast Fourier Transform. In contrast to previous algorithms which proceed either via the Restricted Isometry Property or via filter functions, our approach offers a fresh perspective to the sparse Fourier Transform problem.

cs.DS

Parallel Graph Connectivity in Log Diameter Rounds

We study graph connectivity problem in MPC model. On an undirected graph with $n$ nodes and $m$ edges, $O(\log n)$ round connectivity algorithms have been known for over 35 years. However, no algorithms with better complexity bounds were known. In this work, we give fully scalable, faster algorithms for the connectivity problem, by parameterizing the time complexity as a function of the diameter of the graph. Our main result is a $O(\log D \log\log_{m/n} n)$ time connectivity algorithm for diameter-$D$ graphs, using $Θ(m)$ total memory. If our algorithm can use more memory, it can terminate in fewer rounds, and there is no lower bound on the memory per processor. We extend our results to related graph problems such as spanning forest, finding a DFS sequence, exact/approximate minimum spanning forest, and bottleneck spanning forest. We also show that achieving similar bounds for reachability in directed graphs would imply faster boolean matrix multiplication algorithms. We introduce several new algorithmic ideas. We describe a general technique called double exponential speed problem size reduction which roughly means that if we can use total memory $N$ to reduce a problem from size $n$ to $n/k$, for $k=(N/n)^{Θ(1)}$ in one phase, then we can solve the problem in $O(\log\log_{N/n} n)$ phases. In order to achieve this fast reduction for graph connectivity, we use a multistep algorithm. One key step is a carefully constructed truncated broadcasting scheme where each node broadcasts neighbor sets to its neighbors in a way that limits the size of the resulting neighbor sets. Another key step is random leader contraction, where we choose a smaller set of leaders than many previous works do.

cs.DS

BPTree: an $\ell_2$ heavy hitters algorithm using constant memory

The task of finding heavy hitters is one of the best known and well studied problems in the area of data streams. One is given a list $i_1,i_2,\ldots,i_m\in[n]$ and the goal is to identify the items among $[n]$ that appear frequently in the list. In sub-polynomial space, the strongest guarantee available is the $\ell_2$ guarantee, which requires finding all items that occur at least $ε\|f\|_2$ times in the stream, where the vector $f\in\mathbb{R}^n$ is the count histogram of the stream with $i$th coordinate equal to the number of times~$i$ appears $f_i:=\#\{j\in[m]:i_j=i\}$. The first algorithm to achieve the $\ell_2$ guarantee was the CountSketch of [CCF04], which requires $O(ε^{-2}\log n)$ words of memory and $O(\log n)$ update time and is known to be space-optimal if the stream allows for deletions. The recent work of [BCIW16] gave an improved algorithm for insertion-only streams, using only $O(ε^{-2}\logε^{-1}\log\log n)$ words of memory. In this work, we give an algorithm \bptree for $\ell_2$ heavy hitters in insertion-only streams that achieves $O(ε^{-2}\logε^{-1})$ words of memory and $O(\logε^{-1})$ update time, which is the optimal dependence on $n$ and $m$. In addition, we describe an algorithm for tracking $\|f\|_2$ at all times with $O(ε^{-2})$ memory and update time. Our analyses rely on bounding the expected supremum of a Bernoulli process involving Rademachers with limited independence, which we accomplish via a Dudley-like chaining argument that may have applications elsewhere.

cs.DS

Optimal lower bounds for universal relation, and for samplers and finding duplicates in streams

In the communication problem $\mathbf{UR}$ (universal relation) [KRW95], Alice and Bob respectively receive $x, y \in\{0,1\}^n$ with the promise that $x\neq y$. The last player to receive a message must output an index $i$ such that $x_i\neq y_i$. We prove that the randomized one-way communication complexity of this problem in the public coin model is exactly $Θ(\min\{n,\log(1/δ)\log^2(\frac n{\log(1/δ)})\})$ for failure probability $δ$. Our lower bound holds even if promised $\mathop{support}(y)\subset \mathop{support}(x)$. As a corollary, we obtain optimal lower bounds for $\ell_p$-sampling in strict turnstile streams for $0\le p < 2$, as well as for the problem of finding duplicates in a stream. Our lower bounds do not need to use large weights, and hold even if promised $x\in\{0,1\}^n$ at all points in the stream. We give two different proofs of our main result. The first proof demonstrates that any algorithm $\mathcal A$ solving sampling problems in turnstile streams in low memory can be used to encode subsets of $[n]$ of certain sizes into a number of bits below the information theoretic minimum. Our encoder makes adaptive queries to $\mathcal A$ throughout its execution, but done carefully so as to not violate correctness. This is accomplished by injecting random noise into the encoder's interactions with $\mathcal A$, which is loosely motivated by techniques in differential privacy. Our second proof is via a novel randomized reduction from Augmented Indexing [MNSW98] which needs to interact with $\mathcal A$ adaptively. To handle the adaptivity we identify certain likely interaction patterns and union bound over them to guarantee correct interaction on all of them. To guarantee correctness, it is important that the interaction hides some of its randomness from $\mathcal A$ in the reduction.

cs.CC

Optimal lower bounds for universal relation, samplers, and finding duplicates

In the communication problem $\mathbf{UR}$ (universal relation) [KRW95], Alice and Bob respectively receive $x$ and $y$ in $\{0,1\}^n$ with the promise that $x\neq y$. The last player to receive a message must output an index $i$ such that $x_i\neq y_i$. We prove that the randomized one-way communication complexity of this problem in the public coin model is exactly $Θ(\min\{n, \log(1/δ)\log^2(\frac{n}{\log(1/δ)})\})$ bits for failure probability $δ$. Our lower bound holds even if promised $\mathop{support}(y)\subset \mathop{support}(x)$. As a corollary, we obtain optimal lower bounds for $\ell_p$-sampling in strict turnstile streams for $0\le p < 2$, as well as for the problem of finding duplicates in a stream. Our lower bounds do not need to use large weights, and hold even if it is promised that $x\in\{0,1\}^n$ at all points in the stream. Our lower bound demonstrates that any algorithm $\mathcal{A}$ solving sampling problems in turnstile streams in low memory can be used to encode subsets of $[n]$ of certain sizes into a number of bits below the information theoretic minimum. Our encoder makes adaptive queries to $\mathcal{A}$ throughout its execution, but done carefully so as to not violate correctness. This is accomplished by injecting random noise into the encoder's interactions with $\mathcal{A}$, which is loosely motivated by techniques in differential privacy. Our correctness analysis involves understanding the ability of $\mathcal{A}$ to correctly answer adaptive queries which have positive but bounded mutual information with $\mathcal{A}$'s internal randomness, and may be of independent interest in the newly emerging area of adaptive data analysis with a theoretical computer science lens.

cs.CC