Searcharxiv⌕ Search

arXiv subjects

Jianzhong Zhang

Publications and source records attributed to Jianzhong Zhang.

At least 19 recordsLinked to original sources

EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory

Long-horizon egocentric memory transforms continuous first-person video and audio into a searchable record of past experiences. We demonstrate two bottlenecks in existing systems: indices built from context-poor captions are unreliable for agentic search, while retrieval ignores a question's temporal intent. To address both bottlenecks, we introduce EgoCITE (Egocentric Context-augmented Indexing and Time-aware Evidence retrieval), a long-horizon agentic memory framework for egocentric QA. EgoCITE comprises three components. EgoScheme uses local multimodal context to turn fragmentary video captions and speech transcripts into self-contained atomic memory indices. EgoIndex organizes complementary action, activity, utterance, and conversation representations into searchable multi-view memory indices at multiple granularities. EgoRetrv combines semantic search with question-conditioned temporal relevance scoring and curation of retrieved evidence. We evaluate EgoCITE on EgoLifeQA, EgoMem, and EgoR1-Bench in terms of answer accuracy and target-event retrieval alignment. EgoCITE improves accuracy over agentic memory baselines by at least 4.4--14.2% while achieving 36$\times$ lower cost than long-context LLM agents.

cs.CV↗

StateScribe: Towards Accessible Change Awareness Across Real-World Revisits

Real-world environments evolve continuously, yet blind and low-vision (BLV) individuals often have limited access to understanding how they change over time. Unexpected or relocated objects, layout modifications, and content updates (e.g., price changes) can introduce safety risks and cognitive burden. While existing visual assistive technologies can describe immediate surroundings, they operate as one-off interactions and lack mechanisms to surface meaningful changes across revisits. Informed by a survey of 33 BLV individuals, we develop StateScribe, a system that supports accessible awareness of real-world changes across revisits. StateScribe employs a dual-layer memory architecture that integrates episodic scene memory and object-centric temporal memory to enable scalable and structured change tracking. It provides both live descriptions of the current scene, and descriptions of what has changed, when and where it occurred across revisits, such as "The shop on your right has a "CLOSED" sign; it was open at this time last week.'' Our evaluation shows that StateScribe maintains high accuracy (F1-score=83.1%) across 11 revisits, while remaining low-latency (mean<1.54s) and memory-efficient (<54MB) across 110 revisits. A user study with nine BLV participants demonstrates that StateScribe improves change awareness across revisits in three real-world locations. Finally, we discuss implications for long-term AI-assisted companions that support broader change observation using multimodal sensing, extend beyond changes to other memory capabilities, and adapt to individual users, intents, and contexts.

cs.HC↗

Sim2Field: End-to-End Development of AI RANs for 6G

Following state-of-the-art research results, which showed the potential for significant performance gains by applying AI/ML techniques in the cellular Radio Access Network (RAN), the wireless industry is now broadly pushing for the adoption of AI in 5G and future 6G technology. Despite this enthusiasm, AI-based wireless systems still remain largely untested in the field. Common simulation methods for generating datasets for AI model training suffer from "reality gap" and, as a result, the performance of these simulation-trained models may not carry over to practical cellular systems. Additionally, the cost and complexity of developing high-performance proof-of-concept implementations present major hurdles for evaluating AI wireless systems in the field. In this work, we introduce a methodology which aims to address the challenges of bringing AI to real networks. We discuss how detailed Digital Twin simulations may be employed for training site-specific AI Physical (PHY) layer functions. We further present a powerful testbed for AI-RAN research and demonstrate how it enables rapid prototyping, field testing and data collection. Finally, we evaluate an AI channel estimation algorithm over-the-air with a commercial UE, demonstrating that real-world throughput gains of up to 40% are achievable by incorporating AI in the physical layer.

cs.NI↗

Strong relaxation limit and uniform time asymptotics of the Jin-Xin model in the $L^{p}$ framework

We investigate the time-asymptotic stability of the Jin-Xin model and its diffusive relaxation limit toward viscous conservation laws in $\mathbb{R}^d$ for $d\geq 1$. First, we establish a priori estimates that are uniform with respect to both the time and the relaxation parameter $\varepsilon>0$, for initial data in hybrid Besov spaces based on $L^{p}$-norms. This uniformity enables us to derive $\mathcal{O}(\varepsilon)$ bounds on the difference between solutions of the viscous conservation law and its associated Jin-Xin approximation, thus justifying the strong convergence of the relaxation process. Furthermore, under an additional condition on the initial data, for instance, that the low frequencies belong to $L^{p/2}(\mathbb{R}^{d})$, we show that the $L^{p}(\mathbb{R}^d)$-norm of the solution to the Jin-Xin model decays at the optimal rate $(1+t)^{-d/{2p}}$, and the $L^{p}(\mathbb{R}^d)$-norm of its difference with the solution of the associated viscous conservation law decays at the enhanced rate $\varepsilon(1+t)^{-d/{2p}-1/2}$.

math.AP↗

Revised Optimal design of power electronic transformer based on hybrid MMC under over-modulation operation

The bridge arm of the hybrid modular multilevel converter (MMC) is composed of half-bridge and full-bridge sub-modules cascaded together. Compared with the half-bridge MMC, it can operate in the boost-AC mode, where the modulation index can be higher than 1, and the DC voltage and the AC voltage level are no longer mutually constrained; compared with the full-bridge MMC, it has lower switching device costs and losses. When the hybrid MMC boost-AC mode is used in the power electronic transformer, the degree of freedom in system design is improved, and the cost and volume of the power electronic transformer system can be further reduced. This paper analyzes how to make full use of the newly added modulation index of freedom introduced by the boost-AC hybrid MMC to optimize the power electronic transformer system, and finally gives the optimal modulation index selection scheme of the hybrid MMC for different optimization objectives.

eess.SY↗

A non-invasive fault location method for modular multilevel converters under light load conditions

This paper proposes a non-invasive fault location method for modular multilevel converters (MMC) considering light load conditions. The prior-art fault location methods of the MMC are often developed and verified under full load conditions. However, it is revealed that the faulty arm current will be suppressed to be unipolar when the open-circuit fault happens on the submodule switch under light load. This leads to the capacitor voltage of the healthy and faulty submodules rising or falling with the same variations, increasing the difficulty of fault location. The proposed approach of injecting the second-order circulating current will rebuild the bipolar arm current of the MMC and enlarge the capacitor voltage deviations between the healthy and faulty SMs. As a result, the fault location time is significantly shortened. The simulations are carried out to validate the effectiveness of the proposed approach, showing that the fault location time is reduced to 1/6 compared with the condition without second-order circulating current injection.

eess.SY↗

PolarDenseNet: A Deep Learning Model for CSI Feedback in MIMO Systems

In multiple-input multiple-output (MIMO) systems, the high-resolution channel information (CSI) is required at the base station (BS) to ensure optimal performance, especially in the case of multi-user MIMO (MU-MIMO) systems. In the absence of channel reciprocity in frequency division duplex (FDD) systems, the user needs to send the CSI to the BS. Often the large overhead associated with this CSI feedback in FDD systems becomes the bottleneck in improving the system performance. In this paper, we propose an AI-based CSI feedback based on an auto-encoder architecture that encodes the CSI at UE into a low-dimensional latent space and decodes it back at the BS by effectively reducing the feedback overhead while minimizing the loss during recovery. Our simulation results show that the AI-based proposed architecture outperforms the state-of-the-art high-resolution linear combination codebook using the DFT basis adopted in the 5G New Radio (NR) system.

cs.IT↗

Experimental Investigation of Frequency Domain Channel Extrapolation in Massive MIMO Systems for Zero-Feedback FDD

Estimating downlink (DL) channel state information (CSI) in frequency division duplex (FDD) massive multi-input multi-output (MIMO) systems generally requires downlink pilots and feedback overheads. Accordingly, this paper investigates the feasibility of zero-feedback FDD massive MIMO systems based on channel extrapolation. We use the high-resolution parameter estimation (HRPE), specifically the space-alternating generalized expectation-maximization (SAGE) algorithm, to extrapolate the DL CSI based on the extracted parameters of multipath components in the uplink channel. We apply the HRPE to two different channel models: the vector spatial signature (VSS) model and the direction of arrival (DOA) model. We verify these methods through real-world channel data acquired from channel measurement campaigns with two different types of channel sounders: a) a switched array-based, real-time, time-domain, outdoors setup at 3.5 GHz, and b) a virtual array-based, high-accuracy, frequency-domain, indoors setup at 2.4 and 5-7 GHz. The performance metrics of the extrapolated channels that we evaluate include the mean squared error, beamforming efficiency, and spectral efficiency in multiuser MIMO scenarios. The results show that the HRPE-based channel extrapolation performs best under the simple VSS model, which does not require array calibration, and if the BS is in an open outdoor environment having line-of-sight (LOS) paths to well-separated users.

eess.SP↗

Robust Non-Coherent Beamforming for FDD Downlink Massive MIMO

Designing beamforming techniques for the downlink (DL) of frequency division duplex (FDD) massive MIMO is known to be a challenging problem due to the difficulty of obtaining channel state information (CSI). Indeed, since the uplink-downlink bands are disjoint, the system cannot rely on channel reciprocity to estimate the channel from uplink (UL) pilots as in time division duplexing (TDD) system. Still, in this paper, we propose original designs for robust beamformers that do not require any feedback from the users and only rely on the transmission of UL pilots. The price to pay is that the beamformer is non-coherent in the sense that it does not leverage full knowledge of the phase of each multipath component. A large variety of novel designs are proposed under different criterion and partial phase knowledge.

eess.SP↗

Performance Analysis of Channel Extrapolation in FDD Massive MIMO Systems

Channel estimation for the downlink of frequency division duplex (FDD) massive MIMO systems is well known to generate a large overhead as the amount of training generally scales with the number of transmit antennas in a MIMO system. In this paper, we consider the solution of extrapolating the channel frequency response from uplink pilot estimates to the downlink frequency band, which completely removes the training overhead. We first show that conventional estimators fail to achieve reasonable accuracy. We propose instead to use high-resolution channel estimation. We derive theoretical lower bounds (LB) for the mean squared error (MSE) of the extrapolated channel. Assuming that the paths are well separated, the LB is simplified in an expression that gives considerable physical insight. It is then shown that the MSE is inversely proportional to the number of receive antennas while the extrapolation performance penalty scales with the square of the ratio of the frequency offset and the training bandwidth. The channel extrapolation performance is validated through numeric simulations and experimental measurements taken in an anechoic chamber. Our main conclusion is that channel extrapolation is a viable solution for FDD massive MIMO systems if accurate system calibration is performed and favorable propagation conditions are present.

eess.SP↗

Channel Extrapolation for FDD Massive MIMO: Procedure and Experimental Results

Application of massive multiple-input multiple-output (MIMO) systems to frequency division duplex (FDD) is challenging mainly due to the considerable overhead required for downlink training and feedback. Channel extrapolation, i.e., estimating the channel response at the downlink frequency band based on measurements in the disjoint uplink band, is a promising solution to overcome this bottleneck. This paper presents measurement campaigns obtained by using a wideband (350 MHz) channel sounder at 3.5 GHz composed of a calibrated 64 element antenna array, in both an anechoic chamber and outdoor environment. The Space Alternating Generalized Expectation-Maximization (SAGE) algorithm was used to extract the parameters (amplitude, delay, and angular information) of the multipath components from the attained channel data within the training (uplink) band. The channel in the downlink band is then reconstructed based on these path parameters. The performance of the extrapolated channel is evaluated in terms of mean squared error (MSE) and reduction of beamforming gain (RBG) in comparison to the ground truth, i.e., the measured channel at the downlink frequency. We find strong sensitivity to calibration errors and model mismatch, and also find that performance depends on propagation conditions: LOS performs significantly better than NLOS.

eess.SP↗

Channel Correlation Diversity in MU-MIMO Systems -- Analysis and Measurements

In multiuser multiple-input multiple-output (MU-MIMO) systems, channel correlation is detrimental to system performance. We demonstrate that widely used, yet overly simplified, correlation models that generate identical correlation profiles for each terminal tend to severely underestimate the system performance. In sharp contrast, more physically motivated models that capture variations in the power angular spectra across multiple terminals, generate diverse correlation patterns. This has a significant impact on the system performance. Assuming correlated Rayleigh fading and downlink zero-forcing precoding, tight closed-form approximations for the average signal-to-noise-ratio, and ergodic sum spectral efficiency are derived. Our expressions provide clear insights into the impact of diverse correlation patterns on the above performance metrics. Unlike previous works, the correlation models are parameterized with measured data from a recent 2.53 GHz urban macrocellular campaign in Cologne, Germany. Overall, results from this paper can be treated as a timely re-calibration of performance expectations from practical MU-MIMO systems.

cs.IT↗

How Many Antennas Do We Need for Massive MIMO Channel Sounding? - Validating Through Measurement

This paper investigates the impact of the number of antennas (8 to 64) and the array configuration on massive MIMO channel parameters estimation for multiple propagation scenarios at 3.5 GHz. Different measurement environments are artificially created by placing several reflectors and absorbers in an anechoic chamber. Ground truth channel parameters, e.g, path angles, are obtained by geometry and trigonometric rules. Then, these are compared to the channel parameters extracted by the applying Space-Alternating Generalized Expectation-Maximization (SAGE) algorithm on the measurements. Overall, the estimation errors for various array configurations and the multiple environments are compared. This paper will help to determine the appropriate configuration of the antenna array and the parameter extraction algorithm for outdoor massive MIMO channel sounding campaigns.

eess.SP↗

Channel Extrapolation in FDD Massive MIMO: Theoretical Analysis and Numerical Validation

Downlink channel estimation in massive MIMO systems is well known to generate a large overhead in frequency division duplex (FDD) mode as the amount of training generally scales with the number of transmit antennas. Using instead an extrapolation of the channel from the measured uplink estimates to the downlink frequency band completely removes this overhead. In this paper, we investigate the theoretical limits of channel extrapolation in frequency. We highlight the advantage of basing the extrapolation on high-resolution channel estimation. A lower bound (LB) on the mean squared error (MSE) of the extrapolated channel is derived. A simplified LB is also proposed, giving physical intuition on the SNR gain and extrapolation range that can be expected in practice. The validity of the simplified LB relies on the assumption that the paths are well separated. The SNR gain then linearly improves with the number of receive antennas while the extrapolation performance penalty quadratically scales with the ratio of the frequency and the training bandwidth. The theoretical LB is numerically evaluated using a 3GPP channel model and we show that the LB can be reached by practical high-resolution parameter extraction algorithms. Our results show that there are strong limitations on the extrapolation range than can be expected in SISO systems while much more promising results can be obtained in the multiple-antenna setting as the paths can be more easily separated in the delay-angle domain.

cs.IT↗

Distributive Dynamic Spectrum Access through Deep Reinforcement Learning: A Reservoir Computing Based Approach

Dynamic spectrum access (DSA) is regarded as an effective and efficient technology to share radio spectrum among different networks. As a secondary user (SU), a DSA device will face two critical problems: avoiding causing harmful interference to primary users (PUs), and conducting effective interference coordination with other secondary users. These two problems become even more challenging for a distributed DSA network where there is no centralized controllers for SUs. In this paper, we investigate communication strategies of a distributive DSA network under the presence of spectrum sensing errors. To be specific, we apply the powerful machine learning tool, deep reinforcement learning (DRL), for SUs to learn "appropriate" spectrum access strategies in a distributed fashion assuming NO knowledge of the underlying system statistics. Furthermore, a special type of recurrent neural network (RNN), called the reservoir computing (RC), is utilized to realize DRL by taking advantage of the underlying temporal correlation of the DSA network. Using the introduced machine learning-based strategy, SUs could make spectrum access decisions distributedly relying only on their own current and past spectrum sensing outcomes. Through extensive experiments, our results suggest that the RC-based spectrum access strategy can help the SU to significantly reduce the chances of collision with PUs and other SUs. We also show that our scheme outperforms the myopic method which assumes the knowledge of system statistics, and converges faster than the Q-learning method when the number of channels is large.

cs.LG↗

Real-Time Millimeter-Wave MIMO Channel Sounder for Dynamic Directional Measurements

In this paper, we present a novel real-time multiple-input-multiple-output (MIMO) channel sounder for the 28 GHz band. Until now, most investigations of the directional characteristics of millimeter-wave channels have used mechanically rotating horn antennas. In contrast, the sounder presented here is capable of performing horizontal and vertical beam steering with the help of phased arrays. Due to its fast beam-switching capability, the proposed sounder can perform measurements that are directionally resolved both at the transmitter(TX) and receiver (RX) in 1.44 milliseconds compared to the minutes or even hours required for rotating horn antenna sounders. This not only enables measurement of more TX-RX locations for a better statistical validity but also allows to perform directional analysis in dynamic environments. The short measurement time combined with the high phase stability limits the phase drift between TX and RX, enabling phase-coherent sounding of all beam pairs even when TX and RX have no cabled connection for synchronization without any delay ambiguity. Furthermore, the phase stability over time enables complex RX waveform averaging to improve the signal to noise ratio during high path loss measurements. The paper discusses both the system design as well as the measurements performed for verification of the sounder performance. Furthermore, we present sample results from double directional measurements in dynamic environments.

eess.SP↗

Outdoor to Indoor Penetration Loss at 28 GHz for Fixed Wireless Access

This paper present the results from a 28 GHz channel sounding campaign performed to investigate the effects of outdoor to indoor penetration on the wireless propagation channel characteristics for an urban microcell in a fixed wireless access scenario. The measurements are performed with a real-time channel sounder, which can measure path loss up to 169 dB, and equipped with phased array antennas that allows electrical beam steering for directionally resolved measurements in dynamic environments. Thanks to the short measurement time and the excellent phase stability of the system, we obtain both directional and omnidirectional channel power delay profiles without any delay uncertainty. For outdoor and indoor receiver locations, we compare path loss, delay spreads and angular spreads obtained for two different types of buildings.

cs.IT↗