Searcharxiv⌕ Search

arXiv subjects

Lingzhi Zhao

Publications and source records attributed to Lingzhi Zhao.

11 recordsLinked to original sources

Navigating Sparse Singlet Fission Chemical Space: An Intelligent Generative-Predictive Paradigm

Singlet fission (SF) offers a promising route to surpass the Shockley-Queisser limit by converting a photoexcited singlet exciton into two triplet excitons, thereby enhancing photovoltaic energy conversion efficiency. However, realizing efficient SF process requires stringent energetic requirements among low-lying excited states that render SF molecules intrinsically rare within the vast chemical space. This extreme sparsity presents a grand challenge for molecular discovery. Due to low hit rates and trial-and-error computational waste on nonviable structures, conventional high-throughput virtual screening faces significant constraints, even when accelerated by machine learning models. Here, we establish a synergistic generative-predictive framework for the targeted inverse design of SF molecules by integrating a structure generator, a properties predictor and a multi-criteria validation workflow. By continuously coupling generative exploration with SF predictive models, the framework progressively enriches SF species and achieves a success rate of approximately 90% in generating molecules that satisfy the target SF energetic criteria. High-throughput evaluation of about 100 million generated structures with time-dependent density functional theory (TDDFT) validation of just a random 1% subset confirmed a 90.8% success rate for SF candidates. All together, we constructed an SF database of 283,559 candidates with favorable energetics of excited states and synthetic accessibility. From it, we identified a key fragment strongly associated with the requirements for SF, namely, CN([O])N(C)[O]. These findings establish an efficient route for overcoming the sparsity difficulty in SF molecular discovery and provide interpretable design principles for the development of novel excited-state functional materials.

cond-mat.mtrl-sci↗

QoS-QoE Translation with Large Language Model

QoS-QoE translation is a fundamental problem in multimedia systems because it characterizes how measurable system and network conditions affect user-perceived experience. Although many prior studies have examined this relationship, their findings are often developed for specific setups and remain scattered across papers, experimental settings, and reporting formats, limiting systematic reuse, cross-scenario generalization, and large-scale analysis. To address this gap, we first introduce QoS-QoE Translation dataset, a source-grounded dataset of structured QoS-QoE relationships from the multimedia literature, with a focus on video streaming related tasks. We construct the dataset through an automated pipeline that combines paper curation, QoS-QoE relationship extraction, and iterative data evaluation. Each record preserves the extracted relationship together with parameter definitions, supporting evidence, and contextual metadata. We further evaluate the capability of large language models (LLMs) on QoS-QoE translation, both before and after supervised fine-tuning on our dataset, and show strong performance on both continuous-value and discrete-label prediction in bidirectional translation, from QoS-QoE and QoE-QoS. Our dataset provides a foundation for benchmarking LLMs in QoS-QoE translation and for supporting future LLM-based reasoning for multimedia quality prediction and optimization. The complete dataset and code are publicly available at https://yyu6969.github.io/qos-qoe-translation-page/, for full reproducibility and open access.

cs.MM↗

AquaVLM: Improving Underwater Situation Awareness with Mobile Vision Language Models

Underwater activities like scuba diving enable millions annually to explore marine environments for recreation and scientific research. Maintaining situational awareness and effective communication are essential for diver safety. Traditional underwater communication systems are often bulky and expensive, limiting their accessibility to divers of all levels. While recent systems leverage lightweight smartphones and support text messaging, the messages are predefined and thus restrict context-specific communication. In this paper, we present AquaVLM, a tap-and-send underwater communication system that automatically generates context-aware messages and transmits them using ubiquitous smartphones. Our system features a mobile vision-language model (VLM) fine-tuned on an auto-generated underwater conversation dataset and employs a hierarchical message generation pipeline. We co-design the VLM and transmission, incorporating error-resilient fine-tuning to improve the system's robustness to transmission errors. We develop a VR simulator to enable users to experience AquaVLM in a realistic underwater environment and create a fully functional prototype on the iOS platform for real-world experiments. Both subjective and objective evaluations validate the effectiveness of AquaVLM and highlight its potential for personal underwater communication as well as broader mobile VLM applications.

cs.HC↗

STAC: Leveraging Spatio-Temporal Data Associations For Efficient Cross-Camera Streaming and Analytics

In IoT based distributed network of cameras, real-time multi-camera video analytics is challenged by high bandwidth demands and redundant visual data, creating a fundamental tension where reducing data saves network overhead but can degrade model performance, and vice versa. We present STAC, a cross-cameras surveillance system that leverages spatio-temporal associations for efficient object tracking under constrained network conditions. STAC integrates multi-resolution feature learning, ensuring robustness under variable networked system level optimizations such as frame filtering, FFmpeg-based compression, and Region-of-Interest (RoI) masking, to eliminate redundant content across distributed video streams while preserving downstream model accuracy for object identification and tracking. Evaluated on NVIDIA's AICity Challenge dataset, STAC achieves a 76\% improvement in tracking accuracy and an 8.6x reduction in inference latency over a standard multi-object multi-camera tracking baseline (using YOLOv4 and DeepSORT). Furthermore, 29\% of redundant frames are filtered, significantly reducing data volume without compromising inference quality.

cs.CV↗

AquaScope: Reliable Underwater Image Transmission on Mobile Devices

Underwater communication is essential for both recreational and scientific activities, such as scuba diving. However, existing methods remain highly constrained by environmental challenges and often require specialized hardware, driving research into more accessible underwater communication solutions. While recent acoustic-based communication systems support text messaging on mobile devices, their low data rates severely limit broader applications. We present AquaScope, the first acoustic communication system capable of underwater image transmission on commodity mobile devices. To address the key challenges of underwater environments -- limited bandwidth and high transmission errors -- AquaScope employs and enhances generative image compression to improve compression efficiency, and integrates it with reliability-enhancement techniques at the physical layer to strengthen error resilience. We implemented AquaScope on the Android platform and demonstrated its feasibility for underwater image transmission. Experimental results show that AquaScope enables reliable, low-latency image transmission while preserving perceptual image quality, across various bandwidth-constrained and error-prone underwater conditions.

cs.NI↗

Enhancing Neural Adaptive Wireless Video Streaming via Lower-Layer Information Exposure and Online Tuning

Deep reinforcement learning (DRL) demonstrates its promising potential in the realm of adaptive video streaming and has recently received increasing attention. However, existing DRL-based methods for adaptive video streaming use only application (APP) layer information, adopt heuristic training methods, and train generalized neural networks with pre-collected data. This paper aims to boost the quality of experience (QoE) of adaptive wireless video streaming by using lower-layer information, deriving a rigorous training method, and adopting online tuning with real-time data. First, we formulate a more comprehensive and accurate adaptive wireless video streaming problem as an infinite stage discounted Markov decision process (MDP) problem by additionally incorporating past and lower-layer information, allowing a flexible tradeoff between QoE and costs for obtaining system information and solving the problem. In the offline scenario (only with pre-collected data), we propose an enhanced asynchronous advantage actor-critic (eA3C) method by jointly optimizing the parameters of parameterized policy and value function. Specifically, we build an eA3C network consisting of a policy network and a value network that can utilize cross-layer, past, and current information and jointly train the eA3C network using pre-collected samples. In the online scenario (with additional real-time data), we propose two continual learning-based online tuning methods for designing better policies for a specific user with different QoE and training time tradeoffs. Finally, experimental results show that the proposed offline policy can improve the QoE by 6.8~14.4% compared to the state-of-arts in the offline scenario, and the proposed online policies can further achieve 6~28% gains in QoE over the proposed offline policy in the online scenario.

cs.MM↗

An Optimization Framework for General Rate Splitting for General Multicast

Immersive video, such as virtual reality (VR) and multi-view videos, is growing in popularity. Its wireless streaming is an instance of general multicast, extending conventional unicast and multicast, whose effective design is still open. This paper investigates general rate splitting for general multicast. Specifically, we consider a multi-carrier single-cell wireless network where a multi-antenna base station (BS) communicates to multiple single-antenna users via general multicast. We consider linear beamforming at the BS and joint decoding at each user in the slow fading and fast fading scenarios. In the slow fading scenario, we consider the maximization of the weighted sum average rate, which is a challenging nonconvex stochastic problem with numerous variables. To reduce computational complexity, we decouple the original nonconvex stochastic problem into multiple nonconvex deterministic problems, one for each system channel state. Then, we propose an iterative algorithm for each deterministic problem to obtain a Karush-Kuhn-Tucker (KKT) point using the concave-convex procedure (CCCP). In the fast fading scenario, we consider the maximization of the weighted sum ergodic rate. This problem is more challenging than the one for the slow fading scenario, as it is not separable. First, we propose a stochastic iterative algorithm to obtain a KKT point using stochastic successive convex approximation (SSCA) and the exact penalty method. Then, we propose two low-complexity iterative algorithms to obtain feasible points with promising performance for two cases of channel distributions using approximation and CCCP. The proposed optimization framework generalizes the existing ones for rate splitting for various types of services. Finally, we numerically show substantial gains of the proposed solutions over existing schemes in both scenarios.

cs.IT↗

Rate Splitting for General Multicast

Immersive video, such as virtual reality (VR) and multi-view videos, is growing in popularity. Its wireless streaming is an instance of general multicast, extending conventional unicast and multicast, whose effective design is still open. This paper investigates the optimization of general rate splitting with linear beamforming for general multicast. Specifically, we consider a multi-carrier single-cell wireless network where a multi-antenna base station (BS) communicates to multiple single-antenna users via general multicast. Linear beamforming is adopted at the BS, and joint decoding is adopted at each user. We consider the maximization of the weighted sum rate, which is a challenging nonconvex problem. Then, we propose an iterative algorithm for the problem to obtain a KKT point using the concave-convex procedure (CCCP). The proposed optimization framework generalizes the existing ones for rate splitting for various types of services. Finally, we numerically show substantial gains of the proposed solutions over existing schemes and reveal the design insights of general rate splitting for general multicast.

cs.IT↗

Adaptive Streaming of 360 Videos with Perfect, Imperfect, and Unknown FoV Viewing Probabilities in Wireless Networks

This paper investigates adaptive streaming of one or multiple tiled 360 videos from a multi-antenna base station (BS) to one or multiple single-antenna users, respectively, in a multi-carrier wireless system. We aim to maximize the video quality while keeping rebuffering time small via encoding rate adaptation at each group of pictures (GOP) and transmission adaptation at each (transmission) slot. To capture the impact of field-of-view (FoV) prediction, we consider three cases of FoV viewing probability distributions, i.e., perfect, imperfect, and unknown FoV viewing probability distributions, and use the average total utility, worst average total utility, and worst total utility as the respective performance metrics. In the single-user scenario, we optimize the encoding rates of the tiles, encoding rates of the FoVs, and transmission beamforming vectors for all subcarriers to maximize the total utility in each case. In the multi-user scenario, we adopt rate splitting with successive decoding and optimize the encoding rates of the tiles, encoding rates of the FoVs, rates of the common and private messages, and transmission beamforming vectors for all subcarriers to maximize the total utility in each case. Then, we separate the challenging optimization problem into multiple tractable problems in each scenario. In the single-user scenario, we obtain a globally optimal solution of each problem using transformation techniques and the Karush-Kuhn-Tucker (KKT) conditions. In the multi-user scenario, we obtain a KKT point of each problem using the concave-convex procedure (CCCP). Finally, numerical results demonstrate that the proposed solutions achieve notable gains over existing schemes in all three cases. To the best of our knowledge, this is the first work revealing the impact of FoV prediction on the performance of adaptive streaming of tiled 360 videos.

cs.MM↗

Power-Efficient Wireless Streaming of Multi-Quality Tiled 360 VR Video in MIMO-OFDMA Systems

In this paper, we study the optimal wireless streaming of a multi-quality tiled 360 virtual reality (VR) video from a multi-antenna server to multiple single-antenna users in a multiple-input multiple-output (MIMO)-orthogonal frequency division multiple access (OFDMA) system. In the scenario without user transcoding, we jointly optimize beamforming and subcarrier, transmission power, and rate allocation to minimize the total transmission power. This problem is a challenging mixed discretecontinuous optimization problem. We obtain a globally optimal solution for small multicast groups, an asymptotically optimal solution for a large antenna array, and a suboptimal solution for the general case. In the scenario with user transcoding, we jointly optimize the quality level selection, beamforming, and subcarrier, transmission power, and rate allocation to minimize the weighted sum of the average total transmission power and the transcoding power. This problem is a two-timescale mixed discrete-continuous optimization problem, which is even more challenging than the problem for the scenario without user transcoding. We obtain a globally optimal solution for small multicast groups, an asymptotically optimal solution for a large antenna array, and a low-complexity suboptimal solution for the general case. Finally, numerical results demonstrate the significant gains of proposed solutions over the existing solutions. significant gains of proposed solutions over the existing solutions.

eess.SP↗

Optimal Streaming of 360 VR Videos with Perfect, Imperfect and Unknown FoV Viewing Probabilities

In this paper, we investigate wireless streaming of multi-quality tiled 360 virtual reality (VR) videos from a multi-antenna server to multiple single-antenna users in a multi-carrier system. To capture the impact of field-of-view (FoV) prediction, we consider three cases of FoV viewing probability distributions, i.e., perfect, imperfect and unknown FoV viewing probability distributions, and use the average total utility, worst average total utility and worst total utility as the respective performance metrics. We adopt rate splitting with successive decoding for efficient transmission of multiple sets of tiles of different 360 VR videos to their requesting users. In each case, we optimize the encoding rates of the tiles, minimum encoding rates of the FoVs, rates of the common and private messages and transmission beamforming vectors to maximize the total utility. The problems in the three cases are all challenging nonconvex optimization problems. We successfully transform the problem in each case into a difference of convex (DC) programming problem with a differentiable objective function, and obtain a suboptimal solution using concave-convex procedure (CCCP). Finally, numerical results demonstrate the proposed solutions achieve notable gains over existing schemes in all three cases. To the best of our knowledge, this is the first work revealing the impact of FoV prediction and its accuracy on the performance of streaming of multi-quality tiled 360 VR videos.

cs.IT↗