Searcharxiv⌕ Search

arXiv subjects

Shi Yan

Publications and source records attributed to Shi Yan.

At least 37 records · Page 2Linked to original sources

Den-SOFT: Dense Space-Oriented Light Field DataseT for 6-DOF Immersive Experience

We have built a custom mobile multi-camera large-space dense light field capture system, which provides a series of high-quality and sufficiently dense light field images for various scenarios. Our aim is to contribute to the development of popular 3D scene reconstruction algorithms such as IBRnet, NeRF, and 3D Gaussian splitting. More importantly, the collected dataset, which is much denser than existing datasets, may also inspire space-oriented light field reconstruction, which is potentially different from object-centric 3D reconstruction, for immersive VR/AR experiences. We utilized a total of 40 GoPro 10 cameras, capturing images of 5k resolution. The number of photos captured for each scene is no less than 1000, and the average density (view number within a unit sphere) is 134.68. It is also worth noting that our system is capable of efficiently capturing large outdoor scenes. Addressing the current lack of large-space and dense light field datasets, we made efforts to include elements such as sky, reflections, lights and shadows that are of interest to researchers in the field of 3D reconstruction during the data capture process. Finally, we validated the effectiveness of our provided dataset on three popular algorithms and also integrated the reconstructed 3DGS results into the Unity engine, demonstrating the potential of utilizing our datasets to enhance the realism of virtual reality (VR) and create feasible interactive spaces. The dataset is available at our project website.

cs.CV↗

Explainable Gated Bayesian Recurrent Neural Network for Non-Markov State Estimation

The optimality of Bayesian filtering relies on the completeness of prior models, while deep learning holds a distinct advantage in learning models from offline data. Nevertheless, the current fusion of these two methodologies remains largely ad hoc, lacking a theoretical foundation. This paper presents a novel solution, namely an explainable gated Bayesian recurrent neural network specifically designed to state estimation under model mismatches. Firstly, we transform the non-Markov state-space model into an equivalent first-order Markov model with memory. It is a generalized transformation that overcomes the limitations of the first-order Markov property and enables recursive filtering. Secondly, by deriving a data-assisted joint state-memory-mismatch Bayesian filtering, we design a Bayesian gated framework that includes a memory update gate for capturing the temporal regularities in state evolution, a state prediction gate with the evolution mismatch compensation, and a state update gate with the observation mismatch compensation. The Gaussian approximation implementation of the filtering process within the gated framework is derived, taking into account the computational efficiency. Finally, the corresponding internal neural network structures and end-to-end training methods are designed. The Bayesian filtering theory enhances the interpretability of the proposed gated network, enabling the effective integration of offline data and prior models within functionally explicit gated units. In comprehensive experiments, including simulations and real-world datasets, the proposed gated network demonstrates superior estimation performance compared to benchmark filters and state-of-the-art deep learning filtering methods.

eess.SP↗

LightPainter: Interactive Portrait Relighting with Freehand Scribble

Recent portrait relighting methods have achieved realistic results of portrait lighting effects given a desired lighting representation such as an environment map. However, these methods are not intuitive for user interaction and lack precise lighting control. We introduce LightPainter, a scribble-based relighting system that allows users to interactively manipulate portrait lighting effect with ease. This is achieved by two conditional neural networks, a delighting module that recovers geometry and albedo optionally conditioned on skin tone, and a scribble-based module for relighting. To train the relighting module, we propose a novel scribble simulation procedure to mimic real user scribbles, which allows our pipeline to be trained without any human annotations. We demonstrate high-quality and flexible portrait lighting editing capability with both quantitative and qualitative experiments. User study comparisons with commercial lighting editing tools also demonstrate consistent user preference for our method.

cs.CV↗

Analysis on Blockchain Consensus Mechanism Based on Proof of Work and Proof of Stake

In the white book of Bitcion, Satoshi Nakamoto described a bitcoin system that can realize point-to-point online payment without a third-party organization. After supporting this magical application scenario and subverting the traditional centralized system, the blockchain technology has attracted worldwide attention, triggered a research upsurge of blockchain consensus algorithm, and produced a large number of innovative applications. Although various consensus algorithms continue to evolve with the iteration of blockchain products and applications, Proof of Work (POW) and Proof of Skake (POS) algorithms are still the core of consensus algorithms. This paper discusses two algorithms of POW and POS in blockchain consensus mechanism, and analyzes the advantages and the existing problems of the two consensus mechanisms. Since consensus mechanism is the main focus of blockchain technology and has many influencing factors, this paper discusses the current problems and some improved ideas, and selects some typical algorithms for a more systematic introduction. In addition, some important issues related to safety and performance are also discussed. This paper provides the researchers a great reference on blockchain consensus mechanism.

cs.CR↗

Communication and Computation Assisted Sensing Information Freshness Performance Analysis in Vehicular Networks

The timely sharing of raw sensing information in the vehicular networks (VNETs) is essential to safety. In order to improve the freshness of sensing information, joint scheduling of multi-dimensional resources such as communication and computation is required. However, the complex relevance among multi-dimensional resources is still unclear, and it is difficult to achieve efficient resource utilization. In this paper, we present a theoretical analysis for a novel metric Age of Information (AoI) on a communication and computation assisted spatial-temporal model. An uplink VNETs scenario where Road Side Units (RSUs) are deployed with computational resource is considered. The transmission and computation process is unified into a two-stage tandem queue and the expression of the average AoI is derived. The network interference is analyzed by modeling the VNETs as Cox Poisson Point Process based on stochastic geometry and the closed-form solution of the coverage probability and the expected data rate performance under the constraints of transmission resources is obtained. The simulation results reveal the basic relationship between communication and computation capacity and show that communication and computation should reach a tradeoff to improve resource utilization while ensuring real-time information requirement.

cs.NI↗

Give me a knee radiograph, I will tell you where the knee joint area is: a deep convolutional neural network adventure

Knee pain is undoubtedly the most common musculoskeletal symptom that impairs quality of life, confines mobility and functionality across all ages. Knee pain is clinically evaluated by routine radiographs, where the widespread adoption of radiographic images and their availability at low cost, make them the principle component in the assessment of knee pain and knee pathologies, such as arthritis, trauma, and sport injuries. However, interpretation of the knee radiographs is still highly subjective, and overlapping structures within the radiographs and the large volume of images needing to be analyzed on a daily basis, make interpretation challenging for both naive and experienced practitioners. There is thus a need to implement an artificial intelligence strategy to objectively and automatically interpret knee radiographs, facilitating triage of abnormal radiographs in a timely fashion. The current work proposes an accurate and effective pipeline for autonomous detection, localization, and classification of knee joint area in plain radiographs combining the You Only Look Once (YOLO v3) deep convolutional neural network with a large and fully-annotated knee radiographs dataset. The present work is expected to stimulate more interest from the deep learning computer vision community to this pragmatic and clinical application.

eess.IV↗

Measuring microlensing parallax via simultaneous observations from Chinese Space Station Telescope and Roman Telescope

Simultaneous observations from two spatially well-separated telescopes can lead to the measurements of the microlensing parallax parameter, an important quantity toward the determinations of the lens mass. The separation between Earth and Sun-Earth L2 point, $\sim0.01$ AU, is ideal for parallax measurements of short and ultra-short ($\sim$1\,hr to 10\,days) microlensing events, which are candidates of free-floating planet (FFP) events. In this work, we study the potential of doing so in the context of two proposed space-based missions, the Chinese Space Station Telescope (CSST) in a Leo orbit and the Nancy Grace Roman Space Telescope (\emph{Roman}) at L2. We show that the joint observations of the two can directly measure the microlensing parallax of nearly all FFP events with timescales $t_{\rm E}\lesssim$ 10\,days as well as planetary (and stellar binary) events that show caustic crossing features. The potential of using CSST alone in measuring microlensing parallax is also discussed.

astro-ph.EP↗

A fast method for simultaneous reconstruction and segmentation in X-ray CT application

In this paper, we propose a fast method for simultaneous reconstruction and segmentation (SRS) in X-ray computed tomography (CT). Our work is based on the SRS model where Bayes' rule and the maximum a posteriori (MAP) are used on hidden Markov measure field model (HMMFM). The original method leads to a logarithmic-summation (log-sum) term, which is non-separable to the classification index. The minimization problem in the model was solved by using constrained gradient descend method, Frank-Wolfe algorithm, which is very time-consuming especially when dealing with large-scale CT problems. The starting point of this paper is the commutativity of log-sum operations, where the log-sum problem could be transformed into a sum-log problem by introducing an auxiliary variable. The corresponding sum-log problem for the SRS model is separable. After applying alternating minimization method, this problem turns into several easy-to-solve convex sub-problems. In the paper, we also study an improved model by adding Tikhonov regularization, and give some convergence results. Experimental results demonstrate that the proposed algorithms could produce comparable results with the original SRS method with much less CPU time.

cs.CV↗

Tradeoff between Ergodic Rate and Delivery Latency in Fog Radio Access Networks

Wireless content caching has recently been considered as an efficient way in fog radio access networks (FRANs) to alleviate the heavy burden on capacity-limited fronthaul links and reduce delivery latency. In this paper, an advanced minimal delay association policy is proposed to minimize latency while guaranteeing spectral efficiency in F-RANs. By utilizing stochastic geometry and queueing theory, closed-form expressions of successful delivery probability, average ergodic rate, and average delivery latency are derived, where both the traditional association policy based on accessing the base station with maximal received power and the proposed minimal delay association policy are concerned. Impacts of key operating parameters on the aforementioned performance metrics are exploited. It is shown that the proposed association policy has a better delivery latency than the traditional association policy. Increasing the cache size of fog-computing based access points (F-APs) can more significantly reduce average delivery latency, compared with increasing the density of F-APs. Meanwhile, the latter comes at the expense of decreasing average ergodic rate. This implies the deployment of large cache size at F-APs rather than high density of F-APs can promote performance effectively in F-RANs.

eess.SP↗

Deep Reinforcement Learning Based Mode Selection and Resource Allocation for Cellular V2X Communications

Cellular vehicle-to-everything (V2X) communication is crucial to support future diverse vehicular applications. However, for safety-critical applications, unstable vehicle-to-vehicle (V2V) links and high signalling overhead of centralized resource allocation approaches become bottlenecks. In this paper, we investigate a joint optimization problem of transmission mode selection and resource allocation for cellular V2X communications. In particular, the problem is formulated as a Markov decision process, and a deep reinforcement learning (DRL) based decentralized algorithm is proposed to maximize the sum capacity of vehicle-to-infrastructure users while meeting the latency and reliability requirements of V2V pairs. Moreover, considering training limitation of local DRL models, a two-timescale federated DRL algorithm is developed to help obtain robust model. Wherein, the graph theory based vehicle clustering algorithm is executed on a large timescale and in turn the federated learning algorithm is conducted on a small timescale. Simulation results show that the proposed DRL-based algorithm outperforms other decentralized baselines, and validate the superiority of the two-timescale federated DRL algorithm for newly activated V2V pairs.

cs.NI↗

Mode Selection and Resource Allocation in Sliced Fog Radio Access Networks: A Reinforcement Learning Approach

The mode selection and resource allocation in fog radio access networks (F-RANs) have been advocated as key techniques to improve spectral and energy efficiency. In this paper, we investigate the joint optimization of mode selection and resource allocation in uplink F-RANs, where both of the traditional user equipments (UEs) and fog UEs are served by constructed network slice instances. The concerned optimization is formulated as a mixed-integer programming problem, and both the orthogonal and multiplexed subchannel allocation strategies are proposed to guarantee the slice isolation. Motivated by the development of machine learning, two reinforcement learning based algorithms are developed to solve the original high complexity problem under traditional and fog UEs' specific performance requirements. The basic idea of the proposals is to generate a good mode selection policy according to the immediate reward fed back by an environment. Simulation results validate the benefits of our proposed algorithms and show that a tradeoff between system power consumption and queue delay can be achieved.

cs.NI↗

Convexity Shape Prior for Level Set based Image Segmentation Method

We propose a geometric convexity shape prior preservation method for variational level set based image segmentation methods. Our method is built upon the fact that the level set of a convex signed distanced function must be convex. This property enables us to transfer a complicated geometrical convexity prior into a simple inequality constraint on the function. An active set based Gauss-Seidel iteration is used to handle this constrained minimization problem to get an efficient algorithm. We apply our method to region and edge based level set segmentation models including Chan-Vese (CV) model with guarantee that the segmented region will be convex. Experimental results show the effectiveness and quality of the proposed model and algorithm.

cs.CV↗

Phonemic evidence reveals interwoven evolution of Chinese dialects

Han Chinese experienced substantial population migrations and admixture in history, yet little is known about the evolutionary process of Chinese dialects. Here, we used phylogenetic approaches and admixture inference to explicitly decompose the underlying structure of the diversity of Chinese dialects, based on the total phoneme inventories of 140 dialect samples from seven traditional dialect groups: Mandarin, Wu, Xiang, Gan, Hakka, Min and Yue. We found a north-south gradient of phonemic differences in Chinese dialects induced from historical population migrations. We also quantified extensive horizontal language transfers among these dialects, corresponding to the complicated socio-genetic history in China. We finally identified that the middle latitude dialects of Xiang, Gan and Hakka were formed by admixture with other four dialects. Accordingly, the middle-latitude areas in China were a linguistic melting pot of northern and southern Han populations. Our study provides a detailed phylogenetic and historical context against family-tree model in China.

q-bio.PE↗

FOTS: Fast Oriented Text Spotting with a Unified Network

Incidental scene text spotting is considered one of the most difficult and valuable challenges in the document analysis community. Most existing methods treat text detection and recognition as separate tasks. In this work, we propose a unified end-to-end trainable Fast Oriented Text Spotting (FOTS) network for simultaneous detection and recognition, sharing computation and visual information among the two complementary tasks. Specially, RoIRotate is introduced to share convolutional features between detection and recognition. Benefiting from convolution sharing strategy, our FOTS has little computation overhead compared to baseline text detection network, and the joint training method learns more generic features to make our method perform better than these two-stage methods. Experiments on ICDAR 2015, ICDAR 2017 MLT, and ICDAR 2013 datasets demonstrate that the proposed method outperforms state-of-the-art methods significantly, which further allows us to develop the first real-time oriented text spotting system which surpasses all previous state-of-the-art results by more than 5% on ICDAR 2015 text spotting task while keeping 22.6 fps.

cs.CV↗

An Evolutionary Game for User Access Mode Selection in Fog Radio Access Networks

The fog radio access network (F-RAN) is a promising paradigm for the fifth generation wireless communication systems to provide high spectral efficiency and energy efficiency. Characterizing users to select an appropriate communication mode among fog access point (F-AP), and device-to-device (D2D) in F-RANs is critical for performance optimization. Using evolutionary game theory, we investigate the dynamics of user access mode selection in F-RANs. Specifically, the competition among groups of potential users space is formulated as a dynamic evolutionary game, and the evolutionary equilibrium is the solution to this game. Stochastic geometry tool is used to derive the proposals' payoff expressions for both F-AP and D2D users by taking into account the different nodes locations, cache sizes as well as the delay cost. The analytical results obtained from the game model are evaluated via simulations, which show that the evolutionary game based access mode selection algorithm can reach a much higher payoff than the max rate based algorithm.

cs.IT↗

Joint Beamforming and Power Splitting Control in Downlink Cooperative SWIPT NOMA Systems

This paper investigates the application of simultaneous wireless information and power transfer (SWIPT) to cooperative non-orthogonal multiple access (NOMA). A new cooperative multiple-input single-output (MISO) SWIPT NOMA protocol is proposed, where a user with a strong channel condition acts as an energy-harvesting (EH) relay to help a user with a poor channel condition. The power splitting (PS) scheme is adopted at the EH relay. By jointly optimizing the PS ratio and the beamforming vectors, the design objective is to maximize the data rate of the "strong user" while satisfying the QoS requirement of the "weak user". It boils down to a challenging nonconvex problem. To resolve this issue, the semidefinite relaxation (SDR) technique is applied to relax the quadratic terms related with the beamformers, and then it is solved to its global optimality by two-dimensional exhaustive search. We prove the rank-one optimality, which establishes the equivalence between the relaxed problem and the original one. To further reduce the high complexity due to the exhaustive search, an iterative algorithm based on successive convex approximation (SCA) is proposed, which can at least attain its stationary point efficiently. In view of the potential application scenarios, e.g., IoT, the single-input single-output (SISO) case of the cooperative SWIPT NOMA system is also studied. The formulated problem is proved to be strictly unimodal with respect to the PS ratio. Hence, a golden section search (GSS) based algorithm with closed-form solution at each step is proposed to find the unique global optimal solution. It is worth pointing out that the SCA method can also converge to the optimal solution in SISO cases. In the numerical simulation, the proposed algorithm is numerically shown to converge within a few iterations, and the SWIPT-aided NOMA protocol outperforms the existing transmission protocols.

cs.IT↗

User Access Mode Selection in Fog Computing Based Radio Access Networks

Fog computing based radio access network is a promising paradigm for the fifth generation wireless communication system to provide high spectral and energy efficiency. With the help of the new designed fog computing based access points (F-APs), the user-centric objectives can be achieved through the adaptive technique and will relieve the load of fronthaul and alleviate the burden of base band unit pool. In this paper, we derive the coverage probability and ergodic rate for both F-AP users and device-to-device users by taking into account the different nodes locations, cache sizes as well as user access modes. Particularly, the stochastic geometry tool is used to derive expressions for above performance metrics. Simulation results validate the accuracy of our analysis and we obtain interesting trade-offs that depend on the effect of the cache size, user node density, and the quality of service constrains on the different performance metrics.

cs.IT↗

Inter-tier Interference Suppression in Heterogeneous Cloud Radio Access Networks

Incorporating cloud computing into heterogeneous networks, the heterogeneous cloud radio access network (H-CRAN) has been proposed as a promising paradigm to enhance both spectral and energy efficiencies. Developing interference suppression strategies is critical for suppressing the inter-tier interference between remote radio heads (RRHs) and a macro base station (MBS) in H-CRANs. In this paper, inter-tier interference suppression techniques are considered in the contexts of collaborative processing and cooperative radio resource allocation (CRRA). In particular, interference collaboration (IC) and beamforming (BF) are proposed to suppress the inter-tier interference, and their corresponding performance is evaluated. Closed-form expressions for the overall outage probabilities, system capacities, and average bit error rates under these two schemes are derived. Furthermore, IC and BF based CRRA optimization models are presented to maximize the RRH-accessed users' sum rates via power allocation, which is solved with convex optimization. Simulation results demonstrate that the derived expressions for these performance metrics for IC and BF are accurate; and the relative performance between IC and BF schemes depends on system parameters, such as the number of antennas at the MBS, the number of RRHs, and the target signal-to-interference-plus-noise ratio threshold. Furthermore, it is seen that the sum rates of IC and BF schemes increase almost linearly with the transmit power threshold under the proposed CRRA optimization solution.

cs.IT↗