Searcharxiv⌕ Search

arXiv subjects

Cheng Fang

Publications and source records attributed to Cheng Fang.

At least 37 records · Page 2Linked to original sources

Deep Learning Inference on Heterogeneous Mobile Processors: Potentials and Pitfalls

There is a growing demand to deploy computation-intensive deep learning (DL) models on resource-constrained mobile devices for real-time intelligent applications. Equipped with a variety of processing units such as CPUs, GPUs, and NPUs, the mobile devices hold potential to accelerate DL inference via parallel execution across heterogeneous processors. Various efficient parallel methods have been explored to optimize computation distribution, achieve load balance, and minimize communication cost across processors. Yet their practical effectiveness in the dynamic and diverse real-world mobile environment is less explored. This paper presents a holistic empirical study to assess the capabilities and challenges associated with parallel DL inference on heterogeneous mobile processors. Through carefully designed experiments covering various DL models, mobile software/hardware environments, workload patterns, and resource availability, we identify limitations of existing techniques and highlight opportunities for cross-level optimization.

cs.LG↗

Stochastic parameter reduced-order model based on hybrid machine learning approaches

Establishing appropriate mathematical models for complex systems in natural phenomena not only helps deepen our understanding of nature but can also be used for state estimation and prediction. However, the extreme complexity of natural phenomena makes it extremely challenging to develop full-order models (FOMs) and apply them to studying many quantities of interest. In contrast, appropriate reduced-order models (ROMs) are favored due to their high computational efficiency and ability to describe the key dynamics and statistical characteristics of natural phenomena. Taking the viscous Burgers equation as an example, this paper constructs a Convolutional Autoencoder-Reservoir Computing-Normalizing Flow algorithm framework, where the Convolutional Autoencoder is used to construct latent space representations, and the Reservoir Computing-Normalizing Flow framework is used to characterize the evolution of latent state variables. In this way, a data-driven stochastic parameter reduced-order model is constructed to describe the complex system and its dynamic behavior.

cs.LG↗

Multi-task Meta Label Correction for Time Series Prediction

Time series classification faces two unavoidable problems. One is partial feature information and the other is poor label quality, which may affect model performance. To address the above issues, we create a label correction method to time series data with meta-learning under a multi-task framework. There are three main contributions. First, we train the label correction model with a two-branch neural network in the outer loop. While in the model-agnostic inner loop, we use pre-existing classification models in a multi-task way and jointly update the meta-knowledge so as to help us achieve adaptive labeling on complex time series. Second, we devise new data visualization methods for both image patterns of the historical data and data in the prediction horizon. Finally, we test our method with various financial datasets, including XOM, S\&P500, and SZ50. Results show that our method is more effective and accurate than some existing label correction techniques.

cs.LG↗

A Fingertip Sensor and Algorithms for Pre-touch Distance Ranging and Material Detection in Robotic Grasping

To enhance robotic grasping capabilities, we are developing new contactless fingertip sensors to measure distance in close proximity and simultaneously detect the type of material and the interior structure. These sensors are referred to as pre-touch dual-modal and dual-mechanism (PDM$^2$) sensors, and they operate using both pulse-echo ultrasound (US) and optoacoustic (OA) modalities. We present the design of a PDM$^2$ sensor that utilizes a pulsed laser beam and a customized ultrasound transceiver with a wide acoustic bandwidth for ranging and sensing. Both US and OA signals are collected simultaneously, triggered by the same laser pulse. To validate our design, we have fabricated a prototype of the PDM$^2$ sensor and integrated it into an object scanning system. We have also developed algorithms to enable the sensor, including time-of-flight (ToF) auto estimation, ranging rectification, sensor and system calibration, distance ranging, material/structure detection, and object contour detection and reconstruction. The experimental results demonstrate that the new PDM$^2$ sensor and its algorithms effectively enable the object scanning system to achieve satisfactory ranging and contour reconstruction performances, along with satisfying material/structure detection capabilities. In conclusion, the PDM$^2$ sensor offers a practical and powerful solution to improve grasping of unknown objects with the robotic gripper by providing advanced perception capabilities.

cs.RO↗

Dual Radar: A Multi-modal Dataset with Dual 4D Radar for Autonomous Driving

Radar has stronger adaptability in adverse scenarios for autonomous driving environmental perception compared to widely adopted cameras and LiDARs. Compared with commonly used 3D radars, the latest 4D radars have precise vertical resolution and higher point cloud density, making it a highly promising sensor for autonomous driving in complex environmental perception. However, due to the much higher noise than LiDAR, manufacturers choose different filtering strategies, resulting in an inverse ratio between noise level and point cloud density. There is still a lack of comparative analysis on which method is beneficial for deep learning-based perception algorithms in autonomous driving. One of the main reasons is that current datasets only adopt one type of 4D radar, making it difficult to compare different 4D radars in the same scene. Therefore, in this paper, we introduce a novel large-scale multi-modal dataset featuring, for the first time, two types of 4D radars captured simultaneously. This dataset enables further research into effective 4D radar perception algorithms.Our dataset consists of 151 consecutive series, most of which last 20 seconds and contain 10,007 meticulously synchronized and annotated frames. Moreover, our dataset captures a variety of challenging driving scenarios, including many road conditions, weather conditions, nighttime and daytime with different lighting intensities and periods. Our dataset annotates consecutive frames, which can be applied to 3D object detection and tracking, and also supports the study of multi-modal tasks. We experimentally validate our dataset, providing valuable results for studying different types of 4D radars. This dataset is released on https://github.com/adept-thu/Dual-Radar.

cs.CV↗

Enabling Resource-efficient AIoT System with Cross-level Optimization: A survey

The emerging field of artificial intelligence of things (AIoT, AI+IoT) is driven by the widespread use of intelligent infrastructures and the impressive success of deep learning (DL). With the deployment of DL on various intelligent infrastructures featuring rich sensors and weak DL computing capabilities, a diverse range of AIoT applications has become possible. However, DL models are notoriously resource-intensive. Existing research strives to realize near-/realtime inference of AIoT live data and low-cost training using AIoT datasets on resource-scare infrastructures. Accordingly, the accuracy and responsiveness of DL models are bounded by resource availability. To this end, the algorithm-system co-design that jointly optimizes the resource-friendly DL models and model-adaptive system scheduling improves the runtime resource availability and thus pushes the performance boundary set by the standalone level. Unlike previous surveys on resource-friendly DL models or hand-crafted DL compilers/frameworks with partially fine-tuned components, this survey aims to provide a broader optimization space for more free resource-performance tradeoffs. The cross-level optimization landscape involves various granularity, including the DL model, computation graph, operator, memory schedule, and hardware instructor in both on-device and distributed paradigms. Furthermore, due to the dynamic nature of AIoT context, which includes heterogeneous hardware, agnostic sensing data, varying user-specified performance demands, and resource constraints, this survey explores the context-aware inter-/intra-device controllers for automatic cross-level adaptation. Additionally, we identify some potential directions for resource-efficient AIoT systems. By consolidating problems and techniques scattered over diverse levels, we aim to help readers understand their connections and stimulate further discussions.

cs.LG↗

Reservoir Computing with Error Correction: Long-term Behaviors of Stochastic Dynamical Systems

The prediction of stochastic dynamical systems and the capture of dynamical behaviors are profound problems. In this article, we propose a data-driven framework combining Reservoir Computing and Normalizing Flow to study this issue, which mimics error modeling to improve traditional Reservoir Computing performance and integrates the virtues of both approaches. With few assumptions about the underlying stochastic dynamical systems, this model-free method successfully predicts the long-term evolution of stochastic dynamical systems and replicates dynamical behaviors. We verify the effectiveness of the proposed framework in several experiments, including the stochastic Van der Pal oscillator, El Niño-Southern Oscillation simplified model, and stochastic Lorenz system. These experiments consist of Markov/non-Markov and stationary/non-stationary stochastic processes which are defined by linear/nonlinear stochastic differential equations or stochastic delay differential equations. Additionally, we explore the noise-induced tipping phenomenon, relaxation oscillation, stochastic mixed-mode oscillation, and replication of the strange attractor.

math.DS↗

Spectral Observations and Modeling of a Solar White-light Flare Observed by CHASE

The heating mechanisms of solar white-light flares remain unclear. We present an X1.0 white-light flare on 2022 October 2 (SOL2022-10-02T20:25) observed by the Chinese \ha\ Solar Explorer (CHASE) that provides two-dimensional spectra in the visible light for the full solar disk with a seeing-free condition. The flare shows a prominent enhancement of $\sim$40\% in the photospheric \fe\ line at 6569.2 Å, and the nearby continuum also exhibits a maximum enhancement of $\sim$40\%. For the continuum near the \fe\ line at 6173 Å from the Helioseismic and Magnetic Imager (HMI) on board the Solar Dynamics Observatory (SDO), it is enhanced up to $\sim$20\%. At the white-light kernels, the \fe\ line at 6569.2 Å has a symmetric Gaussian profile that is still in absorption and the H$α$ line at 6562.8 Å displays a very broad emission profile with a central reversal plus a red or blue asymmetry. The white-light kernels are co-spatial with the microwave footpoint sources observed by the Expanded Owens Valley Solar Array (EOVSA) and the time profile of the white-light emission matches that of the hard X-ray emission above 30 keV from the Gamma-ray Burst Monitor (GBM) on Fermi. These facts indicate that the white-light emission is qualitatively related to a nonthermal electron beam. We also perform a radiative hydrodynamic simulation with the electron beam parameters constrained by the hard X-ray observations from Fermi/GBM. The result reveals that the white-light enhancement cannot be well explained by a pure electron-beam heating together with its induced radiative backwarming but may need additional heating sources such as \alfven\ waves.

astro-ph.SR↗

Statistical analysis of the Si I 6560.58 Å line observed by CHASE

The Si I 6560.58 Å line in the H$α$ blue wing is blended with a telluric absorption line from water vapor in ground-based observations. Recent observations with the space-based telescope CHASE provide a new window to study this line. We aim to study the Si I line statistically and to explore possible diagnostics. We select three scannings in the CHASE observations, and measure the equivalent width (EW) and the full width at half maximum (FWHM) for each pixel on the solar disk. We then calculate the theoretical EW and FWHM from the VALC model. An active region is also studied in particular for difference in the quiet Sun and the sunspots. The Si I line is formed at the bottom of the photosphere. The EW of this line increases from the disk center to $μ$ = 0.2, and then decreases toward the solar limb, while the FWHM shows a monotonically increasing trend. Theoretically predicted EW agrees well with observations, while the predicted FWHM is far smaller due to the absence of unresolved turbulence in models. The macroturbulent velocity is estimated to be 2.80 km s$^{-1}$ at the disk center, and increases to 3.52 km s$^{-1}$ at $μ$ = 0.2. We do not find any response to flare heating in current observations. Doppler shifts and line widths of the Si I 6560.58 Å and Fe I 6569.21 Å lines can be used to study the mass flows and turbulence of the different photospheric layers. The Si I line has good potentials to diagnose the dynamics and energy transport in the photosphere.

astro-ph.SR↗

An end-to-end deep learning approach for extracting stochastic dynamical systems with $α$-stable Lévy noise

Recently, extracting data-driven governing laws of dynamical systems through deep learning frameworks has gained a lot of attention in various fields. Moreover, a growing amount of research work tends to transfer deterministic dynamical systems to stochastic dynamical systems, especially those driven by non-Gaussian multiplicative noise. However, lots of log-likelihood based algorithms that work well for Gaussian cases cannot be directly extended to non-Gaussian scenarios which could have high error and low convergence issues. In this work, we overcome some of these challenges and identify stochastic dynamical systems driven by $α$-stable Lévy noise from only random pairwise data. Our innovations include: (1) designing a deep learning approach to learn both drift and diffusion coefficients for Lévy induced noise with $α$ across all values, (2) learning complex multiplicative noise without restrictions on small noise intensity, (3) proposing an end-to-end complete framework for stochastic systems identification under a general input data assumption, that is, $α$-stable random variable. Finally, numerical experiments and comparisons with the non-local Kramers-Moyal formulas with moment generating function confirm the effectiveness of our method.

stat.ML↗

BRIDGE: Byzantine-resilient Decentralized Gradient Descent

Machine learning has begun to play a central role in many applications. A multitude of these applications typically also involve datasets that are distributed across multiple computing devices/machines due to either design constraints (e.g., multiagent systems) or computational/privacy reasons (e.g., learning on smartphone data). Such applications often require the learning tasks to be carried out in a decentralized fashion, in which there is no central server that is directly connected to all nodes. In real-world decentralized settings, nodes are prone to undetected failures due to malfunctioning equipment, cyberattacks, etc., which are likely to crash non-robust learning algorithms. The focus of this paper is on robustification of decentralized learning in the presence of nodes that have undergone Byzantine failures. The Byzantine failure model allows faulty nodes to arbitrarily deviate from their intended behaviors, thereby ensuring designs of the most robust of algorithms. But the study of Byzantine resilience within decentralized learning, in contrast to distributed learning, is still in its infancy. In particular, existing Byzantine-resilient decentralized learning methods either do not scale well to large-scale machine learning models, or they lack statistical convergence guarantees that help characterize their generalization errors. In this paper, a scalable, Byzantine-resilient decentralized machine learning framework termed Byzantine-resilient decentralized gradient descent (BRIDGE) is introduced. Algorithmic and statistical convergence guarantees for one variant of BRIDGE are also provided in the paper for both strongly convex problems and a class of nonconvex problems. In addition, large-scale decentralized learning experiments are used to establish that the BRIDGE framework is scalable and it delivers competitive results for Byzantine-resilient convex and nonconvex learning.

stat.ML↗

The Chinese Hα Solar Explorer (CHASE) mission: An overview

The Chinese Hα Solar Explorer (CHASE), dubbed "Xihe" - Goddess of the Sun, was launched on October 14, 2021 as the first solar space mission of China National Space Administration (CNSA). The CHASE mission is designed to test a newly developed satellite platform and to acquire the spectroscopic observations in the Hα waveband. The Hα Imaging Spectrograph (HIS) is the scientific payload of the CHASE satellite. It consists of two observational modes: raster scanning mode and continuum imaging mode. The raster scanning mode obtains full-Sun or region-of-interest spectral images from 6559.7 to 6565.9 Å and from 6567.8 to 6570.6 Å with 0.024 Å pixel spectral resolution and 1 minute temporal resolution. The continuum imaging mode obtains photospheric images in continuum around 6689 Å with the full width at half maximum of 13.4 Å. The CHASE mission will advance our understanding of the dynamics of solar activity in the photosphere and chromosphere. In this paper, we present an overview of the CHASE mission including the scientific objectives, HIS instrument overview, data calibration flow, and first results of on-orbit observations.

astro-ph.SR↗

Calibration procedures for the CHASE/HIS science data

The Hα line is an important optical line in solar observations containing the information from the photosphere to the chromosphere. To study the mechanisms of solar eruptions and the plasma dynamics in the lower atmosphere, the Chinese Hα Solar Explorer (CHASE) was launched into a Sun-synchronous orbit on October 14, 2021. The scientific payload of the CHASE satellite is the Hα Imaging Spectrograph (HIS). The CHASE/HIS acquires, for the first time, seeing-free Hα spectroscopic observations with high spectral and temporal resolutions. It consists of two observational modes. The raster scanning mode provides full-Sun or region-of-interest spectra at Hα (6559.7-6565.9 Å) and Fe I (6567.8-6570.6 Å) wavebands. The continuum imaging mode obtains full-Sun photospheric images at around 6689 Å. In this paper, we present detailed calibration procedures for the CHASE/HIS science data, including the dark-field and flat-field correction, slit image curvature correction, wavelength and intensity calibration, and coordinate transformation. The higher-level data products can be directly used for scientific research.

astro-ph.SR↗

Refractive index sensing with hybrid surfaces of photonic crystals and dielectric microsphere monolayers

In this work, a refractive index (RI) sensor with an effective integration of colorimetric detection and optical sensing capabilities has been developed. Colorimetric detection relies on the sensitivity of the structural color of photonic crystal (PC) substrates to the changes in background RI, while the optical sensing is performed by measuring the magnification abilities of the dielectric microspheres, which depends on the position of the photonic nanojet. Based on this concept, we have successfully assembled 35 μm-diameter barium titanate glass microspheres, 4.9 μm-diameter polystyrene and silica microsphere monolayers on 1D or 2D PC substrates to perform RI sensing in various liquids. In addition, the developed RI sensor is highly compatible with commercial optical microscopes and applicable for RI sensing in areas as small as tens of square microns.

physics.optics↗

Generalized Conflict-directed Search for Optimal Ordering Problems

Solving planning and scheduling problems for multiple tasks with highly coupled state and temporal constraints is notoriously challenging. An appealing approach to effectively decouple the problem is to judiciously order the events such that decisions can be made over sequences of tasks. As many problems encountered in practice are over-constrained, we must instead find relaxed solutions in which certain requirements are dropped. This motivates a formulation of optimality with respect to the costs of relaxing constraints and the problem of finding an optimal ordering under which this relaxing cost is minimum. In this paper, we present Generalized Conflict-directed Ordering (GCDO), a branch-and-bound ordering method that generates an optimal total order of events by leveraging the generalized conflicts of both inconsistency and suboptimality from sub-solvers for cost estimation and solution space pruning. Due to its ability to reason over generalized conflicts, GCDO is much more efficient in finding high-quality total orders than the previous conflict-directed approach CDITO. We demonstrate this by benchmarking on temporal network configuration problems, which involves managing networks over time and makes necessary tradeoffs between network flows against CDITO and Mixed Integer-Linear Programing (MILP). Our algorithm is able to solve two orders of magnitude more benchmark problems to optimality and twice the problems compared to CDITO and MILP within a runtime limit, respectively.

cs.AI↗

Efficiently Exploring Ordering Problems through Conflict-directed Search

In planning and scheduling, solving problems with both state and temporal constraints is hard since these constraints may be highly coupled. Judicious orderings of events enable solvers to efficiently make decisions over sequences of actions to satisfy complex hybrid specifications. The ordering problem is thus fundamental to planning. Promising recent works have explored the ordering problem as search, incorporating a special tree structure for efficiency. However, such approaches only reason over partial order specifications. Having observed that an ordering is inconsistent with respect to underlying constraints, prior works do not exploit the tree structure to efficiently generate orderings that resolve the inconsistency. In this paper, we present Conflict-directed Incremental Total Ordering (CDITO), a conflict-directed search method to incrementally and systematically generate event total orders given ordering relations and conflicts returned by sub-solvers. Due to its ability to reason over conflicts, CDITO is much more efficient than Incremental Total Ordering. We demonstrate this by benchmarking on temporal network configuration problems that involve routing network flows and allocating bandwidth resources over time.

cs.AI↗

Variations of the 3-D coronal magnetic field associated with the X3.4-class solar flare event of AR 10930

The variations of the 3-D coronal magnetic fields associated with the X3.4-class flare of active region 10930 are studied in this paper. The coronal magnetic field data are reconstructed from the photospheric vector magnetograms obtained by the Hinode satellite and using the nonlinear force-free field extrapolation method developed in our previous work (He et al., 2011). The 3-D force-free factor $α$, 3-D current density, and 3-D magnetic energy density are employed to analyze the coronal data. The distributions of $α$ and current density reveal a prominent magnetic connectivity with strong negative $α$ values and strong current density before the flare. This magnetic connectivity extends along the main polarity inversion line and is found to be totally broken after the flare. The distribution variation of magnetic energy density reveals the redistribution of magnetic energy before and after the flare. In the lower space of the modeling volume the increase of magnetic energy dominates, and in the higher space the decrease of energy dominates. The comparison with the flare onset imaging observation exhibits that the breaking site of the magnetic connectivity and site with the highest values of energy density increase coincide with the location of flare initial eruption. We conclude that a cramped positive $α$ region appearing in the photosphere causes the breaking of the magnetic connectivity. A scenario for flare initial eruption is proposed in which the Lorentz force acting on the isolated electric current at the magnetic connectivity breaking site lifts the associated plasmas and causes the initial ejection.

astro-ph.SR↗

Bidirectional outflows as evidence of magnetic reconnection leading to a solar microflare

Magnetic reconnection is a rapid energy release process that is believed to be responsible for flares on the Sun and stars. Nevertheless, such flare-related reconnection is mostly detected to occur in the corona, while there have been few studies concerning the reconnection in the chromosphere or photosphere. Here we present both spectroscopic and imaging observations of magnetic reconnection in the chromosphere leading to a microflare. During the flare peak time, chromospheric line profiles show significant blueshifted/redshifted components on the two sides of the flaring site, corresponding to upflows and downflows with velocities of $\pm$(70--80) km s$^{-1}$, comparable with the local Alfvén speed as expected by the reconnection in the chromosphere. The three-dimensional nonlinear force-free field configuration further discloses twisted field lines (a flux rope) at a low altitude, cospatial with the dark threads in He I 10830 Å images. The instability of the flux rope may initiate the flare-related reconnection. These observations provide clear evidence of magnetic reconnection in the chromosphere and show the similar mechanisms of a microflare to those of major flares.

astro-ph.SR↗