SearcharxivSearch

arXiv subjects

Shan Lin

Publications and source records attributed to Shan Lin.

At least 73 records · Page 4Linked to original sources

Reducing Annotating Load: Active Learning with Synthetic Images in Surgical Instrument Segmentation

Accurate instrument segmentation in endoscopic vision of robot-assisted surgery is challenging due to reflection on the instruments and frequent contacts with tissue. Deep neural networks (DNN) show competitive performance and are in favor in recent years. However, the hunger of DNN for labeled data poses a huge workload of annotation. Motivated by alleviating this workload, we propose a general embeddable method to decrease the usage of labeled real images, using active generated synthetic images. In each active learning iteration, the most informative unlabeled images are first queried by active learning and then labeled. Next, synthetic images are generated based on these selected images. The instruments and backgrounds are cropped out and randomly combined with each other with blending and fusion near the boundary. The effectiveness of the proposed method is validated on 2 sinus surgery datasets and 1 intraabdominal surgery dataset. The results indicate a considerable improvement in performance, especially when the budget for annotation is small. The effectiveness of different types of synthetic images, blending methods, and external background are also studied. All the code is open-sourced at: https://github.com/HaonanPeng/active_syn_generator.

cs.CV

Multi-frame Feature Aggregation for Real-time Instrument Segmentation in Endoscopic Video

Deep learning-based methods have achieved promising results on surgical instrument segmentation. However, the high computation cost may limit the application of deep models to time-sensitive tasks such as online surgical video analysis for robotic-assisted surgery. Moreover, current methods may still suffer from challenging conditions in surgical images such as various lighting conditions and the presence of blood. We propose a novel Multi-frame Feature Aggregation (MFFA) module to aggregate video frame features temporally and spatially in a recurrent mode. By distributing the computation load of deep feature extraction over sequential frames, we can use a lightweight encoder to reduce the computation costs at each time step. Moreover, public surgical videos usually are not labeled frame by frame, so we develop a method that can randomly synthesize a surgical frame sequence from a single labeled frame to assist network training. We demonstrate that our approach achieves superior performance to corresponding deeper segmentation models on two public surgery datasets.

cs.CV

Dimensional Control of Octahedral Tilt in SrRuO3 via Infinite-layered Oxides

Manipulation of octahedral distortion at atomic length scale is an effective means to tune the physical ground states of functional oxides. Previous work demonstrates that epitaxial strain and film thickness are variable parameters to modify the octahedral rotation and tilt. However, selective control of bonding geometry by structural propagation from adjacent layers is rarely studied. Here we propose a new route to tune the ferromagnetic response in SrRuO3 (SRO) ultrathin layers by oxygen coordination of adjacent SrCuO2 (SCO) layers. The infinite-layered CuO2 in SCO exhibits a structural transformation from "planar-type" to "chain-type" as reducing film thickness. These two orientations dramatically modify the polyhedral connectivity at the interface, thus altering the octahedral distortion of SRO. The local structural variation changes the spin state of Ru and hybridization strength between Ru 4d and O 2p orbitals, leading to a significant change in the magnetoresistance and anomalous Hall resistivity of SRO layers. These findings could launch further investigations into adaptive control of magnetoelectric properties in quantum oxide heterostructures using oxygen coordination.

cond-mat.mtrl-sci

Near-Room Temperature Ferromagnetic Insulating State in Highly Distorted LaCoO2.5 with CoO5 Square Pyramids

Dedicated control of oxygen vacancies is an important route to functionalizing complex oxide films. It is well-known that tensile strain significantly lowers the oxygen vacancy formation energy, whereas compressive strain plays a minor role. Thus, atomically reconstruction by extracting oxygen from a compressive-strained film is challenging. Here we report an unexpected LaCoO2.5 phase with a zigzag-like oxygen vacancy ordering through annealing a compressive-strained LaCoO3 in vacuum. The synergetic tilt and distortion of CoO5 square pyramids with large La and Co shifts are quantified using scanning transmission electron microscopy. The large in-plane expansion of CoO5 square pyramids weaken the crystal-field splitting and facilitated the ordered high-spin state of Co2+, which produces an insulating ferromagnetic state with a Curie temperature of ~284 K and a saturation magnetization of ~0.25 μB/Co. These results demonstrate that extracting targeted oxygen from a compressive-strained oxide provides an opportunity for creating unexpected crystal structures and novel functionalities.

cond-mat.mtrl-sci

Structural Twinning-induced Insulating Phase in CrN (111) Films

Electronic states of a correlated material can be effectively modified by structural variations delivered from a single-crystal substrate. In this letter, we show that the CrN films grown on MgO (001) substrates have a (001) orientation, whereas the CrN films on α-Al2O3 (0001) substrates are oriented along (111) direction parallel to the surface normal. Transport properties of CrN films are remarkably different depending on crystallographic orientations. The critical thickness for the metal-insulator transition (MIT) in CrN 111 films is significantly larger than that of CrN 001 films. In contrast to CrN 001 films without apparent defects, scanning transmission electron microscopy results reveal that CrN 111 films exhibit strain-induced structural defects, e. g. the periodic horizontal twinning domains, resulting in an increased electron scattering facilitating an insulating state. Understanding the key parameters that determine the electronic properties of ultrathin conductive layers is highly desirable for future technological applications.

cond-mat.mtrl-sci

Multi-Domain Adversarial Feature Generalization for Person Re-Identification

With the assistance of sophisticated training methods applied to single labeled datasets, the performance of fully-supervised person re-identification (Person Re-ID) has been improved significantly in recent years. However, these models trained on a single dataset usually suffer from considerable performance degradation when applied to videos of a different camera network. To make Person Re-ID systems more practical and scalable, several cross-dataset domain adaptation methods have been proposed, which achieve high performance without the labeled data from the target domain. However, these approaches still require the unlabeled data of the target domain during the training process, making them impractical. A practical Person Re-ID system pre-trained on other datasets should start running immediately after deployment on a new site without having to wait until sufficient images or videos are collected and the pre-trained model is tuned. To serve this purpose, in this paper, we reformulate person re-identification as a multi-dataset domain generalization problem. We propose a multi-dataset feature generalization network (MMFA-AAE), which is capable of learning a universal domain-invariant feature representation from multiple labeled datasets and generalizing it to `unseen' camera systems. The network is based on an adversarial auto-encoder to learn a generalized domain-invariant latent feature representation with the Maximum Mean Discrepancy (MMD) measure to align the distributions across multiple domains. Extensive experiments demonstrate the effectiveness of the proposed method. Our MMFA-AAE approach not only outperforms most of the domain generalization Person Re-ID methods, but also surpasses many state-of-the-art supervised methods and unsupervised domain adaptation methods by a large margin.

cs.CV

Strong Ferromagnetism Achieved via Breathing Lattices in Atomically Thin Cobaltites

Low-dimensional quantum materials that remain strongly ferromagnetic down to mono layer thickness are highly desired for spintronic applications. Although oxide materials are important candidates for next generation of spintronic, ferromagnetism decays severely when the thickness is scaled to the nano meter regime, leading to deterioration of device performance. Here we report a methodology for maintaining strong ferromagnetism in insulating LaCoO3 (LCO) layers down to the thickness of a single unit cell. We find that the magnetic and electronic states of LCO are linked intimately to the structural parameters of adjacent "breathing lattice" SrCuO2 (SCO). As the dimensionality of SCO is reduced, the lattice constant elongates over 10% along the growth direction, leading to a significant distortion of the CoO6 octahedra, and promoting a higher spin state and long-range spin ordering. For atomically thin LCO layers, we observe surprisingly large magnetic moment (0.5 uB/Co) and Curie temperature (75 K), values larger than previously reported for any mono layer oxide. Our results demonstrate a strategy for creating ultra thin ferromagnetic oxides by exploiting atomic hetero interface engineering,confinement-driven structural transformation, and spin-lattice entanglement in strongly correlated materials.

cond-mat.mtrl-sci

Strain-mediated high conductivity in ultrathin antiferromagnetic metallic nitrides

Strain engineering provides the ability to control the ground states and associated phase transition in the epitaxial films. However, the systematic study of intrinsic characters and their strain dependency in transition-metal nitrides remains challenging due to the difficulty in fabricating the stoichiometric and high-quality films. Here we report the observation of electronic state transition in highly crystalline antiferromagnetic CrN films with strain and reduced dimensionality. Shrinking the film thickness to a critical value of ~ 30 unit cells, a profound conductivity reduction accompanied by unexpected volume expansion is observed in CrN films. The electrical conductivity is observed surprisingly when the CrN layer as thin as single unit cell thick, which is far below the critical thickness of most metallic films. We found that the metallicity of an ultrathin CrN film recovers from an insulating behavior upon the removal of as-grown strain by fabrication of first-ever freestanding nitride films. Both first-principles calculations and linear dichroism measurements reveal that the strain-mediated orbital splitting effectively customizes the relatively small bandgap at the Fermi level, leading to exotic phase transition in CrN. The ability to achieve highly conductive nitride ultrathin films by harness strain-controlling over competing phases can be used for utilizing their exceptional characteristics.

cond-mat.mtrl-sci

LC-GAN: Image-to-image Translation Based on Generative Adversarial Network for Endoscopic Images

Intelligent vision is appealing in computer-assisted and robotic surgeries. Vision-based analysis with deep learning usually requires large labeled datasets, but manual data labeling is expensive and time-consuming in medical problems. We investigate a novel cross-domain strategy to reduce the need for manual data labeling by proposing an image-to-image translation model live-cadaver GAN (LC-GAN) based on generative adversarial networks (GANs). We consider a situation when a labeled cadaveric surgery dataset is available while the task is instrument segmentation on an unlabeled live surgery dataset. We train LC-GAN to learn the mappings between the cadaveric and live images. For live image segmentation, we first translate the live images to fake-cadaveric images with LC-GAN and then perform segmentation on the fake-cadaveric images with models trained on the real cadaveric dataset. The proposed method fully makes use of the labeled cadaveric dataset for live image segmentation without the need to label the live dataset. LC-GAN has two generators with different architectures that leverage the deep feature representation learned from the cadaveric image based segmentation task. Moreover, we propose the structural similarity loss and segmentation consistency loss to improve the semantic consistency during translation. Our model achieves better image-to-image translation and leads to improved segmentation performance in the proposed cross-domain segmentation task.

eess.IV

Towards Better Surgical Instrument Segmentation in Endoscopic Vision: Multi-Angle Feature Aggregation and Contour Supervision

Accurate and real-time surgical instrument segmentation is important in the endoscopic vision of robot-assisted surgery, and significant challenges are posed by frequent instrument-tissue contacts and continuous change of observation perspective. For these challenging tasks more and more deep neural networks (DNN) models are designed in recent years. We are motivated to propose a general embeddable approach to improve these current DNN segmentation models without increasing the model parameter number. Firstly, observing the limited rotation-invariance performance of DNN, we proposed the Multi-Angle Feature Aggregation (MAFA) method, leveraging active image rotation to gain richer visual cues and make the prediction more robust to instrument orientation changes. Secondly, in the end-to-end training stage, the auxiliary contour supervision is utilized to guide the model to learn the boundary awareness, so that the contour shape of segmentation mask is more precise. The proposed method is validated with ablation experiments on the novel Sinus-Surgery datasets collected from surgeons' operations, and is compared to the existing methods on a public dataset collected with a da Vinci Xi Robot.

cs.CV

MPC-guided Imitation Learning of Neural Network Policies for the Artificial Pancreas

Even though model predictive control (MPC) is currently the main algorithm for insulin control in the artificial pancreas (AP), it usually requires complex online optimizations, which are infeasible for resource-constrained medical devices. MPC also typically relies on state estimation, an error-prone process. In this paper, we introduce a novel approach to AP control that uses Imitation Learning to synthesize neural-network insulin policies from MPC-computed demonstrations. Such policies are computationally efficient and, by instrumenting MPC at training time with full state information, they can directly map measurements into optimal therapy decisions, thus bypassing state estimation. We apply Bayesian inference via Monte Carlo Dropout to learn policies, which allows us to quantify prediction uncertainty and thereby derive safer therapy decisions. We show that our control policies trained under a specific patient model readily generalize (in terms of model parameters and disturbance distributions) to patient cohorts, consistently outperforming traditional MPC with state estimation.

cs.LG

Generalization Bounds for Convolutional Neural Networks

Convolutional neural networks (CNNs) have achieved breakthrough performances in a wide range of applications including image classification, semantic segmentation, and object detection. Previous research on characterizing the generalization ability of neural networks mostly focuses on fully connected neural networks (FNNs), regarding CNNs as a special case of FNNs without taking into account the special structure of convolutional layers. In this work, we propose a tighter generalization bound for CNNs by exploiting the sparse and permutation structure of its weight matrices. As the generalization bound relies on the spectral norm of weight matrices, we further study spectral norms of three commonly used convolution operations including standard convolution, depthwise convolution, and pointwise convolution. Theoretical and experimental results both demonstrate that our bounds for CNNs are tighter than existing bounds.

stat.ML

Automated Synthesis of Safe Digital Controllers for Sampled-Data Stochastic Nonlinear Systems

We present a new method for the automated synthesis of digital controllers with formal safety guarantees for systems with nonlinear dynamics, noisy output measurements, and stochastic disturbances. Our method derives digital controllers such that the corresponding closed-loop system, modeled as a sampled-data stochastic control system, satisfies a safety specification with probability above a given threshold. The proposed synthesis method alternates between two steps: generation of a candidate controller pc, and verification of the candidate. pc is found by maximizing a Monte Carlo estimate of the safety probability, and by using a non-validated ODE solver for simulating the system. Such a candidate is therefore sub-optimal but can be generated very rapidly. To rule out unstable candidate controllers, we prove and utilize Lyapunov's indirect method for instability of sampled-data nonlinear systems. In the subsequent verification step, we use a validated solver based on SMT (Satisfiability Modulo Theories) to compute a numerically and statistically valid confidence interval for the safety probability of pc. If the probability so obtained is not above the threshold, we expand the search space for candidates by increasing the controller degree. We evaluate our technique on three case studies: an artificial pancreas model, a powertrain control model, and a quadruple-tank process.

eess.SY

Homogeneous Feature Transfer and Heterogeneous Location Fine-tuning for Cross-City Property Appraisal Framework

Most existing real estate appraisal methods focus on building accuracy and reliable models from a given dataset but pay little attention to the extensibility of their trained model. As different cities usually contain a different set of location features (district names, apartment names), most existing mass appraisal methods have to train a new model from scratch for different cities or regions. As a result, these approaches require massive data collection for each city and the total training time for a multi-city property appraisal system will be extremely long. Besides, some small cities may not have enough data for training a robust appraisal model. To overcome these limitations, we develop a novel Homogeneous Feature Transfer and Heterogeneous Location Fine-tuning (HFT+HLF) cross-city property appraisal framework. By transferring partial neural network learning from a source city and fine-tuning on the small amount of location information of a target city, our semi-supervised model can achieve similar or even superior performance compared to a fully supervised Artificial neural network (ANN) method.

cs.LG

Synthesizing Stealthy Reprogramming Attacks on Cardiac Devices

An Implantable Cardioverter Defibrillator (ICD) is a medical device used for the detection of potentially fatal cardiac arrhythmia and their treatment through the delivery of electrical shocks intended to restore normal heart rhythm. An ICD reprogramming attack seeks to alter the device's parameters to induce unnecessary shocks and, even more egregious, prevent required therapy. In this paper, we present a formal approach for the synthesis of ICD reprogramming attacks that are both effective, i.e., lead to fundamental changes in the required therapy, and stealthy, i.e., involve minimal changes to the nominal ICD parameters. We focus on the discrimination algorithm underlying Boston Scientific devices (one of the principal ICD manufacturers) and formulate the synthesis problem as one of multi-objective optimization. Our solution technique is based on an Optimization Modulo Theories encoding of the problem and allows us to derive device parameters that are optimal with respect to the effectiveness-stealthiness tradeoff (i.e., lie along the corresponding Pareto front). To the best of our knowledge, our work is the first to derive systematic ICD reprogramming attacks designed to maximize therapy disruption while minimizing detection. To evaluate our technique, we employ an extensive dataset of synthetic EGMs (cardiac signals), each generated with a prescribed arrhythmia, allowing us to synthesize attacks tailored to the victim's cardiac condition. Our approach readily generalizes to unseen signals, representing the unknown EGM of the victim patient.

eess.SY

Multi-task Mid-level Feature Alignment Network for Unsupervised Cross-Dataset Person Re-Identification

Most existing person re-identification (Re-ID) approaches follow a supervised learning framework, in which a large number of labelled matching pairs are required for training. Such a setting severely limits their scalability in real-world applications where no labelled samples are available during the training phase. To overcome this limitation, we develop a novel unsupervised Multi-task Mid-level Feature Alignment (MMFA) network for the unsupervised cross-dataset person re-identification task. Under the assumption that the source and target datasets share the same set of mid-level semantic attributes, our proposed model can be jointly optimised under the person's identity classification and the attribute learning task with a cross-dataset mid-level feature alignment regularisation term. In this way, the learned feature representation can be better generalised from one dataset to another which further improve the person re-identification accuracy. Experimental results on four benchmark datasets demonstrate that our proposed method outperforms the state-of-the-art baselines.

cs.CV

Declarative vs Rule-based Control for Flocking Dynamics

The popularity of rule-based flocking models, such as Reynolds' classic flocking model, raises the question of whether more declarative flocking models are possible. This question is motivated by the observation that declarative models are generally simpler and easier to design, understand, and analyze than operational models. We introduce a very simple control law for flocking based on a cost function capturing cohesion (agents want to stay together) and separation (agents do not want to get too close). We refer to it as {\textit declarative flocking} (DF). We use model-predictive control (MPC) to define controllers for DF in centralized and distributed settings. A thorough performance comparison of our declarative flocking with Reynolds' model, and with more recent flocking models that use MPC with a cost function based on lattice structures, demonstrate that DF-MPC yields the best cohesion and least fragmentation, and maintains a surprisingly good level of geometric regularity while still producing natural flock shapes similar to those produced by Reynolds' model. We also show that DF-MPC has high resilience to sensor noise.

cs.MA

Data-Driven Robust Taxi Dispatch under Demand Uncertainties

In modern taxi networks, large amounts of taxi occupancy status and location data are collected from networked in-vehicle sensors in real-time. They provide knowledge of system models on passenger demand and mobility patterns for efficient taxi dispatch and coordination strategies. Such approaches face new challenges: how to deal with uncertainties of predicted customer demand while fulfilling the system's performance requirements, including minimizing taxis' total idle mileage and maintaining service fairness across the whole city; how to formulate a computationally tractable problem. To address this problem, we develop a data-driven robust taxi dispatch framework to consider spatial-temporally correlated demand uncertainties. The robust vehicle dispatch problem we formulate is concave in the uncertain demand and convex in the decision variables. Uncertainty sets of random demand vectors are constructed from data based on theories in hypothesis testing, and provide a desired probabilistic guarantee level for the performance of robust taxi dispatch solutions. We prove equivalent computationally tractable forms of the robust dispatch problem using the minimax theorem and strong duality. Evaluations on four years of taxi trip data for New York City show that by selecting a probabilistic guarantee level at 75%, the average demand-supply ratio error is reduced by 31.7%, and the average total idle driving distance is reduced by 10.13% or about 20 million miles annually, compared with non-robust dispatch solutions.

eess.SY