Searcharxiv⌕ Search

arXiv subjects

Yuning Jiang

Publications and source records attributed to Yuning Jiang.

At least 127 records · Page 7Linked to original sources

Decentralized Optimization over Tree Graphs

This paper presents a decentralized algorithm for non-convex optimization over tree-structured networks. We assume that each node of this network can solve small-scale optimization problems and communicate approximate value functions with its neighbors based on a novel multi-sweep communication protocol. In contrast to existing parallelizable optimization algorithms for non-convex optimization the nodes of the network are neither synchronized nor assign any central entity. None of the nodes needs to know the whole topology of the network, but all nodes know that the network is tree-structured. We discuss conditions under which locally quadratic convergence rates can be achieved. The method is illustrated by running the decentralized asynchronous multi-sweep protocol on a radial AC power network case study.

math.OC↗

SOLO: Segmenting Objects by Locations

We present a new, embarrassingly simple approach to instance segmentation in images. Compared to many other dense prediction tasks, e.g., semantic segmentation, it is the arbitrary number of instances that have made instance segmentation much more challenging. In order to predict a mask for each instance, mainstream approaches either follow the 'detect-thensegment' strategy as used by Mask R-CNN, or predict category masks first then use clustering techniques to group pixels into individual instances. We view the task of instance segmentation from a completely new perspective by introducing the notion of "instance categories", which assigns categories to each pixel within an instance according to the instance's location and size, thus nicely converting instance mask segmentation into a classification-solvable problem. Now instance segmentation is decomposed into two classification tasks. We demonstrate a much simpler and flexible instance segmentation framework with strong performance, achieving on par accuracy with Mask R-CNN and outperforming recent singleshot instance segmenters in accuracy. We hope that this very simple and strong framework can serve as a baseline for many instance-level recognition tasks besides instance segmentation.

cs.CV↗

Controllable Person Image Synthesis with Attribute-Decomposed GAN

This paper introduces the Attribute-Decomposed GAN, a novel generative model for controllable person image synthesis, which can produce realistic person images with desired human attributes (e.g., pose, head, upper clothes and pants) provided in various source inputs. The core idea of the proposed model is to embed human attributes into the latent space as independent codes and thus achieve flexible and continuous control of attributes via mixing and interpolation operations in explicit style representations. Specifically, a new architecture consisting of two encoding pathways with style block connections is proposed to decompose the original hard mapping into multiple more accessible subtasks. In source pathway, we further extract component layouts with an off-the-shelf human parser and feed them into a shared global texture encoder for decomposed latent codes. This strategy allows for the synthesis of more realistic output images and automatic separation of un-annotated attributes. Experimental results demonstrate the proposed method's superiority over the state of the art in pose transfer and its effectiveness in the brand-new task of component attribute transfer.

cs.CV↗

FoveaBox: Beyond Anchor-based Object Detector

We present FoveaBox, an accurate, flexible, and completely anchor-free framework for object detection. While almost all state-of-the-art object detectors utilize predefined anchors to enumerate possible locations, scales and aspect ratios for the search of the objects, their performance and generalization ability are also limited to the design of anchors. Instead, FoveaBox directly learns the object existing possibility and the bounding box coordinates without anchor reference. This is achieved by: (a) predicting category-sensitive semantic maps for the object existing possibility, and (b) producing category-agnostic bounding box for each position that potentially contains an object. The scales of target boxes are naturally associated with feature pyramid representations. In FoveaBox, an instance is assigned to adjacent feature levels to make the model more accurate.We demonstrate its effectiveness on standard benchmarks and report extensive experimental analysis. Without bells and whistles, FoveaBox achieves state-of-the-art single model performance on the standard COCO and Pascal VOC object detection benchmark. More importantly, FoveaBox avoids all computation and hyper-parameters related to anchor boxes, which are often sensitive to the final detection performance. We believe the simple and effective approach will serve as a solid baseline and help ease future research for object detection. The code has been made publicly available at https://github.com/taokong/FoveaBox .

cs.CV↗

Distributed Optimization for Massive Connectivity

Massive device connectivity in Internet of Thing (IoT) networks with sporadic traffic poses significant communication challenges. To overcome this challenge, the serving base station is required to detect the active devices and estimate the corresponding channel state information during each coherence block. The corresponding joint activity detection and channel estimation problem can be formulated as a group sparse estimation problem, also known under the name "Group Lasso". This letter presents a fast and efficient distributed algorithm to solve such Group Lasso problems, which alternates between solving small-scaled problems in parallel and dealing with a linear equation for consensus. Numerical results demonstrate the speedup of this algorithm compared with the state-of-the-art methods in terms of convergence speed and computation time.

eess.SP↗

Optimal Experiment Design for AC Power Systems Admittance Estimation

The integration of renewables into electrical grids calls for the development of tailored control schemes which in turn require reliable grid models. In many cases, the grid topology is known but the actual parameters are not exactly known. This paper proposes a new approach for online parameter estimation in power systems based on optimal experimental design using multiple measurement snapshots. In contrast to conventional methods, our method computes optimal excitations extracting the maximum information in each estimation step to accelerate convergence. The performance of the proposed method is illustrated on a case study.

eess.SY↗

Distributed Optimization using ALADIN for MPC in Smart Grids

This paper presents a distributed optimization algorithm tailored to solve optimization problems arising in smart grids. In detail, we propose a variant of the Augmented Lagrangian based Alternating Direction Inexact Newton (ALADIN) method, which comes along with global convergence guarantees for the considered class of linear-quadratic optimization problems. We establish local quadratic convergence of the proposed scheme and elaborate its advantages compared to the Alternating Direction Method of Multipliers (ADMM). In particular, we show that, at the cost of more communication, ALADIN requires fewer iterations to achieve the desired accuracy. Furthermore, it is numerically demonstrated that the number of iterations is independent of the number of subsystems. The effectiveness of the proposed scheme is illustrated by running both an ALADIN and an ADMM based model predictive controller on a benchmark case study.

eess.SY↗

DeepDualMapper: A Gated Fusion Network for Automatic Map Extraction using Aerial Images and Trajectories

Automatic map extraction is of great importance to urban computing and location-based services. Aerial image and GPS trajectory data refer to two different data sources that could be leveraged to generate the map, although they carry different types of information. Most previous works on data fusion between aerial images and data from auxiliary sensors do not fully utilize the information of both modalities and hence suffer from the issue of information loss. We propose a deep convolutional neural network called DeepDualMapper which fuses the aerial image and trajectory data in a more seamless manner to extract the digital map. We design a gated fusion module to explicitly control the information flows from both modalities in a complementary-aware manner. Moreover, we propose a novel densely supervised refinement decoder to generate the prediction in a coarse-to-fine way. Our comprehensive experiments demonstrate that DeepDualMapper can fuse the information of images and trajectories much more effectively than existing approaches, and is able to generate maps with higher accuracy.

cs.CV↗

Distributed Control Enforcing Group Sparsity in Smart Grids

In modern smart grids, charging of local energy storage devices is coordinated on a residential level to compensate the volatile aggregated power demand on the time interval of interest. However, this results in a perpetual usage of all batteries which reduces their lifetime. We enforce group sparsity by using an $\ell_{p,q}$-regularization on the control to counteract this phenomenon. This leads to a non-smooth convex optimization problem, for which we propose a tailored Alternating Direction Method of Multipliers algorithm. We elaborate further how to embed it in a Model Predictive Control framework. We show that the proposed scheme yields sparse control while achieving reasonable overall peak shaving by numerical simulations.

eess.SY↗

Task-Aware Monocular Depth Estimation for 3D Object Detection

Monocular depth estimation enables 3D perception from a single 2D image, thus attracting much research attention for years. Almost all methods treat foreground and background regions ("things and stuff") in an image equally. However, not all pixels are equal. Depth of foreground objects plays a crucial role in 3D object recognition and localization. To date how to boost the depth prediction accuracy of foreground objects is rarely discussed. In this paper, we first analyse the data distributions and interaction of foreground and background, then propose the foreground-background separated monocular depth estimation (ForeSeE) method, to estimate the foreground depth and background depth using separate optimization objectives and depth decoders. Our method significantly improves the depth estimation performance on foreground objects. Applying ForeSeE to 3D object detection, we achieve 7.5 AP gains and set new state-of-the-art results among other monocular methods. Code will be available at: https://github.com/WXinlong/ForeSeE.

cs.CV↗

Parallel Explicit Tube Model Predictive Control

This paper is about a parallel algorithm for tube-based model predictive control. The proposed control algorithm solves robust model predictive control problems suboptimally, while exploiting their structure. This is achieved by implementing a real-time algorithm that iterates between the evaluation of piecewise affine functions, corresponding to the parametric solution of small-scale robust MPC problems, and the online solution of structured equality constrained QPs. The performance of the associated real-time robust MPC controllers is illustrated by a numerical case study.

math.OC↗

Effective Domain Knowledge Transfer with Soft Fine-tuning

Convolutional neural networks require numerous data for training. Considering the difficulties in data collection and labeling in some specific tasks, existing approaches generally use models pre-trained on a large source domain (e.g. ImageNet), and then fine-tune them on these tasks. However, the datasets from source domain are simply discarded in the fine-tuning process. We argue that the source datasets could be better utilized and benefit fine-tuning. This paper firstly introduces the concept of general discrimination to describe ability of a network to distinguish untrained patterns, and then experimentally demonstrates that general discrimination could potentially enhance the total discrimination ability on target domain. Furthermore, we propose a novel and light-weighted method, namely soft fine-tuning. Unlike traditional fine-tuning which directly replaces optimization objective by a loss function on the target domain, soft fine-tuning effectively keeps general discrimination by holding the previous loss and removes it softly. By doing so, soft fine-tuning improves the robustness of the network to data bias, and meanwhile accelerates the convergence. We evaluate our approach on several visual recognition tasks. Extensive experimental results support that soft fine-tuning provides consistent improvement on all evaluated tasks, and outperforms the state-of-the-art significantly. Codes will be made available to the public.

cs.CV↗

UniVSE: Robust Visual Semantic Embeddings via Structured Semantic Representations

We propose Unified Visual-Semantic Embeddings (UniVSE) for learning a joint space of visual and textual concepts. The space unifies the concepts at different levels, including objects, attributes, relations, and full scenes. A contrastive learning approach is proposed for the fine-grained alignment from only image-caption pairs. Moreover, we present an effective approach for enforcing the coverage of semantic components that appear in the sentence. We demonstrate the robustness of Unified VSE in defending text-domain adversarial attacks on cross-modal retrieval tasks. Such robustness also empowers the use of visual cues to resolve word dependencies in novel sentences.

cs.CV↗

Decomposition of non-convex optimization via bi-level distributed ALADIN

Decentralized optimization algorithms are important in different contexts, such as distributed optimal power flow or distributed model predictive control, as they avoid central coordination and enable decomposition of large-scale problems. In case of constrained non-convex optimization only a few algorithms are currently are available; often their performance is limited, or they lack convergence guarantees. This paper proposes a framework for decentralized non-convex optimization via bi-level distribution of the Augmented Lagrangian Alternating Direction Inexact Newton (ALADIN) algorithm. Bi-level distribution means that the outer ALADIN structure is combined with an inner distribution/decentralization level solving a condensed variant of ALADIN's convex coordination QP by decentralized algorithms. We prove sufficient conditions ensuring local convergence while allowing for inexact decentralized/distributed solutions of the coordination QP. Moreover, we show how a decentralized variant of conjugate gradient or decentralized ADMM schemes can be employed at the inner level. We draw upon case studies from power systems and robotics to illustrate the performance of the proposed framework.

math.OC↗

Distributed State Estimation for AC Power Systems using Gauss-Newton ALADIN

This paper proposes a structure exploiting algorithm for solving non-convex power system state estimation problems in distributed fashion. Because the power flow equations in large electrical grid networks are non-convex equality constraints, we develop a tailored state estimator based on Augmented Lagrangian Alternating Direction Inexact Newton (ALADIN) method, which can handle the nonlinearities efficiently. Here, our focus is on using Gauss-Newton Hessian approximations within ALADIN in order to arrive at at an efficient (computationally and communicationally) variant of ALADIN for network maximum likelihood estimation problems. Analyzing the IEEE 30-Bus system we illustrate how the proposed algorithm can be used to solve highly non-trivial network state estimation problems. We also compare the method with existing distributed parameter estimation codes in order to illustrate its performance.

eess.SY↗

Parallel Explicit Model Predictive Control

This paper is about a real-time model predictive control (MPC) algorithm for large-scale, structured linear systems with polytopic state and control constraints. The proposed controller receives the current state measurement as an input and computes a sub-optimal control reaction by evaluating a finite number of piecewise affine functions that correspond to the explicit solution maps of small-scale parametric quadratic programming (QP) problems. We provide recursive feasibility and asymptotic stability guarantees, which can both be verified offline. The feedback controller is suboptimal on purpose because we are enforcing real-time requirements assuming that it is impossible to solve the given large-scale QP in the given amount of time. In this context, a key contribution of this paper is that we provide a bound on the sub-optimality of the controller. Our numerical simulations illustrate that the proposed explicit real-time scheme easily scales up to systems with hundreds of states and long control horizons, system sizes that are completely out of the scope of existing, non-suboptimal Explicit MPC controllers.

math.OC↗

Consistent Optimization for Single-Shot Object Detection

We present consistent optimization for single stage object detection. Previous works of single stage object detectors usually rely on the regular, dense sampled anchors to generate hypothesis for the optimization of the model. Through an examination of the behavior of the detector, we observe that the misalignment between the optimization target and inference configurations has hindered the performance improvement. We propose to bride this gap by consistent optimization, which is an extension of the traditional single stage detector's optimization strategy. Consistent optimization focuses on matching the training hypotheses and the inference quality by utilizing of the refined anchors during training. To evaluate its effectiveness, we conduct various design choices based on the state-of-the-art RetinaNet detector. We demonstrate it is the consistent optimization, not the architecture design, that yields the performance boosts. Consistent optimization is nearly cost-free, and achieves stable performance gains independent of the model capacities or input scales. Specifically, utilizing consistent optimization improves RetinaNet from 39.1 AP to 40.1 AP on COCO dataset without any bells or whistles, which surpasses the accuracy of all existing state-of-the-art one-stage detectors when adopting ResNet-101 as backbone. The code will be made available.

cs.CV↗

Towards Distributed OPF using ALADIN

The present paper discusses the application of the recently proposed Augmented Lagrangian Alternating Direction Inexact Newton (ALADIN) method to non-convex AC Optimal Power Flow Problems (OPF) in a distributed fashion. In contrast to the often used Alternating Direction of Multipliers Method (ADMM), ALADIN guarantees locally quadratic convergence for AC OPF. Numerical results for 5 to 300 bus test cases indicate that ALADIN is able to outperform ADMM and to reduce the number of iterations by about one order of magnitude. We compare ALADIN to numerical results for ADMM documented in the literature. The improved convergence speed comes at the cost of increasing the communication effort per iteration. Therefore, we propose a variant of ALADIN that uses inexact Hessians to reduce communication. Additionally, we provide a detailed comparison of these ALADIN variants to ADMM from an algorithmic and communication perspective. Moreover, we prove that ALADIN converges locally at quadratic rate even for the relevant case of suboptimally solved local NLPs.

cs.DC↗