Searcharxiv⌕ Search

arXiv subjects

Han Shen

Publications and source records attributed to Han Shen.

32 records · Page 2Linked to original sources

Mapping the Galactic disk with the LAMOST and Gaia Red clump sample: VIII: Mapping the kinematics of the Galactic disk using mono-age and mono-abundance stellar populations

We present a comprehensive study of the kinematic properties of the different Galactic disk populations, as defined by the chemical abundance ratios and stellar ages, across a large disk volume (4.5 $\leq$ R $\leq$ 15.0 kpc and $|Z|$ $\leq$ 3.0 kpc), by using the LAMOST-Gaia red clump sample stars. We determine the median velocities for various spatial and population bins, finding large-scale bulk motions, such as the wave-like behavior in radial velocity, the north-south discrepancy in azimuthal velocity and the warp signal in vertical velocity, and the amplitudes and spatial-dependences of those bulk motions show significant variations for different mono-age and mono-abundance populations. The global spatial behaviors of the velocity dispersions clearly show a signal of spiral arms and, a signal of the disk perturbation event within 4 Gyr, as well as the disk flaring in the outer region (i.e., $R \ge 12$ kpc) mostly for young or alpha-poor stellar populations. Our detailed measurements of age/[$α$/Fe]-velocity dispersion relations for different disk volumes indicate that young/$α$-poor populations are likely originated from dynamically heated by both giant molecular clouds and spiral arms, while old/$α$-enhanced populations require an obvious contribution from other heating mechanisms such as merger and accretion, or born in the chaotic mergers of gas-rich systems and/or turbulent interstellar medium.

astro-ph.GA↗

Joint Unsupervised and Supervised Training for Automatic Speech Recognition via Bilevel Optimization

In this paper, we present a novel bilevel optimization-based training approach to training acoustic models for automatic speech recognition (ASR) tasks that we term {bi-level joint unsupervised and supervised training (BL-JUST)}. {BL-JUST employs a lower and upper level optimization with an unsupervised loss and a supervised loss respectively, leveraging recent advances in penalty-based bilevel optimization to solve this challenging ASR problem with affordable complexity and rigorous convergence guarantees.} To evaluate BL-JUST, extensive experiments on the LibriSpeech and TED-LIUM v2 datasets have been conducted. BL-JUST achieves superior performance over the commonly used pre-training followed by fine-tuning strategy.

cs.CL↗

The tilt of the velocity ellipsoid of different Galactic disk populations

The tilt of the velocity ellipsoid is a helpful tracer of the gravitational potential of the Milky Way. In this paper, we use nearly 140,000 RC stars selected from the LAMOST and Gaia to make a detailed analysis of the tilt of the velocity ellipsoid for various populations, as defined by the stellar ages and chemical information, within 4.5 $\leq$ $R$ $\leq$ 15.0 kpc and $|Z|$ $\leq$ 3.0 kpc. The tilt angles of the velocity ellipsoids of the RC sample stars are accurately described as $α$ = $α_{0}$ $\mathrm{arctan}$ ($Z$/$R$) with $α_{0}$ = (0.68 $\pm$ 0.05). This indicates the alignment of velocity ellipsoids is between cylindrical and spherical, implying that any deviation from the spherical alignment of the velocity ellipsoids may be caused by the gravitational potential of the baryonic disk. The results of various populations suggest that the $α_{0}$ displays an age and population dependence, with the thin and thick disks respectively values $α_{0}$ = (0.72 $\pm$ 0.08) and $α_{0}$ = (0.64 $\pm$ 0.07), and the $α_{0}$ displays a decreasing trend with age (and [$α$/Fe]) increases, meaning that the velocity ellipsoids of the kinematically relaxed stars are mainly dominated by the gravitational potential of the baryonic disk. We determine the $α_{0} - R$ for various populations, finding that the $α_{0}$ displays oscillations with $R$ for all the different populations. The oscillations in $α_{0}$ appear in both kinematically hot and cold populations, indicating that resonances with the Galactic bar are the most likely origin for these oscillations.

astro-ph.GA↗

Alternating Implicit Projected SGD and Its Efficient Variants for Equality-constrained Bilevel Optimization

Stochastic bilevel optimization, which captures the inherent nested structure of machine learning problems, is gaining popularity in many recent applications. Existing works on bilevel optimization mostly consider either unconstrained problems or constrained upper-level problems. This paper considers the stochastic bilevel optimization problems with equality constraints both in the upper and lower levels. By leveraging the special structure of the equality constraints problem, the paper first presents an alternating implicit projected SGD approach and establishes the $\tilde{\cal O}(ε^{-2})$ sample complexity that matches the state-of-the-art complexity of ALSET \citep{chen2021closing} for unconstrained bilevel problems. To further save the cost of projection, the paper presents two alternating implicit projection-efficient SGD approaches, where one algorithm enjoys the $\tilde{\cal O}(ε^{-2}/T)$ upper-level and $\tilde{\cal O}(ε^{-1.5}/T^{\frac{3}{4}})$ lower-level projection complexity with ${\cal O}(T)$ lower-level batch size, and the other one enjoys $\tilde{\cal O}(ε^{-1.5})$ upper-level and lower-level projection complexity with ${\cal O}(1)$ batch size. Application to federated bilevel optimization has been presented to showcase the empirical performance of our algorithms. Our results demonstrate that equality-constrained bilevel optimization with strongly-convex lower-level problems can be solved as efficiently as stochastic single-level optimization problems.

cs.LG↗

A Single-Timescale Analysis For Stochastic Approximation With Multiple Coupled Sequences

Stochastic approximation (SA) with multiple coupled sequences has found broad applications in machine learning such as bilevel learning and reinforcement learning (RL). In this paper, we study the finite-time convergence of nonlinear SA with multiple coupled sequences. Different from existing multi-timescale analysis, we seek for scenarios where a fine-grained analysis can provide the tight performance guarantee for multi-sequence single-timescale SA (STSA). At the heart of our analysis is the smoothness property of the fixed points in multi-sequence SA that holds in many applications. When all sequences have strongly monotone increments, we establish the iteration complexity of $\mathcal{O}(ε^{-1})$ to achieve $ε$-accuracy, which improves the existing $\mathcal{O}(ε^{-1.5})$ complexity for two coupled sequences. When all but the main sequence have strongly monotone increments, we establish the iteration complexity of $\mathcal{O}(ε^{-2})$. The merit of our results lies in that applying them to stochastic bilevel and compositional optimization problems, as well as RL problems leads to either relaxed assumptions or improvements over their existing performance guarantees.

cs.LG↗

Estimating accurate reddening values of LAMOST M dwarfs

M dwarfs are the dominating type of stars in the solar neighbourhood. They serve as excellent tracers for the study of the distribution and properties of the nearby interstellar dust. In this work, we aim to obtain high accuracy reddening values of M dwarf stars from the Large Sky Area Multi-Object Fiber Spectroscopic Telescope (LAMOST) Data Release 8 (DR8). Combining the LAMOST spectra with the high-quality optical photometry from the Gaia Early Data Release 3 (Gaia EDR3), we have estimated the reddening values $E(G_{\rm BP}-G_{\rm RP})$ of 641,426 M dwarfs with the machine-learning algorithm Random Forest regression. The typical reddening uncertainty is only 0.03 mag in $E(G_{\rm BP}-G_{\rm RP})$. We have obtained the reddening coefficient $R_{(G_{\rm BP}-G_{\rm RP})}$, which is a function of the stellar intrinsic colour $(G_{\rm BP}-G_{\rm RP})_0$ and reddening value $E(B-V)$. The values of $E(B-V)$ are also provided for the individual stars in our catalogue. Our resultant high accuracy reddening values of M dwarfs, combined with the Gaia parallaxes, will be very powerful to map the fine structures of the dust in the solar neighbourhood.

astro-ph.SR↗

Towards Understanding Asynchronous Advantage Actor-critic: Convergence and Linear Speedup

Asynchronous and parallel implementation of standard reinforcement learning (RL) algorithms is a key enabler of the tremendous success of modern RL. Among many asynchronous RL algorithms, arguably the most popular and effective one is the asynchronous advantage actor-critic (A3C) algorithm. Although A3C is becoming the workhorse of RL, its theoretical properties are still not well-understood, including its non-asymptotic analysis and the performance gain of parallelism (a.k.a. linear speedup). This paper revisits the A3C algorithm and establishes its non-asymptotic convergence guarantees. Under both i.i.d. and Markovian sampling, we establish the local convergence guarantee for A3C in the general policy approximation case and the global convergence guarantee in softmax policy parameterization. Under i.i.d. sampling, A3C obtains sample complexity of $\mathcal{O}(ε^{-2.5}/N)$ per worker to achieve $ε$ accuracy, where $N$ is the number of workers. Compared to the best-known sample complexity of $\mathcal{O}(ε^{-2.5})$ for two-timescale AC, A3C achieves \emph{linear speedup}, which justifies the advantage of parallelism and asynchrony in AC algorithms theoretically for the first time. Numerical tests on synthetic environment, OpenAI Gym environments and Atari games have been provided to verify our theoretical analysis.

cs.LG↗

Adaptive Temporal Difference Learning with Linear Function Approximation

This paper revisits the temporal difference (TD) learning algorithm for the policy evaluation tasks in reinforcement learning. Typically, the performance of TD(0) and TD($λ$) is very sensitive to the choice of stepsizes. Oftentimes, TD(0) suffers from slow convergence. Motivated by the tight link between the TD(0) learning algorithm and the stochastic gradient methods, we develop a provably convergent adaptive projected variant of the TD(0) learning algorithm with linear function approximation that we term AdaTD(0). In contrast to the TD(0), AdaTD(0) is robust or less sensitive to the choice of stepsizes. Analytically, we establish that to reach an $ε$ accuracy, the number of iterations needed is $\tilde{O}(ε^{-2}\ln^4\frac{1}ε/\ln^4\frac{1}ρ)$ in the general case, where $ρ$ represents the speed of the underlying Markov chain converges to the stationary distribution. This implies that the iteration complexity of AdaTD(0) is no worse than that of TD(0) in the worst case. When the stochastic semi-gradients are sparse, we provide theoretical acceleration of AdaTD(0). Going beyond TD(0), we develop an adaptive variant of TD($λ$), which is referred to as AdaTD($λ$). Empirically, we evaluate the performance of AdaTD(0) and AdaTD($λ$) on several standard reinforcement learning tasks, which demonstrate the effectiveness of our new approaches.

math.OC↗

Byzantine-Resilient Decentralized TD Learning with Linear Function Approximation

This paper considers the policy evaluation problem in a multi-agent reinforcement learning (MARL) environment over decentralized and directed networks. The focus is on decentralized temporal difference (TD) learning with linear function approximation in the presence of unreliable or even malicious agents, termed as Byzantine agents. In order to evaluate the quality of a fixed policy in a common environment, agents usually run decentralized TD($λ$) collaboratively. However, when some Byzantine agents behave adversarially, decentralized TD($λ$) is unable to learn an accurate linear approximation for the true value function. We propose a trimmed-mean based Byzantine-resilient decentralized TD($λ$) algorithm to perform policy evaluation in this setting. We establish the finite-time convergence rate, as well as the asymptotic learning error in the presence of Byzantine agents. Numerical experiments corroborate the robustness of the proposed algorithm.

math.OC↗

Multi-object Tracking via End-to-end Tracklet Searching and Ranking

Recent works in multiple object tracking use sequence model to calculate the similarity score between the detections and the previous tracklets. However, the forced exposure to ground-truth in the training stage leads to the training-inference discrepancy problem, i.e., exposure bias, where association error could accumulate in the inference and make the trajectories drift. In this paper, we propose a novel method for optimizing tracklet consistency, which directly takes the prediction errors into account by introducing an online, end-to-end tracklet search training process. Notably, our methods directly optimize the whole tracklet score instead of pairwise affinity. With sequence model as appearance encoders of tracklet, our tracker achieves remarkable performance gain from conventional tracklet association baseline. Our methods have also achieved state-of-the-art in MOT15~17 challenge benchmarks using public detection and online settings.

cs.CV↗

Real Time Visual Tracking using Spatial-Aware Temporal Aggregation Network

More powerful feature representations derived from deep neural networks benefit visual tracking algorithms widely. However, the lack of exploitation on temporal information prevents tracking algorithms from adapting to appearances changing or resisting to drift. This paper proposes a correlation filter based tracking method which aggregates historical features in a spatial-aligned and scale-aware paradigm. The features of historical frames are sampled and aggregated to search frame according to a pixel-level alignment module based on deformable convolutions. In addition, we also use a feature pyramid structure to handle motion estimation at different scales, and address the different demands on feature granularity between tracking losses and deformation offset learning. By this design, the tracker, named as Spatial-Aware Temporal Aggregation network (SATA), is able to assemble appearances and motion contexts of various scales in a time period, resulting in better performance compared to a single static image. Our tracker achieves leading performance in OTB2013, OTB2015, VOT2015, VOT2016 and LaSOT, and operates at a real-time speed of 26 FPS, which indicates our method is effective and practical. Our code will be made publicly available at \href{https://github.com/ecart18/SATA}{https://github.com/ecart18/SATA}.

cs.CV↗

Object Detection in Video with Spatial-temporal Context Aggregation

Recent cutting-edge feature aggregation paradigms for video object detection rely on inferring feature correspondence. The feature correspondence estimation problem is fundamentally difficult due to poor image quality, motion blur, etc, and the results of feature correspondence estimation are unstable. To avoid the problem, we propose a simple but effective feature aggregation framework which operates on the object proposal-level. It learns to enhance each proposal's feature via modeling semantic and spatio-temporal relationships among object proposals from both within a frame and across adjacent frames. Experiments are carried out on the ImageNet VID dataset. Without any bells and whistles, our method obtains 80.3\% mAP on the ImageNet VID dataset, which is superior over the previous state-of-the-arts. The proposed feature aggregation mechanism improves the single frame Faster RCNN baseline by 5.8% mAP. Besides, under the setting of no temporal post-processing, our method outperforms the previous state-of-the-art by 1.4% mAP.

cs.CV↗

Proposal, Tracking and Segmentation (PTS): A Cascaded Network for Video Object Segmentation

Video object segmentation (VOS) aims at pixel-level object tracking given only the annotations in the first frame. Due to the large visual variations of objects in video and the lack of training samples, it remains a difficult task despite the upsurging development of deep learning. Toward solving the VOS problem, we bring in several new insights by the proposed unified framework consisting of object proposal, tracking and segmentation components. The object proposal network transfers objectness information as generic knowledge into VOS; the tracking network identifies the target object from the proposals; and the segmentation network is performed based on the tracking results with a novel dynamic-reference based model adaptation scheme. Extensive experiments have been conducted on the DAVIS'17 dataset and the YouTube-VOS dataset, our method achieves the state-of-the-art performance on several video object segmentation benchmarks. We make the code publicly available at https://github.com/sydney0zq/PTSNet.

cs.CV↗

Tracklet Association Tracker: An End-to-End Learning-based Association Approach for Multi-Object Tracking

Traditional multiple object tracking methods divide the task into two parts: affinity learning and data association. The separation of the task requires to define a hand-crafted training goal in affinity learning stage and a hand-crafted cost function of data association stage, which prevents the tracking goals from learning directly from the feature. In this paper, we present a new multiple object tracking (MOT) framework with data-driven association method, named as Tracklet Association Tracker (TAT). The framework aims at gluing feature learning and data association into a unity by a bi-level optimization formulation so that the association results can be directly learned from features. To boost the performance, we also adopt the popular hierarchical association and perform the necessary alignment and selection of raw detection responses. Our model trains over 20X faster than a similar approach, and achieves the state-of-the-art performance on both MOT2016 and MOT2017 benchmarks.

cs.CV↗