Searcharxiv⌕ Search

arXiv subjects

He Sun

Publications and source records attributed to He Sun.

At least 91 records · Page 5Linked to original sources

Recovering a Molecule's 3D Dynamics from Liquid-phase Electron Microscopy Movies

The dynamics of biomolecules are crucial for our understanding of their functioning in living systems. However, current 3D imaging techniques, such as cryogenic electron microscopy (cryo-EM), require freezing the sample, which limits the observation of their conformational changes in real time. The innovative liquid-phase electron microscopy (liquid-phase EM) technique allows molecules to be placed in the native liquid environment, providing a unique opportunity to observe their dynamics. In this paper, we propose TEMPOR, a Temporal Electron MicroscoPy Object Reconstruction algorithm for liquid-phase EM that leverages an implicit neural representation (INR) and a dynamical variational auto-encoder (DVAE) to recover time series of molecular structures. We demonstrate its advantages in recovering different motion dynamics from two simulated datasets, 7bcq and Cas9. To our knowledge, our work is the first attempt to directly recover 3D structures of a temporally-varying particle from liquid-phase EM movies. It provides a promising new approach for studying molecules' 3D dynamics in structural biology.

q-bio.QM↗

Nearly-Optimal Hierarchical Clustering for Well-Clustered Graphs

This paper presents two efficient hierarchical clustering (HC) algorithms with respect to Dasgupta's cost function. For any input graph $G$ with a clear cluster-structure, our designed algorithms run in nearly-linear time in the input size of $G$, and return an $O(1)$-approximate HC tree with respect to Dasgupta's cost function. We compare the performance of our algorithm against the previous state-of-the-art on synthetic and real-world datasets and show that our designed algorithm produces comparable or better HC trees with much lower running time.

cs.DS↗

Spectral Toolkit of Algorithms for Graphs: Technical Report (1)

Spectral Toolkit of Algorithms for Graphs (STAG) is an open-source library for efficient spectral graph algorithms, and its development starts in September 2022. We have so far finished the component on local graph clustering, and this technical report presents a user's guide to STAG, showcase studies, and several technical considerations behind our development.

cs.SI↗

Image Reconstruction without Explicit Priors

We consider solving ill-posed imaging inverse problems without access to an explicit image prior or ground-truth examples. An overarching challenge in inverse problems is that there are many undesired images that fit to the observed measurements, thus requiring image priors to constrain the space of possible solutions to more plausible reconstructions. However, in many applications it is difficult or potentially impossible to obtain ground-truth images to learn an image prior. Thus, inaccurate priors are often used, which inevitably result in biased solutions. Rather than solving an inverse problem using priors that encode the explicit structure of any one image, we propose to solve a set of inverse problems jointly by incorporating prior constraints on the collective structure of the underlying images.The key assumption of our work is that the ground-truth images we aim to reconstruct share common, low-dimensional structure. We show that such a set of inverse problems can be solved simultaneously by learning a shared image generator with a low-dimensional latent space. The parameters of the generator and latent embedding are learned by maximizing a proxy for the Evidence Lower Bound (ELBO). Once learned, the generator and latent embeddings can be combined to provide reconstructions for each inverse problem. The framework we propose can handle general forward model corruptions, and we show that measurements derived from only a few ground-truth images (O(10)) are sufficient for image reconstruction without explicit priors.

eess.IV↗

Decision-Aware Conditional GANs for Time Series Data

We introduce the decision-aware time-series conditional generative adversarial network (DAT-CGAN) as a method for time-series generation. The framework adopts a multi-Wasserstein loss on structured decision-related quantities, capturing the heterogeneity of decision-related data and providing new effectiveness in supporting the decision processes of end users. We improve sample efficiency through an overlapped block-sampling method, and provide a theoretical characterization of the generalization properties of DAT-CGAN. The framework is demonstrated on financial time series for a multi-time-step portfolio choice problem. We demonstrate better generative quality in regard to underlying data and different decision-related quantities than strong, GAN-based baselines.

cs.LG↗

Tracing the hot spot motion using the next generation Event Horizon Telescope (ngEHT)

We propose to trace the dynamical motion of a shearing hot spot near the SgrA* source through a dynamical image reconstruction algorithm, StarWarps. Such a hot spot may form as the exhaust of magnetic reconnection in a current sheet near the black hole horizon. A hot spot that is ejected from the current sheet into an orbit in the accretion disk may shear and diffuse due to instabilities at its boundary during its orbit, resulting in a distinct signature. We subdivide the motion to two distinct phases; the first phase refers to the appearance of the hot spot modelled as a bright blob, followed by a subsequent shearing phase simulated as a stretched ellipse. We employ different observational arrays, including EHT(2017,2022) and the next generation event horizon telescope (ngEHTp1, ngEHT) arrays, in which few new additional sites are added to the observational array. We make dynamical image reconstructions for each of these arrays. Subsequently, we infer the hot spot phase in the first phase followed by the axes ratio and the ellipse area in the second phase. We focus on the direct observability of the orbiting hot spot in the sub-mm wavelength. Our analysis demonstrates that newly added dishes may easily trace the first phase as well as part of the second phase, before the flux is reduced substantially. The algorithm used in this work can be extended to any other types of the dynamical motion. Consequently, we conclude that the ngEHT is a key to directly observe the dynamical motions near variable sources, such as SgrA*.

astro-ph.GA↗

Reinforcement Learning with Stepwise Fairness Constraints

AI methods are used in societally important settings, ranging from credit to employment to housing, and it is crucial to provide fairness in regard to algorithmic decision making. Moreover, many settings are dynamic, with populations responding to sequential decision policies. We introduce the study of reinforcement learning (RL) with stepwise fairness constraints, requiring group fairness at each time step. Our focus is on tabular episodic RL, and we provide learning algorithms with strong theoretical guarantees in regard to policy optimality and fairness violation. Our framework provides useful tools to study the impact of fairness constraints in sequential settings and brings up new challenges in RL.

cs.LG↗

Is the Algorithmic Kadison-Singer Problem Hard?

We study the following $\mathsf{KS}_2(c)$ problem: let $c \in\mathbb{R}^+$ be some constant, and $v_1,\ldots, v_m\in\mathbb{R}^d$ be vectors such that $\|v_i\|^2\leq α$ for any $i\in[m]$ and $\sum_{i=1}^m \langle v_i, x\rangle^2 =1$ for any $x\in\mathbb{R}^d$ with $\|x\|=1$. The $\mathsf{KS}_2(c)$ problem asks to find some $S\subset [m]$, such that it holds for all $x \in \mathbb{R}^d$ with $\|x\| = 1$ that \[ \left|\sum_{i \in S} \langle v_i, x\rangle^2 - \frac{1}{2}\right| \leq c\cdot\sqrtα,\] or report no if such $S$ doesn't exist. Based on the work of Marcus et al. and Weaver, the $\mathsf{KS}_2(c)$ problem can be seen as the algorithmic Kadison-Singer problem with parameter $c\in\mathbb{R}^+$. Our first result is a randomised algorithm with one-sided error for the $\mathsf{KS}_2(c)$ problem such that (1) our algorithm finds a valid set $S \subset [m]$ with probability at least $1-2/d$, if such $S$ exists, or (2) reports no with probability $1$, if no valid sets exist. The algorithm has running time \[ O\left(\binom{m}{n}\cdot \mathrm{poly}(m, d)\right)~\mbox{ for }~n = O\left(\frac{d}{ε^2} \log(d) \log\left(\frac{1}{c\sqrtα}\right)\right), \] where $ε$ is a parameter which controls the error of the algorithm. This presents the first algorithm for the Kadison-Singer problem whose running time is quasi-polynomial in $m$, although having exponential dependency on $d$. Moreover, it shows that the algorithmic Kadison-Singer problem is easier to solve in low dimensions. Our second result is on the computational complexity of the $\mathsf{KS}_2(c)$ problem. We show that the $\mathsf{KS}_2(1/(4\sqrt{2}))$ problem is $\mathsf{FNP}$-hard for general values of $d$, and solving the $\mathsf{KS}_2(1/(4\sqrt{2}))$ problem is as hard as solving the $\mathsf{NAE\mbox{-}3SAT}$ problem.

cs.CC↗

Differential Liquidity Provision in Uniswap v3 and Implications for Contract Design

Decentralized exchanges (DEXs) provide a means for users to trade pairs of assets on-chain without the need for a trusted third party to effectuate a trade. Amongst these, constant function market maker DEXs such as Uniswap handle the most volume of trades between ERC-20 tokens. With the introduction of Uniswap v3, liquidity providers can differentially allocate liquidity to trades that occur within specific price intervals. In this paper, we formalize the profit and loss that liquidity providers can earn when providing specific liquidity allocations to a v3 contract. We give a convex stochastic optimization problem for computing optimal liquidity allocation for a liquidity provider who holds a belief on how prices will evolve over time and use this to study the design question regarding how v3 contracts should partition the price space for permissible liquidity allocations. Our results show that making a greater diversity of price-space partitions available to a contract designer can simultaneously benefit both liquidity providers and traders.

cs.GT↗

Advanced wavefront sensing and control demonstration with MagAO-X

The search for exoplanets is pushing adaptive optics systems on ground-based telescopes to their limits. Currently, we are limited by two sources of noise: the temporal control error and non-common path aberrations. First, the temporal control error of the AO system leads to a strong residual halo. This halo can be reduced by applying predictive control. We will show and described the performance of predictive control with the 2K BMC DM in MagAO-X. After reducing the temporal control error, we can target non-common path wavefront aberrations. During the past year, we have developed a new model-free focal-plane wavefront control technique that can reach deep contrast (<1e-7 at 5 $λ$/D) on MagAO-X. We will describe the performance and discuss the on-sky implementation details and how this will push MagAO-X towards imaging planets in reflected light. The new data-driven predictive controller and the focal plane wavefront controller will be tested on-sky in April 2022.

astro-ph.IM↗

A Tighter Analysis of Spectral Clustering, and Beyond

This work studies the classical spectral clustering algorithm which embeds the vertices of some graph $G=(V_G, E_G)$ into $\mathbb{R}^k$ using $k$ eigenvectors of some matrix of $G$, and applies $k$-means to partition $V_G$ into $k$ clusters. Our first result is a tighter analysis on the performance of spectral clustering, and explains why it works under some much weaker condition than the ones studied in the literature. For the second result, we show that, by applying fewer than $k$ eigenvectors to construct the embedding, spectral clustering is able to produce better output for many practical instances; this result is the first of its kind in spectral clustering. Besides its conceptual and theoretical significance, the practical impact of our work is demonstrated by the empirical analysis on both synthetic and real-world datasets, in which spectral clustering produces comparable or better results with fewer than $k$ eigenvectors.

cs.DS↗

Nondestructive Quality Control in Powder Metallurgy using Hyperspectral Imaging

Measuring the purity in the metal powder is critical for preserving the quality of additive manufacturing products. Contamination is one of the most headache problems which can be caused by multiple reasons and lead to the as-built components cracking and malfunctions. Existing methods for metallurgical condition assessment are mostly time-consuming and mainly focus on the physical integrity of structure rather than material composition. Through capturing spectral data from a wide frequency range along with the spatial information, hyperspectral imaging (HSI) can detect minor differences in terms of temperature, moisture and chemical composition. Therefore, HSI can provide a unique way to tackle this challenge. In this paper, with the use of a near-infrared HSI camera, applications of HSI for the non-destructive inspection of metal powders are introduced. Technical assumptions and solutions on three step-by-step case studies are presented in detail, including powder characterization, contamination detection, and band selection analysis. Experimental results have fully demonstrated the great potential of HSI and related AI techniques for NDT of powder metallurgy, especially the potential to satisfy the industrial manufacturing environment.

cs.CV↗

End-to-End Sequential Sampling and Reconstruction for MRI

Accelerated MRI shortens acquisition time by subsampling in the measurement $κ$-space. Recovering a high-fidelity anatomical image from subsampled measurements requires close cooperation between two components: (1) a sampler that chooses the subsampling pattern and (2) a reconstructor that recovers images from incomplete measurements. In this paper, we leverage the sequential nature of MRI measurements, and propose a fully differentiable framework that jointly learns a sequential sampling policy simultaneously with a reconstruction strategy. This co-designed framework is able to adapt during acquisition in order to capture the most informative measurements for a particular target. Experimental results on the fastMRI knee dataset demonstrate that the proposed approach successfully utilizes intermediate information during the sampling process to boost reconstruction performance. In particular, our proposed method can outperform the current state-of-the-art learned $κ$-space sampling baseline on over 96% of test samples. We also investigate the individual and collective benefits of the sequential sampling and co-design strategies.

eess.IV↗

Finding Bipartite Components in Hypergraphs

Hypergraphs are important objects to model ternary or higher-order relations of objects, and have a number of applications in analysing many complex datasets occurring in practice. In this work we study a new heat diffusion process in hypergraphs, and employ this process to design a polynomial-time algorithm that approximately finds bipartite components in a hypergraph. We theoretically prove the performance of our proposed algorithm, and compare it against the previous state-of-the-art through extensive experimental analysis on both synthetic and real-world datasets. We find that our new algorithm consistently and significantly outperforms the previous state-of-the-art across a wide range of hypergraphs.

cs.DS↗

Hybrid Contrastive Learning with Cluster Ensemble for Unsupervised Person Re-identification

Unsupervised person re-identification (ReID) aims to match a query image of a pedestrian to the images in gallery set without supervision labels. The most popular approaches to tackle unsupervised person ReID are usually performing a clustering algorithm to yield pseudo labels at first and then exploit the pseudo labels to train a deep neural network. However, the pseudo labels are noisy and sensitive to the hyper-parameter(s) in clustering algorithm. In this paper, we propose a Hybrid Contrastive Learning (HCL) approach for unsupervised person ReID, which is based on a hybrid between instance-level and cluster-level contrastive loss functions. Moreover, we present a Multi-Granularity Clustering Ensemble based Hybrid Contrastive Learning (MGCE-HCL) approach, which adopts a multi-granularity clustering ensemble strategy to mine priority information among the pseudo positive sample pairs and defines a priority-weighted hybrid contrastive loss for better tolerating the noises in the pseudo positive samples. We conduct extensive experiments on two benchmark datasets Market-1501 and DukeMTMC-reID. Experimental results validate the effectiveness of our proposals.

cs.CV↗

alpha-Deep Probabilistic Inference (alpha-DPI): efficient uncertainty quantification from exoplanet astrometry to black hole feature extraction

Inference is crucial in modern astronomical research, where hidden astrophysical features and patterns are often estimated from indirect and noisy measurements. Inferring the posterior of hidden features, conditioned on the observed measurements, is essential for understanding the uncertainty of results and downstream scientific interpretations. Traditional approaches for posterior estimation include sampling-based methods and variational inference. However, sampling-based methods are typically slow for high-dimensional inverse problems, while variational inference often lacks estimation accuracy. In this paper, we propose alpha-DPI, a deep learning framework that first learns an approximate posterior using alpha-divergence variational inference paired with a generative neural network, and then produces more accurate posterior samples through importance re-weighting of the network samples. It inherits strengths from both sampling and variational inference methods: it is fast, accurate, and scalable to high-dimensional problems. We apply our approach to two high-impact astronomical inference problems using real data: exoplanet astrometry and black hole feature extraction.

astro-ph.IM↗

Hierarchical Clustering: $O(1)$-Approximation for Well-Clustered Graphs

Hierarchical clustering studies a recursive partition of a data set into clusters of successively smaller size, and is a fundamental problem in data analysis. In this work we study the cost function for hierarchical clustering introduced by Dasgupta, and present two polynomial-time approximation algorithms: Our first result is an $O(1)$-approximation algorithm for graphs of high conductance. Our simple construction bypasses complicated recursive routines of finding sparse cuts known in the literature. Our second and main result is an $O(1)$-approximation algorithm for a wide family of graphs that exhibit a well-defined structure of clusters. This result generalises the previous state-of-the-art, which holds only for graphs generated from stochastic models. The significance of our work is demonstrated by the empirical analysis on both synthetic and real-world data sets, on which our presented algorithm outperforms the previously proposed algorithm for graphs with a well-defined cluster structure.

cs.DS↗

Constraining the Orbit and Mass of epsilon Eridani b with Radial Velocities, Hipparcos IAD-Gaia DR2 Astrometry, and Multi-epoch Vortex Coronagraphy Upper Limits

$ε$~Eridani is a young planetary system hosting a complex multi-belt debris disk and a confirmed Jupiter-like planet orbiting at 3.48 AU from its host star. Its age and architecture are thus reminiscent of the early Solar System. The most recent study of Mawet et al. 2019, which combined radial velocity (RV) data and Ms-band direct imaging upper limits, started to constrain the planet's orbital parameters and mass, but are still affected by large error bars and degeneracies. Here we make use of the most recent data compilation from three different techniques to further refine $ε$~Eridani~b's properties: RVs, absolute astrometry measurements from the Hipparcos~and Gaia~missions, and new Keck/NIRC2 Ms-band vortex coronagraph images. We combine this data in a Bayesian framework. We find a new mass, $M_b$ = $0.66_{-0.09}^{+0.12}$~M$_{Jup}$, and inclination, $i$ = $77.95_{-21.06}^{\circ+28.50}$, with at least a factor 2 improvement over previous uncertainties. We also report updated constraints on the longitude of the ascending node, the argument of the periastron, and the time of periastron passage. With these updated parameters, we can better predict the position of the planet at any past and future epoch, which can greatly help define the strategy and planning of future observations and with subsequent data analysis. In particular, these results can assist the search for a direct detection with JWST and the Nancy Grace Roman Space Telescope's coronagraph instrument (CGI).

astro-ph.EP↗