Searcharxiv⌕ Search

arXiv subjects

Hanbing Liu

Publications and source records attributed to Hanbing Liu.

29 records · Page 2Linked to original sources

Refined Temporal Pyramidal Compression-and-Amplification Transformer for 3D Human Pose Estimation

Accurately estimating the 3D pose of humans in video sequences requires both accuracy and a well-structured architecture. With the success of transformers, we introduce the Refined Temporal Pyramidal Compression-and-Amplification (RTPCA) transformer. Exploiting the temporal dimension, RTPCA extends intra-block temporal modeling via its Temporal Pyramidal Compression-and-Amplification (TPCA) structure and refines inter-block feature interaction with a Cross-Layer Refinement (XLR) module. In particular, TPCA block exploits a temporal pyramid paradigm, reinforcing key and value representation capabilities and seamlessly extracting spatial semantics from motion sequences. We stitch these TPCA blocks with XLR that promotes rich semantic representation through continuous interaction of queries, keys, and values. This strategy embodies early-stage information with current flows, addressing typical deficits in detail and stability seen in other transformer-based methods. We demonstrate the effectiveness of RTPCA by achieving state-of-the-art results on Human3.6M, HumanEva-I, and MPI-INF-3DHP benchmarks with minimal computational overhead. The source code is available at https://github.com/hbing-l/RTPCA.

cs.CV↗

PoSynDA: Multi-Hypothesis Pose Synthesis Domain Adaptation for Robust 3D Human Pose Estimation

Existing 3D human pose estimators face challenges in adapting to new datasets due to the lack of 2D-3D pose pairs in training sets. To overcome this issue, we propose \textit{Multi-Hypothesis \textbf{P}ose \textbf{Syn}thesis \textbf{D}omain \textbf{A}daptation} (\textbf{PoSynDA}) framework to bridge this data disparity gap in target domain. Typically, PoSynDA uses a diffusion-inspired structure to simulate 3D pose distribution in the target domain. By incorporating a multi-hypothesis network, PoSynDA generates diverse pose hypotheses and aligns them with the target domain. To do this, it first utilizes target-specific source augmentation to obtain the target domain distribution data from the source domain by decoupling the scale and position parameters. The process is then further refined through the teacher-student paradigm and low-rank adaptation. With extensive comparison of benchmarks such as Human3.6M and MPI-INF-3DHP, PoSynDA demonstrates competitive performance, even comparable to the target-trained MixSTE model\cite{zhang2022mixste}. This work paves the way for the practical application of 3D human pose estimation in unseen domains. The code is available at https://github.com/hbing-l/PoSynDA.

cs.CV↗

KeyPosS: Plug-and-Play Facial Landmark Detection through GPS-Inspired True-Range Multilateration

Accurate facial landmark detection is critical for facial analysis tasks, yet prevailing heatmap and coordinate regression methods grapple with prohibitive computational costs and quantization errors. Through comprehensive theoretical analysis and experimentation, we identify and elucidate the limitations of existing techniques. To overcome these challenges, we pioneer the application of True-Range Multilateration, originally devised for GPS localization, to facial landmark detection. We propose KeyPoint Positioning System (KeyPosS) - the first framework to deduce exact landmark coordinates by triangulating distances between points of interest and anchor points predicted by a fully convolutional network. A key advantage of KeyPosS is its plug-and-play nature, enabling flexible integration into diverse decoding pipelines. Extensive experiments on four datasets demonstrate state-of-the-art performance, with KeyPosS outperforming existing methods in low-resolution settings despite minimal computational overhead. By spearheading the integration of Multilateration with facial analysis, KeyPosS marks a paradigm shift in facial landmark detection. The code is available at https://github.com/zhiqic/KeyPosS.

cs.CV↗

HDFormer: High-order Directed Transformer for 3D Human Pose Estimation

Human pose estimation is a challenging task due to its structured data sequence nature. Existing methods primarily focus on pair-wise interaction of body joints, which is insufficient for scenarios involving overlapping joints and rapidly changing poses. To overcome these issues, we introduce a novel approach, the High-order Directed Transformer (HDFormer), which leverages high-order bone and joint relationships for improved pose estimation. Specifically, HDFormer incorporates both self-attention and high-order attention to formulate a multi-order attention module. This module facilitates first-order "joint$\leftrightarrow$joint", second-order "bone$\leftrightarrow$joint", and high-order "hyperbone$\leftrightarrow$joint" interactions, effectively addressing issues in complex and occlusion-heavy situations. In addition, modern CNN techniques are integrated into the transformer-based architecture, balancing the trade-off between performance and efficiency. HDFormer significantly outperforms state-of-the-art (SOTA) models on Human3.6M and MPI-INF-3DHP datasets, requiring only 1/10 of the parameters and significantly lower computational costs. Moreover, HDFormer demonstrates broad real-world applicability, enabling real-time, accurate 3D pose estimation. The source code is in https://github.com/hyer/HDFormer

cs.CV↗

ChatPipe: Orchestrating Data Preparation Program by Optimizing Human-ChatGPT Interactions

Orchestrating a high-quality data preparation program is essential for successful machine learning (ML), but it is known to be time and effort consuming. Despite the impressive capabilities of large language models like ChatGPT in generating programs by interacting with users through natural language prompts, there are still limitations. Specifically, a user must provide specific prompts to iteratively guide ChatGPT in improving data preparation programs, which requires a certain level of expertise in programming, the dataset used and the ML task. Moreover, once a program has been generated, it is non-trivial to revisit a previous version or make changes to the program without starting the process over again. In this paper, we present ChatPipe, a novel system designed to facilitate seamless interaction between users and ChatGPT. ChatPipe provides users with effective recommendation on next data preparation operations, and guides ChatGPT to generate program for the operations. Also, ChatPipe enables users to easily roll back to previous versions of the program, which facilitates more efficient experimentation and testing. We have developed a web application for ChatPipe and prepared several real-world ML tasks from Kaggle. These tasks can showcase the capabilities of ChatPipe and enable VLDB attendees to easily experiment with our novel features to rapidly orchestrate a high-quality data preparation program.

cs.DB↗

Stabilizability of linear systems with discrete observation mode

For linear control systems, the usual state feedback stabilizability has two components: one is a continuous observation mode (i.e., to observe solutions continuously in time), and the other is a class of feedback laws (which is usually the space of all of the linear and bounded operators from a state space to a control space). This paper studies the stabilizability for abstract linear control systems, with a discrete observation mode (i.e., to observe solutions discretely in time) and two different classes of feedback laws. We first characterize these types of stabilizabilities via some weak observability inequalities for the dual systems. Then, we use these characterizations to reveal the connections between these types of stabilizabilities and those with continuous observation mode. Finally, we show some applications of the aforementioned weak observability inequalities.

math.OC↗

Characterizations of complete stabilizability

We present several characterizations, via some weak observability inequalities, on the complete stabilizability for a control system $[A,B]$, i.e., $y'(t)=Ay(t)+Bu(t)$, $t\geq 0$, where $A$ generates a $C_0$-semigroup on a Hilbert space $X$ and $B$ is a linear and bounded operator from another Hilbert space $U$ to $X$. We then extend the aforementioned characterizations in two directions: first, the control operator $B$ is unbounded; second, the control system is time-periodic. We also give some sufficient conditions, from the perspective of the spectral projections, to ensure the weak observability inequalities. As applications, we provide several examples, which are not null controllable, but can be verified, via the weak observability inequalities, to be completely stabilizable.

math.OC↗

Boundary sampled-data feedback stabilization for parabolic equations

The aim of this work is to design an explicit finite dimensional boundary feedback controller of sampled-data form for locally exponentially stabilizing the equilibrium solutions to semilinear parabolic equations. The feedback controller is expressed in terms of the eigenfunctions corresponding to unstable eigenvalues of the linearized equation. This stabilizing procedure is applicable for any sampling rate, not necessary to be small enough, and it tends to the continuous-times version when the sampling period tends to zero.

math.OC↗

Output feedback stabilization for heat equations with sampled-data controls

In this paper, we build up an output feedback law to stabilize a sampled-data controlled heat equation (with a potential) in a bounded domain $Ω$. The feedback law abides the following rules: First, we divide equally the time interval $[0,+\infty)$ into infinitely many disjoint time periods, and divide each time period into three disjoint subintervals. Second, for each time period, we observe a solution over an open subset of $Ω$ in the first subinterval, take sample from outputs at one time point of the first subinterval, add a time-invariant output feedback control over another open subset of $Ω$ in the second subinterval; let the equation evolve free in the last subinterval. Thus, the corresponding feedback control is of sampled-data. Our feedback law has the following advantages: the sampling period (which is the length of the above time period) can be arbitrarily taken; the feedback law has an explicit expression in terms of the sampling period; the behaviors of the norm of the feedback law, when the sampling period goes to zero or infinity, are clear. The construction of the feedback law is based on two kinds of approximate null-controllability for heat equations. One has time-invariant controls, while another has impulse controls. The studies of the aforementioned controllability with time-invariant controls need a new observability inequality for heat equations built up in the current work.

math.OC↗

Boundary feedback stabilization of Fisher's equation

The aim of this work is to design an explicit finite dimensional boundary feedback controller for locally exponentially stabilizing the equilibrium solutions to Fisher's equation in both $L^2(0,1)$ and $H^1(0,1)$. The feedback controller is expressed in terms of the eigenfunctions corresponding to unstable eigenvalues of the linearized equation. This stabilizing procedure is applicable for any level of instability, which extends the result of \cite{02} for nonlinear parabolic equations. The effectiveness of the approach is illustrated by a numerical simulation.

math.OC↗