SearcharxivSearch

arXiv subjects

Hongyi Fan

Publications and source records attributed to Hongyi Fan.

11 recordsLinked to original sources

Multiview Image-Based Localization

The image retrieval (IR) approach to image localization has distinct advantages to the 3D and the deep learning (DNN) approaches: it is seen-agnostic, simpler to implement and use, has no privacy issues, and is computationally efficient. The main drawback of this approach is relatively poor localization in both position and orientation of the query camera when compared to the competing approaches. This paper represents a hybrid approach that stores only image features in the database like some IR methods, but relies on a latent 3D reconstruction, like 3D methods but without retaining a 3D scene reconstruction. The approach is based on two ideas: {\em (i)} a novel proposal where query camera center estimation relies only on relative translation estimates but not relative rotation estimates through a decoupling of the two, and {\em (ii)} a shift from computing optimal pose from estimated relative pose to computing optimal pose from multiview correspondences, thus cutting out the ``middle-man''. Our approach shows improved performance on the 7-Scenes and Cambridge Landmarks datasets while also improving on timing and memory footprint as compared to state-of-the-art.

cs.CV

Integrating 3D Slicer with a Dynamic Simulator for Situational Aware Robotic Interventions

Image-guided robotic interventions represent a transformative frontier in surgery, blending advanced imaging and robotics for improved precision and outcomes. This paper addresses the critical need for integrating open-source platforms to enhance situational awareness in image-guided robotic research. We present an open-source toolset that seamlessly combines a physics-based constraint formulation framework, AMBF, with a state-of-the-art imaging platform application, 3D Slicer. Our toolset facilitates the creation of highly customizable interactive digital twins, that incorporates processing and visualization of medical imaging, robot kinematics, and scene dynamics for real-time robot control. Through a feasibility study, we showcase real-time synchronization of a physical robotic interventional environment in both 3D Slicer and AMBF, highlighting low-latency updates and improved visualization.

cs.RO

Condition numbers in multiview geometry, instability in relative pose estimation, and RANSAC

In this paper, we introduce a general framework for analyzing the numerical conditioning of minimal problems in multiple view geometry, using tools from computational algebra and Riemannian geometry. Special motivation comes from the fact that relative pose estimation, based on standard 5-point or 7-point Random Sample Consensus (RANSAC) algorithms, can fail even when no outliers are present and there is enough data to support a hypothesis. We argue that these cases arise due to the intrinsic instability of the 5- and 7-point minimal problems. We apply our framework to characterize the instabilities, both in terms of the world scenes that lead to infinite condition number, and directly in terms of ill-conditioned image data. The approach produces computational tests for assessing the condition number before solving the minimal problem. Lastly, synthetic and real data experiments suggest that RANSAC serves not only to remove outliers, but in practice it also selects for well-conditioned image data, which is consistent with our theory.

cs.CV

On the Instability of Relative Pose Estimation and RANSAC's Role

In this paper we study the numerical instabilities of the 5- and 7-point problems for essential and fundamental matrix estimation in multiview geometry. In both cases we characterize the ill-posed world scenes where the condition number for epipolar estimation is infinite. We also characterize the ill-posed instances in terms of the given image data. To arrive at these results, we present a general framework for analyzing the conditioning of minimal problems in multiview geometry, based on Riemannian manifolds. Experiments with synthetic and real-world data then reveal a striking conclusion: that Random Sample Consensus (RANSAC) in Structure-from-Motion (SfM) does not only serve to filter out outliers, but RANSAC also selects for well-conditioned image data, sufficiently separated from the ill-posed locus that our theory predicts. Our findings suggest that, in future work, one could try to accelerate and increase the success of RANSAC by testing only well-conditioned image data.

cs.CV

Benchmarking Pedestrian Odometry: The Brown Pedestrian Odometry Dataset (BPOD)

We present the Brown Pedestrian Odometry Dataset (BPOD) for benchmarking visual odometry algorithms in head-mounted pedestrian settings. This dataset was captured using synchronized global and rolling shutter stereo cameras in 12 diverse indoor and outdoor locations on Brown University's campus. Compared to existing datasets, BPOD contains more image blur and self-rotation, which are common in pedestrian odometry but rare elsewhere. Ground-truth trajectories are generated from stick-on markers placed along the pedestrian's path, and the pedestrian's position is documented using a third-person video. We evaluate the performance of representative direct, feature-based, and learning-based VO methods on BPOD. Our results show that significant development is needed to successfully capture pedestrian trajectories. The link to the dataset is here: \url{https://doi.org/10.26300/c1n7-7p93

cs.CV

GPU-Based Homotopy Continuation for Minimal Problems in Computer Vision

Systems of polynomial equations arise frequently in computer vision, especially in multiview geometry problems. Traditional methods for solving these systems typically aim to eliminate variables to reach a univariate polynomial, e.g., a tenth-order polynomial for 5-point pose estimation, using clever manipulations, or more generally using Grobner basis, resultants, and elimination templates, leading to successful algorithms for multiview geometry and other problems. However, these methods do not work when the problem is complex and when they do, they face efficiency and stability issues. Homotopy Continuation (HC) can solve more complex problems without the stability issues, and with guarantees of a global solution, but they are known to be slow. In this paper we show that HC can be parallelized on a GPU, showing significant speedups up to 26 times on polynomial benchmarks. We also show that GPU-HC can be generically applied to a range of computer vision problems, including 4-view triangulation and trifocal pose estimation with unknown focal length, which cannot be solved with elimination template but they can be efficiently solved with HC. GPU-HC opens the door to easy formulation and solution of a range of computer vision problems.

cs.CV

Operator transpose within normal ordering and its applications for quantifying entanglement

Partial transpose is an important operation for quantifying the entanglement, here we study the (partial) transpose of any single (two-mode) operators. Using the Fock-basis expansion, it is found that the transposed operator of an arbitrary operator can be obtained by replacement of a^{{\dag}}(a) by a(a^{{\dag}}) instead of c-number within normal ordering form. The transpose of displacement operator and Wigner operator are studied, from which the relation of Wigner function, characteristics function and average values such as covariance matrix are constructed between density operator and transposed density operator. These observations can be further extended to multi-mode cases. As applications, the partial transpose of two-mode squeezed operator and the entanglement of two-mode squeezed vacuum through a laser channel are considered.

quant-ph

Trifocal Relative Pose from Lines at Points and its Efficient Solution

We present a method for solving two minimal problems for relative camera pose estimation from three views, which are based on three view correspondences of i) three points and one line and the novel case of ii) three points and two lines through two of the points. These problems are too difficult to be efficiently solved by the state of the art Groebner basis methods. Our method is based on a new efficient homotopy continuation (HC) solver framework MINUS, which dramatically speeds up previous HC solving by specializing HC methods to generic cases of our problems. We characterize their number of solutions and show with simulated experiments that our solvers are numerically robust and stable under image noise, a key contribution given the borderline intractable degree of nonlinearity of trinocular constraints. We show in real experiments that i) SIFT feature location and orientation provide good enough point-and-line correspondences for three-view reconstruction and ii) that we can solve difficult cases with too few or too noisy tentative matches, where the state of the art structure from motion initialization fails.

cs.CV

Partition of unity with mixed quantum states

The completeness of quantum state space, is usually expressed as \sum_{m=0}^{\infty}|m> } is selected set of quantum states (basis). Density matrix |m><m| describes a pure quantum state. In this paper, by virtue of the summation method within the normally ordered product of operators we propose and show that the completeness relation can also be represented or partitioned in terms of some mixed states, such as binomial states and negative binomial states. Thus the view on the structure of Fock space is widen and the connotation of Fock space is enriched. See the matter in this sight, experimentalists may have interests to prepare the binomial- and negative binomial states.

quant-ph

Weyl correspondence method to construct multipartite entangled quantum state

Via the Weyl correspondence approach, we construct multipartite entangled state which is the common eigenvector of their center-of-mass coordinate and mass-weighted relative momenta. This approach is concise and effective for setting up the Fock representation of continuous multipartite entangled states. The technique of integration within an ordered product (IWOP) of operators is also essential in our derivation.

quant-ph

Higher-order properties and Bell-inequality violation for the three-mode enhanced squeezed state

By extending the usual two-mode squeezing operator $S_{2}=\exp [ i\lambda (Q_{1}P_{2}+Q_{2}P_{1}) ] $ to the three-mode squeezing operator $S_{3}=\exp {i\lambda [ Q_{1}(P_{2}+P_{3}) +Q_{2}(P_{1}+P_{3}) +Q_{3}(P_{1}+P_{2}) ]} $, we obtain the corresponding three-mode squeezed coherent state. The state's higher-order properties, such as higher-order squeezing and higher-order sub-Possonian photon statistics, are investigated. It is found that the new squeezed state not only can be squeezed to all even orders but also exhibits squeezing enhancement comparing with the usual cases. In addition, we examine the violation of Bell-inequality for the three-mode squeezed states by using the formalism of Wigner representation.

quant-ph