arXiv · 2609.33973
FINGR: Learning Dexterous Hand Control for Real-World Rubik's Cube Solving
Abstract
Manipulating a Rubik's Cube with a single dexterous hand is a challenging test of sustained, contact-rich control: the hand must execute successive layer turns while keeping the cube secure. Each turn requires some fingers to support the cube while others push a moving layer, release contact, and reset for the next move. To learn this coordination, we introduce FINGR (Future-supervised Interaction Network with Geometric Representations), a policy that combines finger-relative geometry with future interaction prediction. A shared point encoder expresses the cube relative to each fingertip and aggregates its points without depending on cubie indexing. Learned future tokens share the observation encoder and receive supervision for contact-force changes, layer-turn progress, and finger joint displacement at multiple time scales. The resulting representation conditions a flow policy that directly generates finger actions. On a real dexterous hand, our policy achieves 99.0% success over 300 turn attempts, compared with 79.7% for the base flow policy. Integrated with grasping and table-assisted regrasping, the policy solves all ten scrambled $2\times2\times2$ cubes in a mean complete-system time of approximately 137 seconds. The project website is available at https://www.lyt0112.com/projects/FINGR
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yutong Liang, Quanquan Peng, Matthew Kim, Xiaolong Wang. 2026-09-27. FINGR: Learning Dexterous Hand Control for Real-World Rubik's Cube Solving. https://arxiv.org/abs/2609.33973
Cite the original work for its findings. Save a collection to share your selection of sources.