SearcharxivSearch

arXiv subjects

Jiyang Gao

Publications and source records attributed to Jiyang Gao.

At least 19 recordsLinked to original sources

RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation

Despite the critical role of bimanual manipulation in endowing robots with human-like dexterity, large-scale and diverse datasets remain scarce due to the significant hardware heterogeneity across bimanual robotic platforms. To bridge this gap, we introduce RoboCOIN, a large-scale multi-embodiment bimanual manipulation dataset comprising over 180,000 demonstrations collected from 15 distinct robotic platforms. Spanning 16 diverse environments-including residential, commercial, and industrial settings-the dataset features 421 bimanual tasks systematically categorized by 39 bimanual collaboration actions and 432 objects. A key innovation of our work is the hierarchical capability pyramid, which provides granular annotations ranging from trajectory-level concepts to segment-level subtasks and frame-level kinematics. Furthermore, we present CoRobot, an efficient data processing pipeline powered by the Robot Trajectory Markup Language (RTML), designed to facilitate quality assessment, automated annotation, and unified multi-embodiment and data management. Extensive experiments demonstrate the effectiveness of RoboCOIN in enhancing the performance of various bimanual manipulation models across a wide spectrum of robotic embodiments. The entire dataset and codebase are fully open-sourced, providing a valuable resource for advancing research in bimanual and multi-embodiment manipulation.

cs.RO

Rigidity matroids and linear algebraic matroids with applications to matrix completion and tensor codes

We establish a connection between problems studied in rigidity theory and matroids arising from linear algebraic constructions like tensor products and symmetric products. A special case of this correspondence identifies the problem of giving a description of the correctable erasure patterns in a maximally recoverable tensor code with the problem of describing bipartite rigid graphs or low-rank completable matrix patterns. Additionally, we relate dependencies among symmetric products of generic vectors to graph rigidity and symmetric matrix completion. With an eye toward applications to computer science, we study the dependency of these matroids on the characteristic by giving new combinatorial descriptions in several cases, including the first description of the correctable patterns in an (m, n, a=2, b=2) maximally recoverable tensor code.

math.CO

Tilted Richardson Varieties

The study of the flag variety $\mathrm{Fl}_n$ and its subvarieties, including Schubert and Richardson varieties, plays a fundamental role in algebraic geometry and algebraic combinatorics. In this paper, we introduce and develop the theory of tilted Richardson varieties $\mathrm{T}_{u,v}$, a new family of subvarieties of the flag variety that provides a geometric framework for the quantum Bruhat graphs. These varieties are defined for all pairs of permutations $u$ and $v$, extending the classical Richardson varieties in the case where $u\leq v$ in the Bruhat order. We establish their fundamental geometric properties, proving irreducibility and providing explicit dimension formulas. Moreover, we show that they have a well-defined stratification indexed by tilted Bruhat intervals, a generalization of classical Bruhat intervals previously introduced by Brenti, Fomin, and Postnikov. Additionally, we introduce a tilted generalization of the classical Deodhar decomposition of Richardson varieties, which leads to a combinatorial formula for tilted Kazhdan--Lusztig R-polynomials, a notion that arises naturally in our framework. We further develop a theory of total positivity for tilted Richardson varieties. In particular, we define and study the totally nonnegative parts of tilted Richardson varieties, proving they form a CW complex. This generalizes earlier results on the totally nonnegative flag variety and answers Björner's questions regarding geometric realizations of tilted Bruhat intervals. Finally, we establish explicit connections between tilted Richardson varieties and quantum Schubert calculus. Specifically, we prove that $\mathrm{T}_{u,v}$ coincides with minimal-degree two-point curve neighborhoods. As a result, we compute their cohomology classes and derive new relationships among Gromov--Witten invariants of the flag variety.

math.CO

Galaxea Open-World Dataset and G0 Dual-System VLA Model

We present Galaxea Open-World Dataset, a large-scale, diverse collection of robot behaviors recorded in authentic human living and working environments. All demonstrations are gathered using a consistent robotic embodiment, paired with precise subtask-level language annotations to facilitate both training and evaluation. Building on this dataset, we introduce G0, a dual-system framework that couples a Vision-Language Model (VLM) for multimodal planning with a Vision-Language-Action (VLA) model for fine-grained execution. G0 is trained using a three-stage curriculum: cross-embodiment pre-training, single-embodiment pre-training, and task-specific post-training. A comprehensive benchmark spanning tabletop manipulation, few-shot learning, and long-horizon mobile manipulation, demonstrates the effectiveness of our approach. In particular, we find that the single-embodiment pre-training stage, together with the Galaxea Open-World Dataset, plays a critical role in achieving strong performance.

cs.RO

Characterizing positroid quotients of uniform matroids

We study two-step flag positroids $(P_1, P_2)$, where $P_1$ is a quotient of $P_{2}$. We provide a complete characterization of all two-step flag positroids that contain a uniform matroid, extending and completing a partial result by Benedetti, Chávez, and Jiménez. To contrast general positroids with the special case of lattice path matroids, we show that the containment relations of Grassmann necklaces and conecklaces fully characterize flag lattice path matroids, but are insufficient for general flag positroids. Additionally, we prove that the decorated permutations of any elementary quotient pair are related by a cyclic shift, resolving a conjecture of Benedetti, Chávez and Jiménez.

math.CO

On two notions of total positivity for generalized partial flag varieties of classical Lie types

For Grassmannians, Lusztig's notion of total positivity coincides with positivity of the Plucker coordinates. This coincidence underpins the rich interaction between matroid theory, tropical geometry, and the theory of total positivity. Bloch and Karp furthermore characterized the (type A) partial flag varieties for which the two notions of positivity similarly coincide. We characterize the symplectic (type C) and odd-orthogonal (type B) partial flag varieties for which Lusztig's total positivity coincides with Plucker positivity.

math.CO

Sandpile Groups of Cayley Graphs of $\mathbb{F}_2^r$

The sandpile group of a connected graph $G$, defined to be the torsion part of the cokernel of the graph Laplacian, is a subtle graph invariant with combinatorial, algebraic, and geometric descriptions. Extending and improving previous works on the sandpile group of hypercubes, we study the sandpile groups of the Cayley graphs of $\mathbb{F}_2^r$, focusing on their poorly understood Sylow-$2$ component. We find the number of Sylow-$2$ cyclic factors for "generic" Cayley graphs and deduce a bound for the non-generic ones. Moreover, we provide a sharp upper bound for their largest Sylow-$2$ cyclic factors. In the case of hypercubes, we give exact formulae for the largest $n-1$ Sylow-$2$ cyclic factors. Some key ingredients of our work include the natural ring structure on these sandpile groups from representation theory, and calculation of the $2$-adic valuations of binomial sums via the combinatorics of carries.

math.CO

Quantum Bruhat graphs and tilted Richardson varieties

Quantum Bruhat graph is a weighted directed graph on a finite Weyl group first defined by Brenti-Fomin-Postnikov. It encodes quantum Monk's rule and can be utilized to study the $3$-point Gromov-Witten invariants of the flag variety. In this paper, we provide an explicit formula for the minimal weights between any pair of permutations on the quantum Bruhat graph, and consequently obtain an Ehresmann-like characterization for the tilted Bruhat order. Moreover, for any ordered pair of permutations $u$ and $v$, we define the tilted Richardson variety $T_{u,v}$, with a stratification that gives a geometric meaning to intervals in the tilted Bruhat order. We provide a few equivalent definitions to this new family of varieties that include Richardson varieties, and establish some fundamental geometric properties including their dimensions and closure relations.

math.CO

Acyclic Orientations and the Chromatic Polynomial of Signed Graphs

We present a new correspondence between acyclic orientations and coloring of a signed graph (symmetric graph). Goodall et al. introduced a bivariate chromatic polynomial $χ_G(k,l)$ that counts the number of signed colorings using colors $0,\pm1,\dots,\pm k$ along with $l-1$ symmetric colors $0_1,\dots,0_{l-1}$. We show that the evaluation of the bivariate chromatic polynomial $|χ_G(-1,2)|$ is equal to the number of acyclic orientations of the signed graph modulo the equivalence relation generated by swapping sources and sinks. We present three proofs of this fact, a proof using toric hyperplane arrangements, a proof using deletion-contraction, and a direct proof.

math.CO

Balanced shifted tableaux

We introduce balanced shifted tableaux, as an analogue of balanced tableaux of Edelman and Greene, from the perspective of root systems of type B and C. We show that they are equinumerous to standard Young tableaux of the corresponding shifted shape by presenting an explicit bijection.

math.CO

HDMapGen: A Hierarchical Graph Generative Model of High Definition Maps

High Definition (HD) maps are maps with precise definitions of road lanes with rich semantics of the traffic rules. They are critical for several key stages in an autonomous driving system, including motion forecasting and planning. However, there are only a small amount of real-world road topologies and geometries, which significantly limits our ability to test out the self-driving stack to generalize onto new unseen scenarios. To address this issue, we introduce a new challenging task to generate HD maps. In this work, we explore several autoregressive models using different data representations, including sequence, plain graph, and hierarchical graph. We propose HDMapGen, a hierarchical graph generation model capable of producing high-quality and diverse HD maps through a coarse-to-fine approach. Experiments on the Argoverse dataset and an in-house dataset show that HDMapGen significantly outperforms baseline methods. Additionally, we demonstrate that HDMapGen achieves high scalability and efficiency.

cs.CV

TNT: Target-driveN Trajectory Prediction

Predicting the future behavior of moving agents is essential for real world applications. It is challenging as the intent of the agent and the corresponding behavior is unknown and intrinsically multimodal. Our key insight is that for prediction within a moderate time horizon, the future modes can be effectively captured by a set of target states. This leads to our target-driven trajectory prediction (TNT) framework. TNT has three stages which are trained end-to-end. It first predicts an agent's potential target states $T$ steps into the future, by encoding its interactions with the environment and the other agents. TNT then generates trajectory state sequences conditioned on targets. A final stage estimates trajectory likelihoods and a final compact set of trajectory predictions is selected. This is in contrast to previous work which models agent intents as latent variables, and relies on test-time sampling to generate diverse trajectories. We benchmark TNT on trajectory prediction of vehicles and pedestrians, where we outperform state-of-the-art on Argoverse Forecasting, INTERACTION, Stanford Drone and an in-house Pedestrian-at-Intersection dataset.

cs.CV

Virtual Complete Intersections in $\mathbb{P}^1 \times \mathbb{P}^1$

The minimal free resolution of the coordinate ring of a complete intersection in projective space is a Koszul complex on a regular sequence. In the product of projective spaces $\mathbb{P}^1 \times \mathbb{P}^1$, we investigate which sets of points have a virtual resolution that is a Koszul complex on a regular sequence. This paper provides conditions on sets of points; some of which guarantee the points have this property, and some of which guarantee the points do not have this property.

math.AG

STINet: Spatio-Temporal-Interactive Network for Pedestrian Detection and Trajectory Prediction

Detecting pedestrians and predicting future trajectories for them are critical tasks for numerous applications, such as autonomous driving. Previous methods either treat the detection and prediction as separate tasks or simply add a trajectory regression head on top of a detector. In this work, we present a novel end-to-end two-stage network: Spatio-Temporal-Interactive Network (STINet). In addition to 3D geometry modeling of pedestrians, we model the temporal information for each of the pedestrians. To do so, our method predicts both current and past locations in the first stage, so that each pedestrian can be linked across frames and the comprehensive spatio-temporal information can be captured in the second stage. Also, we model the interaction among objects with an interaction graph, to gather the information among the neighboring objects. Comprehensive experiments on the Lyft Dataset and the recently released large-scale Waymo Open Dataset for both object detection and future trajectory prediction validate the effectiveness of the proposed method. For the Waymo Open Dataset, we achieve a bird-eyes-view (BEV) detection AP of 80.73 and trajectory prediction average displacement error (ADE) of 33.67cm for pedestrians, which establish the state-of-the-art for both tasks.

cs.CV

VectorNet: Encoding HD Maps and Agent Dynamics from Vectorized Representation

Behavior prediction in dynamic, multi-agent systems is an important problem in the context of self-driving cars, due to the complex representations and interactions of road components, including moving agents (e.g. pedestrians and vehicles) and road context information (e.g. lanes, traffic lights). This paper introduces VectorNet, a hierarchical graph neural network that first exploits the spatial locality of individual road components represented by vectors and then models the high-order interactions among all components. In contrast to most recent approaches, which render trajectories of moving agents and road context information as bird-eye images and encode them with convolutional neural networks (ConvNets), our approach operates on a vector representation. By operating on the vectorized high definition (HD) maps and agent trajectories, we avoid lossy rendering and computationally intensive ConvNet encoding steps. To further boost VectorNet's capability in learning context features, we propose a novel auxiliary task to recover the randomly masked out map entities and agent trajectories based on their context. We evaluate VectorNet on our in-house behavior prediction benchmark and the recently released Argoverse forecasting dataset. Our method achieves on par or better performance than the competitive rendering approach on both benchmarks while saving over 70% of the model parameters with an order of magnitude reduction in FLOPs. It also outperforms the state of the art on the Argoverse dataset.

cs.CV

CPARR: Category-based Proposal Analysis for Referring Relationships

The task of referring relationships is to localize subject and object entities in an image satisfying a relationship query, which is given in the form of \texttt{ }. This requires simultaneous localization of the subject and object entities in a specified relationship. We introduce a simple yet effective proposal-based method for referring relationships. Different from the existing methods such as SSAS, our method can generate a high-resolution result while reducing its complexity and ambiguity. Our method is composed of two modules: a category-based proposal generation module to select the proposals related to the entities and a predicate analysis module to score the compatibility of pairs of selected proposals. We show state-of-the-art performance on the referring relationship task on two public datasets: Visual Relationship Detection and Visual Genome.

cs.CV

End-to-End Multi-View Fusion for 3D Object Detection in LiDAR Point Clouds

Recent work on 3D object detection advocates point cloud voxelization in birds-eye view, where objects preserve their physical dimensions and are naturally separable. When represented in this view, however, point clouds are sparse and have highly variable point density, which may cause detectors difficulties in detecting distant or small objects (pedestrians, traffic signs, etc.). On the other hand, perspective view provides dense observations, which could allow more favorable feature encoding for such cases. In this paper, we aim to synergize the birds-eye view and the perspective view and propose a novel end-to-end multi-view fusion (MVF) algorithm, which can effectively learn to utilize the complementary information from both. Specifically, we introduce dynamic voxelization, which has four merits compared to existing voxelization methods, i) removing the need of pre-allocating a tensor with fixed size; ii) overcoming the information loss due to stochastic point/voxel dropout; iii) yielding deterministic voxel embeddings and more stable detection outcomes; iv) establishing the bi-directional relationship between points and voxels, which potentially lays a natural foundation for cross-view feature fusion. By employing dynamic voxelization, the proposed feature fusion architecture enables each point to learn to fuse context information from different views. MVF operates on points and can be naturally extended to other approaches using LiDAR point clouds. We evaluate our MVF model extensively on the newly released Waymo Open Dataset and on the KITTI dataset and demonstrate that it significantly improves detection accuracy over the comparable single-view PointPillars baseline.

cs.CV

MAC: Mining Activity Concepts for Language-based Temporal Localization

We address the problem of language-based temporal localization in untrimmed videos. Compared to temporal localization with fixed categories, this problem is more challenging as the language-based queries not only have no pre-defined activity list but also may contain complex descriptions. Previous methods address the problem by considering features from video sliding windows and language queries and learning a subspace to encode their correlation, which ignore rich semantic cues about activities in videos and queries. We propose to mine activity concepts from both video and language modalities by applying the actionness score enhanced Activity Concepts based Localizer (ACL). Specifically, the novel ACL encodes the semantic concepts from verb-obj pairs in language queries and leverages activity classifiers' prediction scores to encode visual concepts. Besides, ACL also has the capability to regress sliding windows as localization results. Experiments show that ACL significantly outperforms state-of-the-arts under the widely used metric, with more than 5% increase on both Charades-STA and TACoS datasets.

cs.CV