Searcharxiv⌕ Search

arXiv · 2610.02922

TwinJEPA: Action-Preferred Predictive Representations for Goal-Conditioned Control

Abstract

Joint-Embedding Predictive Architectures (JEPAs) are a promising paradigm for representation learning by predicting future latent states without reconstructing observations. Recent work has adapted JEPA-style latent prediction to offline zero-shot control through action-conditioned temporal prediction, enabling representations to capture long-horizon dynamics from fixed behavioral data. However, transition-level predictive objectives supervise actions in isolation and provide limited information about which actions are preferable when similar states admit different goal-conditioned outcomes. We introduce TwinJEPA, a framework for learning action-preferred predictive representations by augmenting JEPA-based control with offline-mined action-preference supervision. TwinJEPA identifies approximately matched states across offline trajectories and constructs preference pairs via goal-conditioned reward relabeling. It then learns two complementary objectives: reward-gap regression, which preserves the magnitude of outcome differences, and preference classification, which captures the relative ordering of alternative actions. Both objectives are used only during training and incur no additional inference-time cost. We evaluate TwinJEPA on long-horizon navigation and continuous-control benchmarks with state- and pixel-based observations. TwinJEPA yields positive benchmark-level mean differences across all matched state-based evaluations, while analyses across domains and observation modalities indicate that larger gains tend to arise when local action alternatives provide more informative outcome contrasts. The results suggest that local action-comparison supervision can complement temporal prediction by encouraging JEPA-based representations to retain both long-horizon temporal structure and fine-grained distinctions among locally observed actions for offline zero-shot control.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Feiran You, Hongyang Du. 2026-10-02. TwinJEPA: Action-Preferred Predictive Representations for Goal-Conditioned Control. https://arxiv.org/abs/2610.02922

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

From Inference to Control: Structure-Guided Control of Hypergraph Dynamics

Controllability determines whether a system's state can be guided toward any desired configuration, making it a fundamental prerequisite for designing effective control strategies. In the context of networked systems, controllability is a well-established concept. However, many real-world systems, from biological collectives to engineered infrastructures, exhibit higher-order interactions that cannot be captured by simple graphs. Moreover, the interaction structures might be unknown and difficult to measure directly. Here, we close this gap by combining hypergraph inference with the identification of controllable nodes. Building on the inferred structure, we design a controller that, given a set of controllable nodes, steers the system toward a desired configuration. We formulate analytical controllability guarantees for polynomial systems. For non-polynomial dynamics on hypergraphs, we propose a heuristic method for identifying controllable nodes and validate the proposed approach using Kuramoto oscillators.

eess.SY↗

Input Dexterity and Output Negotiation in Feedback-Linearizable Nonlinear Systems

We introduce a task-relative taxonomy of actuator inputs for nonlinear systems within the input-output feedback-linearization framework. Given a flat output specifying the task, inputs are classified as essential, redundant, or dexterity: essential inputs are required for exact linearization, redundant inputs can be removed without effect, and dexterity inputs can be deactivated while preserving exact linearization of a reduced task. We show that a subset is dexterity if and only if, under a suitable dynamic prolongation, it can appear as additional output channels (flat-input complement) on a common validity set. Whenever a family of systems obtained by (de)activating dexterity inputs admits a common prolongation, the family can be interpreted as a single prolonged system endowed with different output selections. This enables a unified linearizing controller that negotiates between full and reduced tasks without transients on shared outputs under compatibility and dwell-time conditions. Simulations on a fully actuated aerial platform illustrate graceful task downgrades from six-dimensional pose tracking as lateral-force channels are deactivated.

eess.SY↗

Input-to-state stabilization of linear systems under data-rate constraints

We study feedback stabilization of linear systems under data-rate constraints in the presence of completely unknown disturbances. A communication and control strategy is proposed based on sampled and quantized state measurements, where the quantization range is dynamically adjusted using reachable-set approximations and a disturbance estimate derived from quantization parameters. The strategy alternates between stabilizing and searching stages to recapture the state after escapes from the quantization range. Under a data-rate condition, it guarantees input-to-state stability (ISS) with respect to the disturbance. An additional quantization symbol is introduced to establish ISS near the equilibrium. A simulation example illustrates the effectiveness of the proposed approach.

eess.SY↗