SearcharxivSearch

arXiv subjects

Junho Cho

Publications and source records attributed to Junho Cho.

15 recordsLinked to original sources

RLDX-1 Technical Report

While Vision-Language-Action models (VLAs) have shown remarkable progress toward human-like generalist robotic policies through the versatile intelligence (i.e. broad scene understanding and language-conditioned generalization) inherited from pre-trained Vision-Language Models, they still struggle with complex real-world tasks requiring broader functional capabilities (e.g. motion awareness, long-term memory, and physical sensing). To address this, we introduce RLDX-1, a general-purpose robotic policy for dexterous manipulation built on the Multi-Stream Action Transformer (MSAT), an architecture that unifies these capabilities by integrating heterogeneous modalities through modality-specific streams with cross-modal joint self-attention. RLDX-1 further combines this architecture with system-level design choices, including data synthesis for rare manipulation scenarios, learning procedures specialized for human-like manipulation, and inference optimizations for real-time deployment. Through empirical evaluation, we show that RLDX-1 consistently outperforms recent frontier VLAs (e.g. $\pi_{0.5}$ and GR00T N1.6) across both simulation benchmarks and real-world tasks that require broad functional capabilities beyond general versatility. In particular, RLDX-1 shows superiority in ALLEX humanoid tasks by achieving success rates of 86.8% while $\pi_{0.5}$ and GR00T N1.6 achieve around 40%, highlighting the ability of RLDX-1 to control a high-DoF humanoid robot under diverse functional demands. Together, these results position RLDX-1 as a promising step toward reliable VLAs for complex, contact-rich, and dynamic real-world dexterous manipulation.

cs.RO

How Can Objects Help Video-Language Understanding?

Do we still need to represent objects explicitly in multimodal large language models (MLLMs)? To one extreme, pre-trained encoders convert images into visual tokens, with which objects and spatiotemporal relationships may be implicitly modeled. To the other extreme, image captions by themselves provide strong empirical performances for understanding tasks, despite missing fine-grained spatiotemporal information. To answer this question, we introduce ObjectMLLM, a framework capable of leveraging arbitrary computer vision algorithm to extract and integrate structured visual representation. Through extensive evaluations on six video question answering benchmarks, we confirm that explicit integration of object-centric representation remains necessary. Surprisingly, we observe that the simple approach of quantizing the continuous, structured object information and representing them as plain text performs the best, offering a data-efficient approach to integrate other visual perception modules into MLLM design. Our code and models are released at https://github.com/brown-palm/ObjectMLLM.

cs.CV

Font Representation Learning via Paired-glyph Matching

Fonts can convey profound meanings of words in various forms of glyphs. Without typography knowledge, manually selecting an appropriate font or designing a new font is a tedious and painful task. To allow users to explore vast font styles and create new font styles, font retrieval and font style transfer methods have been proposed. These tasks increase the need for learning high-quality font representations. Therefore, we propose a novel font representation learning scheme to embed font styles into the latent space. For the discriminative representation of a font from others, we propose a paired-glyph matching-based font representation learning model that attracts the representations of glyphs in the same font to one another, but pushes away those of other fonts. Through evaluations on font retrieval with query glyphs on new fonts, we show our font representation learning scheme achieves better generalization performance than the existing font representation learning techniques. Finally on the downstream font style transfer and generation tasks, we confirm the benefits of transfer learning with the proposed method. The source code is available at https://github.com/junhocho/paired-glyph-matching.

cs.CV

On Digital Subcarrier Multiplexing under A Bandwidth Limitation and ASE Noise

We show that digital subcarrier multiplexing (DSM) systems require much greater complexity for Nyquist pulse shaping than single-carrier (SC) systems, and it is a misconception that both systems use the same bandwidth when using the same pulse shaping. Through back-to-back (B2B) experiments with realistic transmitter (TX) modules and amplified spontaneous emission (ASE) noise loading, we show that even with optimized waterfilling and entropy loading, DSM does not achieve a larger net data rate (NDR) compared to SC when only ASE noise exists in the channel in long-haul transmission scenarios.

eess.SP

On the Kurtosis of Modulation Formats for Characterizing the Nonlinear Fiber Propagation

Knowing only two high-order statistical moments of modulation symbols, often represented by the fourth moment called "kurtosis", the overestimation of nonlinear interference (NLI) in a Gaussian noise (GN) model due to Gaussian signaling assumption can be corrected through an enhanced GN (EGN) model. However, in some modern optical communication systems where the transmitted modulation symbols are statistically correlated, such as in systems that use probabilistic constellation shaping (PCS) with finite-length sphere shaping, the kurtosis-based EGN model produces significant inaccuracies in analytical prediction of NLI. In this paper, we show that for correlated modulation symbols, the NLI can be more accurately estimated by substituting a statistical measure called windowed kurtosis into the EGN model, instead of the conventional kurtosis. Remarkably, the optimal window length for windowed kurtosis is found to be consistent with the self-phase modulation (SPM) and cross-phase modulation (XPM) characteristic times in various system configurations. The findings can be used in practice to analytically evaluate and design NLI-tolerant modulation formats.

eess.SP

Single-ended Coherent Receiver

Commercial coherent receivers utilize balanced photodetectors (PDs) with high single-port rejection ratio (SPRR) to mitigate the signal-signal beat interference (SSBI) due to the square-law detection process. As the symbol rates of coherent transponders are increased to 100 Gbaud and beyond, maintaining a high SPRR in a cost-effective manner becomes more and more challenging. One potential approach for solving this problem is to leverage the concept of single-ended coherent receiver (SER) where single-ended PDs are used instead of the balanced PDs. In this case, the resulting SSBI should be mitigated in the digital domain. In this paper, we show that SSBI can be effectively mitigated using various low-complexity techniques, such as the direct filed reconstruction (DFR), clipped iterative SSBI cancellation (CIC) and gradient decent (GD). In addition, we present a self-calibration technique for SERs which can be extended for characterizing the optical-to-electrical (O/E) response of a conventional balanced coherent receiver (BR). Using the developed techniques, we then experimentally demonstrate a 90 Gbaud probabilistically constellation shaped 64-QAM (PCS-64QAM) transmission using a SER, achieving a net data rate of 882 Gb/s over 100 km of standard single mode fiber (SSMF). The sensitivity penalty compared to the BR is below 0.5 dB. We expect that when the symbol rate is increased further, a SER can potentially outperform a BR, especially when applied to cost-sensitive commercial pluggable coherent transceivers

eess.SP

Unsupervised Hyperbolic Representation Learning via Message Passing Auto-Encoders

Most of the existing literature regarding hyperbolic embedding concentrate upon supervised learning, whereas the use of unsupervised hyperbolic embedding is less well explored. In this paper, we analyze how unsupervised tasks can benefit from learned representations in hyperbolic space. To explore how well the hierarchical structure of unlabeled data can be represented in hyperbolic spaces, we design a novel hyperbolic message passing auto-encoder whose overall auto-encoding is performed in hyperbolic space. The proposed model conducts auto-encoding the networks via fully utilizing hyperbolic geometry in message passing. Through extensive quantitative and qualitative analyses, we validate the properties and benefits of the unsupervised hyperbolic representations. Codes are available at https://github.com/junhocho/HGCAE.

cs.LG

Does Probabilistic Constellation Shaping Benefit IM-DD Systems without Optical Amplifiers?

Probabilistic constellation shaping (PCS) has been widely applied to amplified coherent optical transmissions owing to its shaping gain over the uniform signaling and fine-grained rate adaptation to the underlying fiber channel condition. These merits stimulate the study of applying PCS to short-reach applications dominated by intensity modulation (IM) direct detection (DD) systems. As commercial IM-DD systems typically do not employ optical amplification to save the cost and power consumption, they are no longer subject to an average power constraint (APC) but a peak power constraint (PPC), which poses unique challenges to take full advantages of PCS. This paper provides a comprehensive investigation of PCS in IM-DD systems without optical amplifiers. In particular, we reveal that if the transmitter enhances the peak-to-average power ratio of the signal, a PPC system can be partially or even fully converted to an APC system in which the classical PCS offers its merits. The findings are verified through an IM-DD experiment using 4- and 8-ary pulse amplitude modulations.

eess.SP

Supply-Power-Constrained Cable Capacity Maximization Using Multi-Layer Neural Networks

We experimentally solve the problem of maximizing capacity under a total supply power constraint in a massively parallel submarine cable context, i.e., for a spatially uncoupled system in which fiber Kerr nonlinearity is not a dominant limitation. By using multi-layer neural networks trained with extensive measurement data acquired from a 12-span 744-km optical fiber link as an accurate digital twin of the true optical system, we experimentally maximize fiber capacity with respect to the transmit signal's spectral power distribution based on a gradient-descent algorithm. By observing convergence to approximately the same maximum capacity and power distribution for almost arbitrary initial conditions, we conjecture that the capacity surface is a concave function of the transmit signal power distribution. We then demonstrate that eliminating gain flattening filters (GFFs) from the optical amplifiers results in substantial capacity gains per Watt of electrical supply power compared to a conventional system that contains GFFs.

eess.SP

Prefix-Free Code Distribution Matching for 5G New Radio

We use prefix-free code distribution matching (PCDM) for rate matching (RM) in some 5G New Radio (NR) deployment scenarios, realizing a wide range of information rates from 1.4 to 6.0 bit/symbol in fine granularity of 0.2 bit/symbol. We study the performance and implementation of the PCDM-based RM, in comparison with the low-density parity-check (LDPC)-based RM, as defined in the 5G NR standard. Simulations in the additive white Gaussian noise channel show that up to 2.16 dB gain in the signal-to-noise ratio can be obtained with the PCDM-based RM at a block error rate of 10-2 when compared to LDPC-based RM in the tested scenarios, potentially at a smaller hardware cost.

eess.SP

Prefix-Free Code Distribution Matching for Probabilistic Constellation Shaping

In this work, we construct energy-efficient variable-to-fixed length (V2F), fixed-to-variable length (F2V), and variable-to-variable length (V2V) prefix-free codes, which are optimal (or near-optimal) in the sense that no (or few) other codes with the size can achieve a smaller energy per code letter for the same entropy rate. Under stringent constraints of 4096 entries or below per codebook, the constructed codes yield an energy per code letter within a few tenths of a dB of the unconstrained theoretic lower bound, across a wide range of entropy rates with a very fine granularity. We also propose a framing method that allows variable-length codes to be transmitted using a fixed-length frame. The penalty caused by framing is studied using simulations and analysis, showing that the energy per code letter is kept within 0.2 dB of the unconstrained theoretic limit for some tested codes with a large frame length. When framed prefix-free codes are used to implement probabilistic constellation shaping (PCS) for communications in the additive white Gaussian noise channel, simulations show that 1.1 dB and 0.65 dB of shaping gains are achieved relative to uniform 8- and 16-quadrature amplitude modulation (QAM), respectively.

cs.IT