SearcharxivSearch

arXiv subjects

Junhyeong Park

Publications and source records attributed to Junhyeong Park.

8 recordsLinked to original sources

RLDX-1 Technical Report

While Vision-Language-Action models (VLAs) have shown remarkable progress toward human-like generalist robotic policies through the versatile intelligence (i.e. broad scene understanding and language-conditioned generalization) inherited from pre-trained Vision-Language Models, they still struggle with complex real-world tasks requiring broader functional capabilities (e.g. motion awareness, long-term memory, and physical sensing). To address this, we introduce RLDX-1, a general-purpose robotic policy for dexterous manipulation built on the Multi-Stream Action Transformer (MSAT), an architecture that unifies these capabilities by integrating heterogeneous modalities through modality-specific streams with cross-modal joint self-attention. RLDX-1 further combines this architecture with system-level design choices, including data synthesis for rare manipulation scenarios, learning procedures specialized for human-like manipulation, and inference optimizations for real-time deployment. Through empirical evaluation, we show that RLDX-1 consistently outperforms recent frontier VLAs (e.g. $π_{0.5}$ and GR00T N1.6) across both simulation benchmarks and real-world tasks that require broad functional capabilities beyond general versatility. In particular, RLDX-1 shows superiority in ALLEX humanoid tasks by achieving success rates of 86.8% while $π_{0.5}$ and GR00T N1.6 achieve around 40%, highlighting the ability of RLDX-1 to control a high-DoF humanoid robot under diverse functional demands. Together, these results position RLDX-1 as a promising step toward reliable VLAs for complex, contact-rich, and dynamic real-world dexterous manipulation.

cs.RO

LogoDiffuser: Training-Free Multilingual Logo Generation and Stylization via Letter-Aware Attention Control

Recent advances in text-to-image generation have been remarkable, but generating multilingual design logos that harmoniously integrate visual and textual elements remains a challenging task. Existing methods often distort character geometry when applying creative styles and struggle to support multilingual text generation without additional training. To address these challenges, we propose LogoDiffuser, a training-free method that synthesizes multilingual logo designs using the multimodal diffusion transformer. Instead of using textual prompts, we input the target characters as images, enabling robust character structure control regardless of language. We first analyze the joint attention mechanism to identify core tokens, which are tokens that strongly respond to textual structures. With this observation, our method integrates character structure and visual design by injecting the most informative attention maps. Furthermore, we perform layer-wise aggregation of attention maps to mitigate attention shifts across layers and obtain consistent core tokens. Extensive experiments and user studies demonstrate that our method achieves state-of-the-art performance in multilingual logo generation.

cs.CV

InfoCausalQA:Can Models Perform Non-explicit Causal Reasoning Based on Infographic?

Recent advances in Vision-Language Models (VLMs) have demonstrated impressive capabilities in perception and reasoning. However, the ability to perform causal inference -- a core aspect of human cognition -- remains underexplored, particularly in multimodal settings. In this study, we introduce InfoCausalQA, a novel benchmark designed to evaluate causal reasoning grounded in infographics that combine structured visual data with textual context. The benchmark comprises two tasks: Task 1 focuses on quantitative causal reasoning based on inferred numerical trends, while Task 2 targets semantic causal reasoning involving five types of causal relations: cause, effect, intervention, counterfactual, and temporal. We manually collected 494 infographic-text pairs from four public sources and used GPT-4o to generate 1,482 high-quality multiple-choice QA pairs. These questions were then carefully revised by humans to ensure they cannot be answered based on surface-level cues alone but instead require genuine visual grounding. Our experimental results reveal that current VLMs exhibit limited capability in computational reasoning and even more pronounced limitations in semantic causal reasoning. Their significantly lower performance compared to humans indicates a substantial gap in leveraging infographic-based information for causal inference. Through InfoCausalQA, we highlight the need for advancing the causal reasoning abilities of multimodal AI systems.

cs.CL

Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists?

Any-to-any generative models aim to enable seamless interpretation and generation across multiple modalities within a unified framework, yet their ability to preserve relationships across modalities remains uncertain. Do unified models truly achieve cross-modal coherence, or is this coherence merely perceived? To explore this, we introduce ACON, a dataset of 1,000 images (500 newly contributed) paired with captions, editing instructions, and Q&A pairs to evaluate cross-modal transfers rigorously. Using three consistency criteria-cyclic consistency, forward equivariance, and conjugated equivariance-our experiments reveal that any-to-any models do not consistently demonstrate greater cross-modal consistency than specialized models in pointwise evaluations such as cyclic consistency. However, equivariance evaluations uncover weak but observable consistency through structured analyses of the intermediate latent space enabled by multiple editing operations. We release our code and data at https://github.com/JiwanChung/ACON.

cs.CL

FMCW SAR with New Synthesis Method Based on A-SPC Technique

Frequency modulated continuous wave (FMCW) radar is emerging as a trendy radar system for synthetic aperture radar (SAR). This letter proposes a novel method for the extraction of the SAR image with the FMCW radar. The proposed method can improve the quality of the SAR image. For the verification, we built an automobile SAR (AutoSAR) system and conducted experiments to extract the SAR map by using the AutoSAR system. Then, we synthesized SAR images through both the conventional method and the proposed method to demonstrate the performance of the proposed method. The experimental results show that the SAR image has been successfully improved by the proposed method.

eess.SP

Small Drone Classification with Light CNN and New Micro-Doppler Signature Extraction Method Based on A-SPC Technique

As the threats of small drones increase, not only the detection but also the classification of small drones has become important. Many recent studies have applied an approach to utilize the micro-Doppler signature (MDS) for the small drone classification by using frequency modulated continuous wave (FMCW) radars. In this letter, we propose a novel method to extract the MDS images of the small drones with the FMCW radar. Moreover, we propose a light convolutional neural network (CNN) whose structure is straightforward, and the number of parameters is quite small for fast classification. The proposed method contributes to increasing the classification accuracy by improving the quality of MDS images. We classified the small drones with the MDS images extracted by the conventional method and the proposed method through the proposed CNN. The experimental results showed that the total classification accuracy was increased by 10.00 % due to the proposed method. The total classification accuracy was recorded at 97.14 % with the proposed MDS extraction method and the proposed light CNN.

eess.SP

Advanced Stationary Point Concentration Technique for Leakage Mitigation and Small Drone Detection with FMCW Radar

As the threats of small drones have grown, developing radars to detect the small drones has become an important issue. In earlier studies, we proposed the stationary point concentration (SPC) technique for the small drone detection with frequency-modulated continuous-wave (FMCW) radar. The SPC technique is a new approach to mitigate the leakage that is an inherent problem in the FMCW radar. The SPC technique improves the signal-to-noise ratio of the small drones by reducing the noise floor and provides accurate distance and velocity information of the small drones. However, the SPC technique has shortcomings in realizing it. In this paper, we present the drawbacks of the SPC technique clearly and propose an advanced SPC (A-SPC) technique. The A-SPC technique can overcome the drawbacks of the SPC technique while taking all the good effects of the SPC technique. The experimental results verify the proposed A-SPC technique and show its robustness and usefulness.

eess.SP

Leakage Mitigation and Internal Delay Compensation in FMCW Radar for Small Drone Detection

One of the notorious problems of frequency modulated continuous-wave (FMCW) radar is leakage between the transmitter and the receiver. The phase noise of the leakage is expressed as a skirt around the leakage signal on power spectrum. It causes the deterioration of the dynamic range, especially, in the near-distance region. Therefore, although FMCW radar has an advantage over pulse radar in terms of near-distance target detection due to its way of operation, the advantage of FMCW radar can be lost because of the leakage. Another problem of FMCW radar is internal delay in the radar system. It leads to the decrease of the maximum detectable range. In this paper, a novel down-conversion technique which resolves these problems is proposed. Detailed theory and procedures to implement the proposed technique are explained. Then, performances of it are verified with the experiment results. The proposed technique can be implemented through frequency planning and digital signal processing without additional parts. The results show that the proposed technique lowers the noise floor about 7.0 dB in the near-distance region and 2.1 dB even in the far-distance region. Also, the results demonstrate the proposed technique recover the reduced maximum detectable range by compensating the internal delay.

eess.SP