SearcharxivSearch

arXiv subjects

Youngin Kim

Publications and source records attributed to Youngin Kim.

4 recordsLinked to original sources

Identifiable Token Correspondence for World Models

Token-based transformer world models have shown strong performance in visual reinforcement learning, but often suffer from temporal inconsistency in long-horizon rollouts, including object duplication, disappearance, and transmutation. A key reason is that most existing approaches treat next-frame prediction purely as a token generation problem, without considering the persistence of tokens across time. We introduce Identifiable Token Correspondence (ITC), a decoding step for token-based transformer world models that formulates next-frame prediction as a structured assignment problem with latent token correspondence variables: each next-frame token is explained either by copying a token from the previous frame or by generating a new one. ITC leaves the transformer architecture and training procedure unchanged and can be added on top of existing backbones. Our experiments show state-of-the-art performance on 4 challenging benchmarks. The proposed method achieves a return of 72.5% and a score of 35.6% on the Craftax-classic benchmark, significantly surpassing the previous best of 67.4% and 27.9%. We release our source code on https://github.com/snu-mllab/Identifiable-Token-Correspondence.

cs.LG

QuadStretcher: A Forearm-Worn Skin Stretch Display for Bare-Hand Interaction in AR/VR

The paradigm of bare-hand interaction has become increasingly prevalent in Augmented Reality (AR) and Virtual Reality (VR) environments, propelled by advancements in hand tracking technology. However, a significant challenge arises in delivering haptic feedback to users' hands, due to the necessity for the hands to remain bare. In response to this challenge, recent research has proposed an indirect solution of providing haptic feedback to the forearm. In this work, we present QuadStretcher, a skin stretch display featuring four independently controlled stretching units surrounding the forearm. While achieving rich haptic expression, our device also eliminates the need for a grounding base on the forearm by using a pair of counteracting tactors, thereby reducing bulkiness. To assess the effectiveness of QuadStretcher in facilitating immersive bare-hand experiences, we conducted a comparative user evaluation (n = 20) with a baseline solution, Squeezer. The results confirmed that QuadStretcher outperformed Squeezer in terms of expressing force direction and heightening the sense of realism, particularly in 3-DoF VR interactions such as pulling a rubber band, hooking a fishing rod, and swinging a tennis racket. We further discuss the design insights gained from qualitative user interviews, presenting key takeaways for future forearm-haptic systems aimed at advancing AR/VR bare-hand experiences.

cs.HC

Electronic-Photonic Interface for Multiuser Optical Wireless Communication

We demonstrate an electronic-photonic (EP) interface for multiuser optical wireless communication (OWC), consisting of a multibeam optical phased array (MBOPA) along with co-integrated electro-optic (EO) modulators and high-speed CMOS drivers. The MBOPA leverages a path-length difference in the optical phased array (OPA) along with wavelength-division multiplexing technology for spatial carrier aggregation and multiplexing. To generate two and four pulsed amplitude modulation signals, and transmit them to multiple users, we employ an optical digital-to-analog converter technique by using two traveling-wave electrode Mach-Zehnder modulators, which are monolithically integrated with high-speed, wide-output-swing CMOS drivers. The MBOPA and monolithic EO modulator are implemented by silica wafer through planar lightwave circuit fabrication process and a 45-nm monolithic silicon photonics technology, respectively. We measured and analyzed two-channel parallel communication at a data rate of 54 Gbps per user over the wireless distance of 1 m. To the best of our knowledge, this is the first system level demonstration of the multi-user OWC using the in-house-designed photonic and monolithically integrated chips. Finally, we suggest best modulation format for different data rate and the number of multibeams, considering effects of the proposed OPA and the monolithic modulator.

physics.optics

AutoCycle-VC: Towards Bottleneck-Independent Zero-Shot Cross-Lingual Voice Conversion

This paper proposes a simple and robust zero-shot voice conversion system with a cycle structure and mel-spectrogram pre-processing. Previous works suffer from information loss and poor synthesis quality due to their reliance on a carefully designed bottleneck structure. Moreover, models relying solely on self-reconstruction loss struggled with reproducing different speakers' voices. To address these issues, we suggested a cycle-consistency loss that considers conversion back and forth between target and source speakers. Additionally, stacked random-shuffled mel-spectrograms and a label smoothing method are utilized during speaker encoder training to extract a time-independent global speaker representation from speech, which is the key to a zero-shot conversion. Our model outperforms existing state-of-the-art results in both subjective and objective evaluations. Furthermore, it facilitates cross-lingual voice conversions and enhances the quality of synthesized speech.

cs.SD