SearcharxivSearch

arXiv subjects

Sunghyun Lim

Publications and source records attributed to Sunghyun Lim.

2 recordsLinked to original sources

Deployable Human Preference Alignment in Robotics: Learning Representative Rewards from Diverse Human Preferences

Aligning robot policies with human preferences is essential for deployment to diverse end users. In per-user alignment approach, preference feedback is often sparse, so learning becomes unstable and vulnerable to human preference noise, and a growing number of individualized policies makes validation difficult before deployment. A single shared policy approach to user alignment avoids this cost but fails to capture heterogeneous preferences and often neglects minority preferences. To address these challenges, we introduce Preference-based REward Clustering (PREC), a novel framework that learns a compact set of policies from binary preference labels provided by diverse users. From a dataset of user trajectories and their preference labels, PREC first sets the labels aside and aggregates trajectories across users to learn a population-level shared trajectory encoder, alleviating limited per-user coverage and avoiding label noise during representation learning. Using this representation, PREC jointly assigns users to preference-coherent clusters and learns a representative reward model per cluster using preference labels, from which a policy is optimized for each cluster. Clustering similar users compensates for the limited number of labels available from each user and mitigates the effect of label noise. At the same time, maintaining a manageable number of reward models reduces the validation burden at deployment. Experiments across diverse simulated locomotion environments show that PREC groups users who label different trajectory subsets into preference-coherent clusters more accurately than baseline methods. Under sparse and noisy feedback, policies trained with PREC improve all three social welfare metrics over an existing single shared-policy user-alignment approach and even outperform per-user alignment approaches.

cs.RO

Quantum simulation approach to ultra-weak magnetic anisotropy in a frustrated spin-1/2 antiferromagnet

The intrinsic equivalence between electron spin and qubit offers a natural foundation for quantum simulations of magnetic materials. However, incorporating magnetocrystalline anisotropy (MCA), a key feature of real magnets, remains a major challenge. Here, we develop a quantum simulation framework for MCA in CuSb2O6, a spin-1/2 antiferromagnet with alternating ferromagnetic chains arising from frustrated, anisotropic exchange interactions in a nearly square lattice. The $\mathrm{Cu}^{2+}$ spin network is modeled as a four-qubit square lattice, with four paired ancilla qubits introduced to encode angle-dependent MCA. This two-qubit representation per spin site resolves the limitation that squared Pauli operators yield only the identity, enabling MCA terms to be faithfully embedded into quantum circuits. Using the variational quantum eigensolver, we determine an exceptionally small easy-axis MCA constant, just 0.00022% of the nearest-neighbor exchange interaction, yet sufficient to drive a spin-flop transition with $90^{\circ}$ spin reorientation and strong angular variation in magnetic torque. Beyond this regime, the simulations uncover a half-saturated magnetic phase at ultra-high fields, stabilized by anisotropic next-nearest-neighbor interactions. Our findings demonstrate the feasibility of resource-efficient quantum simulations of complex magnetic phenomena in real materials.

quant-ph