TY - RPRT TI - Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models AU - Martin Pawelczyk AU - Lillian Sun AU - Zhenting Qi AU - Aounon Kumar AU - Himabindu Lakkaraju PY - 2024 UR - https://arxiv.org/abs/2501.00418 ID - 2501.00418 ER -