TY - RPRT TI - One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-Attention AU - Arvind Mahankali AU - Tatsunori B. Hashimoto AU - Tengyu Ma PY - 2023 UR - https://arxiv.org/abs/2307.03576 ID - 2307.03576 ER -