SearcharxivSearch

arXiv subjects

Sophia Zhou

Publications and source records attributed to Sophia Zhou.

2 recordsLinked to original sources

Rethinking Reward Supervision: Rubric-Conditioned Self-Distillation

Post-training of reasoning language models is commonly driven by supervised distillation and reinforcement learning with verifiable rewards. Distillation often relies on chain-of-thought annotations that are expensive to obtain and may themselves be noisy, incomplete, or partially incorrect; even when the final solution is correct, an imperfect rationale can interfere with learning. Reinforcement learning with verified rewards, on the other hand, typically compresses evaluative feedback into a scalar signal, obscuring which aspects of a response should be improved. We propose \textbf{Rubric-Conditioned Self-Distillation}, a framework that incorporates rubrics as structured, fine-grained feedback for on-policy self-distillation. Our method conditions the teacher model on criterion-level rubrics and uses it to provide token-level guidance on the student's own sampled trajectories. This design avoids treating a single reference rationale as the sole supervision target. Instead, rubrics specify what a strong response should satisfy, enabling more fine-grained credit assignment over the reasoning process than scalar reward optimization. We instantiate this framework with a two-stage pipeline that first learns to generate task-specific rubrics and then trains a rubric-guided reasoner. We evaluate on a diverse suite of science reasoning benchmarks and results show that rubric-conditioned self-distillation effectively converts rubric-level criteria into token-level guidance over the reasoning process, surpassing GRPO by 1.0 points and OPSD by 0.9 points on average.

cs.AI

Slope gap distribution of the double heptagon and an algorithm for determining winning vectors

In this paper, we study the distribution of renormalized gaps between slopes of saddle connections on translation surfaces. Specifically, we describe a procedure for finding the "winning holonomy vectors" as defined by Kumanduri-Sanchez-Wang in arXiv:2102.10069, which constitutes a key step in calculating the slope gap distribution for an arbitrary Veech surface. We then apply this method to explicitly compute the gap distribution for the regular double heptagon translation surface. This extends work of Athreya-Chaika-Lelievre in arXiv:1308.4203 on the gap distribution for the "golden L" translation surface, which is equivalent to the regular double pentagon surface.

math.DS