SearcharxivSearch

arXiv subjects

Rongyi Lin

Publications and source records attributed to Rongyi Lin.

2 recordsLinked to original sources

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit criteria. Yet in GRPO-style pipelines, these structured judgments are reduced to a scalar response-level reward and converted into a response-level advantage, which is broadcast uniformly to all generated tokens. This leaves no explicit mechanism for allocating credit within a response, even when different criteria are grounded in different spans, formatting decisions, or semantic choices. We propose CoRT, a token-level credit weighting method for rubric-conditioned GRPO. Instead of training an auxiliary token scoring model, CoRT uses counterfactual replay to rescore the same sampled response under the original rubric-conditioned prompt and a matched criteria-free prompt. The resulting tokenwise log-likelihood contrasts serve as a proxy for dependence on the rubric context. CoRT maps these contrasts to bounded, response-normalized weights and uses them to redistribute the signed GRPO advantage across tokens, without introducing an auxiliary scorer or changing the response-level reward. Experiments across instruction-tuned models and reward granularities show that CoRT improves over matched response-level GRPO in the vast majority of comparisons, with an average gain of 4.4 percentage points. The method remains competitive with learned token-level credit baselines while avoiding a separate relevance-learning stage. These results suggest that policy-internal counterfactual likelihood contrasts provide an effective training signal for within-response credit allocation while retaining the simplicity and stability of GRPO.

cs.AI

Physically interpretable diffractive optical networks for high-dimensional vortex mode sorting

Despite the significant progress achieved by diffractive optical networks in diverse computing tasks, such as mode multiplexing and demultiplexing, investigations into the physical meanings behind complex diffractive networks at the layer level have been quite limited. Here, for highdimensional vortex mode sorting tasks, we show how various physical transformation rules for each layer within trained diffractive networks can be revealed under properly defined input/output mode relations. An intriguing physical transformation division phenomenon, associated with the saturated sorting performance of the system, has been observed with an increasing number of masks. In addition, we have also demonstrated the use of physical interpretation for efficiently designing parameter-varying networks with high performance. These physically interpretable optical networks resolve the contradiction between rigorous physical theorems and operationally vague network structures, paving the way for designing and understanding systems for various mode conversion tasks, and inspiring further interpretation of diffractive networks in advanced tasks and other network structures.

physics.optics