SearcharxivSearch

arXiv subjects

Yunyi Yang

Publications and source records attributed to Yunyi Yang.

At least 19 recordsLinked to original sources

Open Rubric System: Scaling Reinforcement Learning with Pairwise Adaptive Rubric

Scalar reward models compress multi-dimensional human preferences into a single opaque score, creating an information bottleneck that often leads to brittleness and reward hacking in open-ended alignment. We argue that robust alignment for non-verifiable tasks is fundamentally a principle generalization problem: reward should not be a learned function internalized into a judge, but an explicit reasoning process executed under inspectable principles. To operationalize this view, we present the Open Rubric System (OpenRS), a plug-and-play, rubrics-based LLM-as-a-Judge framework built around Pairwise Adaptive Meta-Rubrics (PAMR) and lightweight Pointwise Verifiable Rubrics (PVRs), which provide both hard-constraint guardrails and verifiable reward components when ground-truth or programmatic checks are available. OpenRS uses an explicit meta-rubric -- a constitution-like specification that governs how rubrics are instantiated, weighted, and enforced -- and instantiates adaptive rubrics on the fly by conditioning on the semantic differences between two candidate responses. It then performs criterion-wise pairwise comparisons and aggregates criterion-level preferences externally, avoiding pointwise weighted scalarization while improving discriminability in open-ended settings. To keep principles consistent yet editable across various domains, we introduce a two-level meta-rubric refinement pipeline (automated evolutionary refinement for general principles and a reproducible human-in-the-loop procedure for domain principles), complemented with pointwise verifiable rubrics that act as both guardrails against degenerate behaviors and a source of verifiable reward for objective sub-tasks. Finally, we instantiate OpenRS as reward supervision in pairwise RL training.

cs.CL

Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards

Reinforcement learning with verifiable rewards (RLVR) has enabled large language models (LLMs) to achieve remarkable breakthroughs in reasoning tasks with objective ground-truth answers, such as mathematics and code generation. However, a significant gap remains for non-verifiable tasks, like creative writing and open-ended dialogue, where quality assessment is inherently subjective and lacks definitive references. Existing approaches for these domains often rely on scalar reward models trained with human preferences, which suffer from limited generalization and are prone to reward hacking, such as over-explanation and length bias. In this work, we propose a unified RLVR-based training paradigm that bridges the gap between non-verifiable tasks and verifiable rewards. We introduce a writing-principle-based pairwise Generative Reward Model (GenRM) and a novel Bootstrapped Relative Policy Optimization (BRPO) algorithm. The pairwise writing GenRM leverages self-principled critique to transform subjective assessments into reliable, verifiable rewards, while BRPO enables dynamic, reference-free pairwise comparison by leveraging a bootstrapped response as temporary reference from within group rollouts during RL training. Our approach empowers LLMs to develop robust writing capabilities without supervised fine-tuning, as demonstrated by Writing-Zero, which shows consistent improvement and strong resistance to reward hacking compared to scalar reward baselines. Furthermore, our method achieves competitive results on both in-house and open-source writing benchmarks. Our findings suggest the potential to unify rule-based, reference-based, and reference-free reward modeling under the RLVR framework, thus paving the way for a comprehensive and scalable RL training paradigm applicable across all language tasks.

cs.CL

Waveguide optical parametric amplifiers in silicon nitride with 2D graphene oxide films

Optical parametric amplification (OPA) represents a powerful solution to achieve broadband amplification in wavelength ranges beyond the scope of conventional gain media, for generating high-power optical pulses, optical microcombs, entangled photon pairs and a wide range of other applications. Here, we demonstrate optical parametric amplifiers based on silicon nitride (Si3N4) waveguides integrated with two-dimensional (2D) layered graphene oxide (GO) films. We achieve precise control over the thickness, length, and position of the GO films using a transfer-free, layer-by-layer coating method combined with accurate window opening in the chip cladding using photolithography. Detailed OPA measurements with a pulsed pump for the fabricated devices with different GO film thicknesses and lengths show a maximum parametric gain of ~24.0 dB, representing a ~12.2 dB improvement relative to the device without GO. We perform a theoretical analysis of the device performance, achieving good agreement with experiment and showing that there is substantial room for further improvement. This work represents the first demonstration of integrating 2D materials on chips to enhance the OPA performance, providing a new way of achieving high performance photonic integrated OPA by incorporating 2D materials.

physics.optics

UBARv2: Towards Mitigating Exposure Bias in Task-Oriented Dialogs

This paper studies the exposure bias problem in task-oriented dialog systems, where the model's generated content over multiple turns drives the dialog context away from the ground-truth distribution at training time, introducing error propagation and damaging the robustness of the TOD system. To bridge the gap between training and inference for multi-turn task-oriented dialogs, we propose session-level sampling which explicitly exposes the model to sampled generated content of dialog context during training. Additionally, we employ a dropout-based consistency regularization with the masking strategy R-Mask to further improve the robustness and performance of the model. The proposed UBARv2 achieves state-of-the-art performance on the standardized evaluation benchmark MultiWOZ and extensive experiments show the effectiveness of the proposed methods.

cs.CL

Enhanced self-phase modulation in silicon nitride waveguides integrated with 2D graphene oxide films

We experimentally demonstrate enhanced self-phase modulation (SPM) in silicon nitride (Si3N4) waveguides integrated with 2D graphene oxide (GO) films. GO films are integrated onto Si3N4 waveguides using a solution-based, transfer-free coating method that enables precise control of the film thickness. Detailed SPM measurements are carried out using both picosecond and femtosecond optical pulses. Owing to the high Kerr nonlinearity of GO, the hybrid waveguides show significantly improved spectral broadening compared to the uncoated waveguide, achieving a broadening factor of up to ~3.4 for a device with 2 layers of GO. By fitting the experimental results with theory, we obtain an improvement in the waveguide nonlinear parameter by a factor of up to 18.4 and a Kerr coefficient (n2) of GO that is about 5 orders of magnitude higher than Si3N4. Finally, we provide a theoretical analysis for the influence of GO film length, coating position, and its saturable absorption on the SPM performance. These results verify the effectiveness of on-chip integrating 2D GO films to enhance the nonlinear optical performance of Si3N4 devices.

physics.optics

Enhanced spectral broadening via self-phase modulation with femtosecond optical pulses in silicon nanowires integrated with 2D graphene oxide films

We experimentally demonstrate enhanced spectral broadening of femtosecond optical pulses af-ter propagation through silicon-on-insulator (SOI) nanowire waveguides integrated with two-dimensional (2D) graphene oxide (GO) films. Owing to the strong mode overlap between the SOI nanowires and the GO films with a high Kerr nonlinearity, the self-phase modulation (SPM) process in the hybrid waveguides is significantly enhanced, resulting in greatly improved spectral broadening of the femtosecond optical pulses. A solution-based, transfer-free coating method is used to integrate GO films onto the SOI nanowires with precise control of the film thickness. Detailed SPM measurements using femtosecond optical pulses are carried out, achieving a broadening factor of up to ~4.3 for a device with 0.4-mm-long, 2 layers of GO. By fit-ting the experimental results with theory, we obtain an improvement in the waveguide nonlin-ear parameter by a factor of ~3.5 and the effective nonlinear figure of merit (FOM) by a factor of ~3.8, relative to the uncoated waveguide. Finally, we discuss the influence of GO film length on the spectral broadening and compare the nonlinear optical performance of different integrated waveguides coated with GO films. These results confirm the improved nonlinear optical per-formance for silicon devices integrated with 2D GO films.

physics.optics

Towards Building an Open-Domain Dialogue System Incorporated with Internet Memes

In recent years, Internet memes have been widely used in online chatting. Compared with text-based communication, conversations become more expressive and attractive when Internet memes are incorporated. This paper presents our solutions for the Meme incorporated Open-domain Dialogue (MOD) Challenge of DSTC10, where three tasks are involved: text response modeling, meme retrieval, and meme emotion classification. Firstly, we leverage a large-scale pre-trained dialogue model for coherent and informative response generation. Secondly, based on interaction-based text-matching, our approach can retrieve appropriate memes with good generalization ability. Thirdly, we propose to model the emotion flow (EF) in conversations and introduce an auxiliary task of emotion description prediction (EDP) to boost the performance of meme emotion classification. Experimental results on the MOD dataset demonstrate that our methods can incorporate Internet memes into dialogue systems effectively.

cs.CL

Amendable Generation for Dialogue State Tracking

In task-oriented dialogue systems, recent dialogue state tracking methods tend to perform one-pass generation of the dialogue state based on the previous dialogue state. The mistakes of these models made at the current turn are prone to be carried over to the next turn, causing error propagation. In this paper, we propose a novel Amendable Generation for Dialogue State Tracking (AG-DST), which contains a two-pass generation process: (1) generating a primitive dialogue state based on the dialogue of the current turn and the previous dialogue state, and (2) amending the primitive dialogue state from the first pass. With the additional amending generation pass, our model is tasked to learn more robust dialogue state tracking by amending the errors that still exist in the primitive dialogue state, which plays the role of reviser in the double-checking process and alleviates unnecessary error propagation. Experimental results show that AG-DST significantly outperforms previous works in two active DST datasets (MultiWOZ 2.2 and WOZ 2.0), achieving new state-of-the-art performances.

cs.CL

Directed Acyclic Graph Network for Conversational Emotion Recognition

The modeling of conversational context plays a vital role in emotion recognition from conversation (ERC). In this paper, we put forward a novel idea of encoding the utterances with a directed acyclic graph (DAG) to better model the intrinsic structure within a conversation, and design a directed acyclic neural network, namely DAG-ERC, to implement this idea. In an attempt to combine the strengths of conventional graph-based neural models and recurrence-based neural models, DAG-ERC provides a more intuitive way to model the information flow between long-distance conversation background and nearby context. Extensive experiments are conducted on four ERC benchmarks with state-of-the-art models employed as baselines for comparison. The empirical results demonstrate the superiority of this new model and confirm the motivation of the directed acyclic graph architecture for ERC.

cs.CL

Retrieve & Memorize: Dialog Policy Learning with Multi-Action Memory

Dialogue policy learning, a subtask that determines the content of system response generation and then the degree of task completion, is essential for task-oriented dialogue systems. However, the unbalanced distribution of system actions in dialogue datasets often causes difficulty in learning to generate desired actions and responses. In this paper, we propose a retrieve-and-memorize framework to enhance the learning of system actions. Specially, we first design a neural context-aware retrieval module to retrieve multiple candidate system actions from the training set given a dialogue context. Then, we propose a memory-augmented multi-decoder network to generate the system actions conditioned on the candidate actions, which allows the network to adaptively select key information in the candidate actions and ignore noises. We conduct experiments on the large-scale multi-domain task-oriented dialogue dataset MultiWOZ 2.0 and MultiWOZ 2.1. Experimental results show that our method achieves competitive performance among several state-of-the-art models in the context-to-response generation task.

cs.CL

UBAR: Towards Fully End-to-End Task-Oriented Dialog Systems with GPT-2

This paper presents our task-oriented dialog system UBAR which models task-oriented dialogs on a dialog session level. Specifically, UBAR is acquired by fine-tuning the large pre-trained unidirectional language model GPT-2 on the sequence of the entire dialog session which is composed of user utterance, belief state, database result, system act, and system response of every dialog turn. Additionally, UBAR is evaluated in a more realistic setting, where its dialog context has access to user utterances and all content it generated such as belief states, system acts, and system responses. Experimental results on the MultiWOZ datasets show that UBAR achieves state-of-the-art performances in multiple settings, improving the combined score of response generation, policy optimization, and end-to-end modeling by 4.7, 3.5, and 9.4 points respectively. Thorough analyses demonstrate that the session-level training sequence formulation and the generated dialog context are essential for UBAR to operate as a fully end-to-end task-oriented dialog system in real life. We also examine the transfer ability of UBAR to new domains with limited data and provide visualization and a case study to illustrate the advantages of UBAR in modeling on a dialog session level.

cs.CL

High performance integrated polarizers achieved by incorporating 2D layered graphene oxide films

Polarizers and polarization selective resonant cavities (e.g., ring resonators, gratings), are key components for applications to photography, coherent optical detection, polarization-division-multiplexing, optical sensing and liquid crystal displays. We demonstrate waveguide polarizers and polarization discriminating micro-ring resonators (MRRs) by integrating them with 2D graphene oxide (GO) layered thin films. We achieve precise control of the thickness, placement, and size of the films integrated onto photonic devices with a solution based, layer-by-layer transfer-free coating method combined with photolithography and lift-off. This overcomes limitations of layer transfer methods for 2D materials and is a significant advance to manufacturing integrated photonic devices incorporated with 2D materials. We measure the waveguide polarizer for different film thicknesses and lengths versus wavelength, polarization, and power, measuring a high polarization dependent loss (PDL) of ~ 53.8 dB. For GO-coated MRRs, we achieve an extinction ratio difference for TE/TM polarizations of 8.3-dB. We also present measurements of the linear optical properties of 2D layered GO films that yield the material loss anisotropy of the GO films and relative contribution of film loss anisotropy versus polarization-dependent mode overlap. Our results offer interesting physical insights into the transition of the layered GO films from 2D behaviour to quasi bulk like behavior and confirm the high performance of GO based integrated polarization selective devices.

physics.optics

BiOBr 2D materials for integrated nonlinear photonics devices

As a new group of advanced 2D layered materials, bismuth oxyhalides, i.e., BiOX (X = Cl, Br, I), have recently become of great interest. In this work, we characterize the third-order optical nonlinearities of BiOBr, an important member of the BiOX family. The nonlinear absorption and Kerr nonlinearity of BiOBr nanoflakes at both 800 nm and 1550 nm are characterized via the Z-Scan technique. Experimental results show that BiOBr nanoflakes exhibit a large nonlinear absorption coefficient = \b{eta} = 10-7 m/W as well as a large Kerr coefficient n2 = 10-14 m2/W. We also note that the n2 of BiOBr reverses sign from negative to positive as the wavelength is changed from 800 nm to 1550 nm. We further characterize the thickness-dependent nonlinear optical properties of BiOBr nanoflakes, finding that the magnitudes of \b{eta} and n2 increase with decreasing thickness of the BiOBr nanoflakes. Finally, we integrate BiOBr nanoflakes into silicon integrated waveguides and measure their insertion loss, with the extracted waveguide propagation loss showing good agreement with mode simulations based on ellipsometry measurements. These results confirm the strong potential of BiOBr as a promising nonlinear optical material for high-performance hybrid integrated photonic devices.

physics.optics

Enhanced four-wave-mixing and 3rd order optical nonlinearity in SiN nanowires integrated with graphene oxide films

Layered 2D graphene oxide (GO) films are integrated with silicon nitride (SiN) waveguides to experimentally demonstrate an enhanced Kerr nonlinearity via four wave mixing (FWM). Owing to the strong light matter interaction between the SiN waveguides and the highly nonlinear GO films, the FWM performance of the hybrid waveguides is significantly improved. SiN waveguides with both uniformly coated and patterned GO films are fabricated based on a transfer free, layer by layer GO coating method together with standard photolithography and lift off processes, yielding precise control of the film thickness, placement and coating length. Detailed FWM measurements are carried out for the fabricated devices with different numbers of GO layers and at different pump powers. By optimizing the trade off between the nonlinearity and loss, we obtain a significant improvement in the FWM conversion efficiency of 7.3 dB for a uniformly coated device with 1 layer of GO and 9.1 dB for a patterned device with 5 layers of GO. We also obtain a significant increase in FWM bandwidth for the patterned devices. A detailed analysis of the influence of pattern length and position on the FWM performance is performed. Based on the FWM measurements, the dependence of GOs third-order nonlinearity on layer number and pump power is also extracted, revealing interesting physical insights about the 2D layered GO films. Finally, we obtain an enhancement in the effective nonlinear parameter of the hybrid waveguides by over a factor of 100. These results verify the enhanced nonlinear optical performance of SiN waveguides achievable by incorporating 2D layered GO films.

physics.optics

Relational Graph Attention Network for Aspect-based Sentiment Analysis

Aspect-based sentiment analysis aims to determine the sentiment polarity towards a specific aspect in online reviews. Most recent efforts adopt attention-based neural network models to implicitly connect aspects with opinion words. However, due to the complexity of language and the existence of multiple aspects in a single sentence, these models often confuse the connections. In this paper, we address this problem by means of effective encoding of syntax information. Firstly, we define a unified aspect-oriented dependency tree structure rooted at a target aspect by reshaping and pruning an ordinary dependency parse tree. Then, we propose a relational graph attention network (R-GAT) to encode the new tree structure for sentiment prediction. Extensive experiments are conducted on the SemEval 2014 and Twitter datasets, and the experimental results confirm that the connections between aspects and opinion words can be better established with our approach, and the performance of the graph attention network (GAT) is significantly improved as a consequence.

cs.CL

Enhanced nonlinear optical figure-of-merit at 1550nm for silicon nanowires integrated with graphene oxide layered films

Layered 2D GO films are integrated with silicon on insulator (SOI) nanowire waveguides to experimentally demonstrate an enhanced Kerr nonlinearity, observed through selfphase modulation (SPM). The GO films are integrated with SOI nanowires using a large area, transfer free, layer by layer coating method that yields precise control of the film thickness. The film placement and coating length are controlled by opening windows in the silica cladding of the SOI nanowires. Owing to the strong mode overlap between the SOI nanowires and the highly nonlinear GO films, the Kerr nonlinearity of the hybrid waveguides is significantly enhanced. Detailed SPM measurements using picosecond optical pulses show significant spectral broadening enhancement for SOI nanowires coated with 2.2 mm long films of 1 to 3 layers of GO, and 0.4 mm long films with 5 to 20 layers of GO. By fitting the experimental results with theory, the dependence of the n2 for GO on layer number and pulse energy is obtained, showing interesting physical insights and trends of the layered GO films from 2D monolayers to quasi bulk like behavior. Finally, we show that by coating SOI nanowires with GO films the effective nonlinear parameter of SOI nanowires is increased 16 times, with the effective nonlinear figure of merit (FOM) increasing by about 20 times to greater than 5. These results reveal the strong potential of using layered GO films to improve the Kerr nonlinear optical performance of silicon photonic devices.

physics.optics

Enhanced four-wave-mixing with 2D layered graphene oxide films integrated with CMOS compatible micro-ring resonators

Layered 2D graphene oxide (GO) films are integrated with microring resonators (MRRs) to experimentally demonstrate enhanced nonlinear optics in the form of four wave mixing (FWM). Both uniformly coated and patterned GO films are integrated on CMOS compatible doped silica MRRs using a large area, transfer free, layer by layer GO coating method together with photolithography and lift off processes, yielding precise control of the film thickness, placement, and coating length. The high Kerr nonlinearity and low loss of the GO films combined with the strong light matter interaction within the MRRs results in a significant improvement in the FWM efficiency in the hybrid MRRs. Detailed FWM measurements are performed at different pump powers and resonant wavelengths for the uniformly coated MRRs with 1 to 5 layers of GO as well as the patterned devices with 10 to 50 layers of GO. The experimental results show good agreement with theory, achieving up to 7.6 dB enhancement in the FWM conversion efficiency (CE) for an MRR uniformly coated with 1 layer of GO and 10.3 dB for a patterned device with 50 layers of GO. By fitting the measured CE as a function of pump power for devices with different numbers of GO layers, we also extract the dependence of the third-order nonlinearity on layer number and pump power, revealing interesting physical insights about the evolution of the layered GO films from 2D monolayers to quasi bulk like behavior. These results confirm the high nonlinear optical performance of integrated photonic resonators incorporated with 2D layered GO films.

physics.optics

High third-order Kerr optical nonlinearity in BiOBr 2D films measured by the Z-scan method

We investigate the nonlinear optical properties of BiOBr nanoflakes, a novel two-dimensional (2D) layered material from the Bismuth oxyhalide family. We measure the nonlinear absorption and Kerr nonlinearity of BiOBr nanoflakes at both 800 nm and 1550 nm via the Z scan technique. We observe a large nonlinear absorption coefficient beta = 10^-7 m/W as well as a large Kerr coefficient n2 = 10^-14 m^2/W. We also observe strong dispersion in n2, with it reversing sign from negative to positive as the wavelength varies from 800 nm to 1550 nm. In addition, we characterize the thickness-dependence of the nonlinear optical properties of BiOBr nanoflakes, observing that both the magnitudes of beta and n2 increase for very thin flakes. Finally, we integrate BiOBr nanoflakes onto silicon integrated waveguides and characterize the linear optical properties of the resulting hybrid integrated devices, with the measurements agreeing with calculated parameters using independent ellipsometry measurements. These results verify the strong potential of BiOBr as an advanced nonlinear optical material for high-performance hybrid integrated photonic devices.

physics.optics