SearcharxivSearch

arXiv subjects

Jinmei Liu

Publications and source records attributed to Jinmei Liu.

7 recordsLinked to original sources

Beyond the Dirac Delta: Mitigating Diversity Collapse in Reinforcement Fine-Tuning for Versatile Image Generation

Reinforcement learning (RL) has emerged as a powerful paradigm for fine-tuning large-scale generative models, such as diffusion and flow models, to align with complex human preferences and user-specified tasks. A fundamental limitation remains \textit{the curse of diversity collapse}, where the objective formulation and optimization landscape inherently collapse the policy to a Dirac delta distribution. To address this challenge, we propose \textbf{DRIFT} (\textbf{D}ive\textbf{R}sity-\textbf{I}ncentivized Reinforcement \textbf{F}ine-\textbf{T}uning for Versatile Image Generation), an innovative framework that systematically incentivizes output diversity throughout the on-policy fine-tuning process, reconciling strong task alignment with high generation diversity to enhance versatility essential for applications that demand diverse candidate generations. We approach the problem across three representative perspectives: i) \textbf{sampling} a reward-concentrated subset that filters out reward outliers to prevent premature collapse; ii) \textbf{prompting} with stochastic variations to expand the conditioning space, and iii) \textbf{optimization} of the intra-group diversity with a potential-based reward shaping mechanism. Experimental results show that DRIFT achieves superior Pareto dominance regarding task alignment and generation diversity, yielding a $ 9.08\%\!\sim\! 43.46\%$ increase in diversity at equivalent alignment levels and a $ 59.65\% \!\sim\! 65.86\%$ increase in alignment at equivalent levels of diversity.

cs.LG

Scalable In-Context Q-Learning

Recent advancements in language models have demonstrated remarkable in-context learning abilities, prompting the exploration of in-context reinforcement learning (ICRL) to extend the promise to decision domains. Due to involving more complex dynamics and temporal correlations, existing ICRL approaches may face challenges in learning from suboptimal trajectories and achieving precise in-context inference. In the paper, we propose \textbf{S}calable \textbf{I}n-\textbf{C}ontext \textbf{Q}-\textbf{L}earning (\textbf{S-ICQL}), an innovative framework that harnesses dynamic programming and world modeling to steer ICRL toward efficient reward maximization and task generalization, while retaining the scalability and stability of supervised pretraining. We design a prompt-based multi-head transformer architecture that simultaneously predicts optimal policies and in-context value functions using separate heads. We pretrain a generalized world model to capture task-relevant information, enabling the construction of a compact prompt that facilitates fast and precise in-context inference. During training, we perform iterative policy improvement by fitting a state value function to an upper-expectile of the Q-function, and distill the in-context value functions into policy extraction using advantage-weighted regression. Extensive experiments across a range of discrete and continuous environments show consistent performance gains over various types of baselines, especially when learning from suboptimal data. Our code is available at \textcolor{magenta}{\href{https://github.com/NJU-RL/SICQL}{https://github.com/NJU-RL/SICQL}}.

cs.AI

Continual Offline Reinforcement Learning via Diffusion-based Dual Generative Replay

We study continual offline reinforcement learning, a practical paradigm that facilitates forward transfer and mitigates catastrophic forgetting to tackle sequential offline tasks. We propose a dual generative replay framework that retains previous knowledge by concurrent replay of generated pseudo-data. First, we decouple the continual learning policy into a diffusion-based generative behavior model and a multi-head action evaluation model, allowing the policy to inherit distributional expressivity for encompassing a progressive range of diverse behaviors. Second, we train a task-conditioned diffusion model to mimic state distributions of past tasks. Generated states are paired with corresponding responses from the behavior generator to represent old tasks with high-fidelity replayed samples. Finally, by interleaving pseudo samples with real ones of the new task, we continually update the state and behavior generators to model progressively diverse behaviors, and regularize the multi-head critic via behavior cloning to mitigate forgetting. Experiments demonstrate that our method achieves better forward transfer with less forgetting, and closely approximates the results of using previous ground-truth data due to its high-fidelity replay of the sample space. Our code is available at \href{https://github.com/NJU-RL/CuGRO}{https://github.com/NJU-RL/CuGRO}.

cs.LG

Efficient Bayesian Policy Reuse with a Scalable Observation Model in Deep Reinforcement Learning

Bayesian policy reuse (BPR) is a general policy transfer framework for selecting a source policy from an offline library by inferring the task belief based on some observation signals and a trained observation model. In this paper, we propose an improved BPR method to achieve more efficient policy transfer in deep reinforcement learning (DRL). First, most BPR algorithms use the episodic return as the observation signal that contains limited information and cannot be obtained until the end of an episode. Instead, we employ the state transition sample, which is informative and instantaneous, as the observation signal for faster and more accurate task inference. Second, BPR algorithms usually require numerous samples to estimate the probability distribution of the tabular-based observation model, which may be expensive and even infeasible to learn and maintain, especially when using the state transition sample as the signal. Hence, we propose a scalable observation model based on fitting state transition functions of source tasks from only a small number of samples, which can generalize to any signals observed in the target task. Moreover, we extend the offline-mode BPR to the continual learning setting by expanding the scalable observation model in a plug-and-play fashion, which can avoid negative transfer when faced with new unknown tasks. Experimental results show that our method can consistently facilitate faster and more efficient policy transfer.

cs.LG

Fast-light Assisted Four-Wave-Mixing in Photonic Bandgap

Since the forward and backward waves are coupled with each other and a standing wave with no net propagation of energy is formed in the photonic bandgap, it is a commonsense of basic physics that, any kinds of effects associated with wave propagation including four-wave-mixing (FWM) are thought to be impossible. However, we lay great emphasis here on explaining that this commonsense could be broken under specific circumstances. In this article, we report with the first experimental observation of the energy conversion in the photonic bandgap into other channel via FWM. Owing to the phase manipulation by fast light effect in the photonic bandgap, we manage to achieve the phase-match condition and thus occurred FWM transfer energy into other channels outside the photonic bandgap efficiently. As one-dimensional photonic crystal, simulations on fiber Bragg grating (FBG) with and without fast light were conducted respectively, and an enhanced FWM in photonic bandgap of FBG was observed. The experimental result shows great agreement with the analysis.

physics.optics

Watt-level ultrahigh-OSNR single-longitudinal-mode tunable Brillouin fiber laser

A watt-level ultrahigh optical signal-to-noise ratio (OSNR) single-longitudinal-mode (SLM) tunable Brillouin fiber laser (BFL) has been demonstrated.By optimizing the length of the single mode fiber (SMF) cavity at 11m and its output ratio at sixty percent, 1.04 W output power, as well as stable SLM operation is obtained at 2.24 W pump power. The single pass cavity BFL has the advantage that Brillouin pump frequency doesn't need to match the cavity mode, thus the stability is greatly improved. As only SMF is used in the cavity, the operate wavelength can be tunable without the restriction from self lasing cavity mode. Furthermore, it proves that core-pumped single frequency fiber laser is able to generate watt-level power. The laser has excellent performance in terms of noise, linewidth, and stability.

physics.optics

Superluminal signal conversion in stimulated Brillouin scattering via an optical fiber ring resonator

We report the superluminal phenomenon of both Stokes and pump light in stimulated Brillouin scattering (SBS) via an optical-fiber ring lasing resonator. In our experiment, the superluminal generation of Stokes light firstly delayed with 33.79 ns and then propagated with the advancement of 93.34 ns within an 8-m single mode fiber (SMF), when the pump power increased from the SBS threshold to a higher power. This proves the first evidence that the optical interaction is determined by the group velocity even at the negative group velocity superluminal propagation. It is important that, this implies the possibility of superluminal information interchange because the results also indicate that the signal conversion between different wavelengths can be realized at the negative group velocity propagation.

physics.optics