SearcharxivSearch

arXiv subjects

Yunjian Xu

Publications and source records attributed to Yunjian Xu.

14 recordsLinked to original sources

DASH-OPD: Discrepancy-Aware Switching with Hysteresis for On-Policy Distillation

While on-policy distillation (OPD) reduces exposure bias by training student language models on their own rollouts, early student errors in long-horizon agentic scenarios can lead to contexts unfamiliar to the teacher. To improve trajectory quality, recent work on agentic OPD introduces teacher intervention into training rollouts by switching the executor between the student and the teacher. However, existing methods determine how much teacher intervention is needed---but not when. To address this limitation, we propose DASH-OPD (Discrepancy-Aware Switching with Hysteresis for OPD), the first agentic OPD method to perform adaptive, bidirectional executor switching. At each turn, DASH-OPD measures teacher--student discrepancy using a mean log-probability ratio over action tokens. Student-to-teacher ratios on student turns serve as drift signals, while teacher-to-student ratios on teacher turns serve as recovery signals. These signals are accumulated over multiple turns to form drift and recovery evidence, respectively. DASH-OPD switches executors when either type of evidence exceeds its corresponding switching threshold, introducing hysteresis that prevents rapid switching triggered by transient discrepancy fluctuations. Across three benchmarks and two student model sizes, DASH-OPD outperforms five baselines in all 14 task performance comparisons, while requiring the fewest interaction turns in nine of ten efficiency comparisons. Code, models, and training logs are available at https://github.com/Lucian1115/DASH-OPD

cs.LG

ADG: Ambient Diffusion-Guided Dataset Recovery for Corruption-Robust Offline Reinforcement Learning

Real-world datasets collected from sensors or human inputs are prone to noise and errors, posing significant challenges for applying offline reinforcement learning (RL). While existing methods have made progress in addressing corrupted actions and rewards, they remain insufficient for handling corruption in high-dimensional state spaces and for cases where multiple elements in the dataset are corrupted simultaneously. Diffusion models, known for their strong denoising capabilities, offer a promising direction for this problem-but their tendency to overfit noisy samples limits their direct applicability. To overcome this, we propose Ambient Diffusion-Guided Dataset Recovery (ADG), a novel approach that pioneers the use of diffusion models to tackle data corruption in offline RL. First, we introduce Ambient Denoising Diffusion Probabilistic Models (DDPM) from approximated distributions, which enable learning on partially corrupted datasets with theoretical guarantees. Second, we use the noise-prediction property of Ambient DDPM to distinguish between clean and corrupted data, and then use the clean subset to train a standard DDPM. Third, we employ the trained standard DDPM to refine the previously identified corrupted data, enhancing data quality for subsequent offline RL training. A notable strength of ADG is its versatility-it can be seamlessly integrated with any offline RL algorithm. Experiments on a range of benchmarks, including MuJoCo, Kitchen, and Adroit, demonstrate that ADG effectively mitigates the impact of corrupted data and improves the robustness of offline RL under various noise settings, achieving state-of-the-art results.

cs.LG

GAS: Generative Auto-bidding with Post-training Search

Auto-bidding is essential in facilitating online advertising by automatically placing bids on behalf of advertisers. Generative auto-bidding, which generates bids based on an adjustable condition using models like transformers and diffusers, has recently emerged as a new trend due to its potential to learn optimal strategies directly from data and adjust flexibly to preferences. However, generative models suffer from low-quality data leading to a mismatch between the condition, like return to go, and true action value, especially in long sequential decision-making. Besides, the majority preference in the dataset may hinder models' generalization ability on minority advertisers' preferences. While it is possible to collect high-quality data and retrain multiple models for different preferences, the high cost makes it unaffordable, hindering the advancement of auto-bidding into the era of large foundation models. To address this, we propose a flexible and practical Generative Auto-bidding scheme using post-training Search, termed GAS, to refine a base policy model's output and adapt to various preferences. We use weak-to-strong search alignment by training small critics for different preferences and an MCTS-inspired search to refine the model's output. Specifically, a novel voting mechanism with transformer-based critics trained with policy indications could enhance search alignment performance. Additionally, utilizing the search, we provide a fine-tuning method for high-frequency preference scenarios considering computational efficiency. Extensive experiments conducted on the real-world dataset and online A/B test on the Kuaishou advertising platform demonstrate the effectiveness of GAS, achieving significant improvements, e.g., 4.60% increment of target cost.

cs.AI

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs

Reinforcement learning (RL) has become a cornerstone for enhancing the reasoning capabilities of large language models (LLMs), with recent innovations such as Group Relative Policy Optimization (GRPO) demonstrating exceptional effectiveness. In this study, we identify a critical yet underexplored issue in RL training: low-probability tokens disproportionately influence model updates due to their large gradient magnitudes. This dominance hinders the effective learning of high-probability tokens, whose gradients are essential for LLMs' performance but are substantially suppressed. To mitigate this interference, we propose two novel methods: Advantage Reweighting and Low-Probability Token Isolation (Lopti), both of which effectively attenuate gradients from low-probability tokens while emphasizing parameter updates driven by high-probability tokens. Our approaches promote balanced updates across tokens with varying probabilities, thereby enhancing the efficiency of RL training. Experimental results demonstrate that they substantially improve the performance of GRPO-trained LLMs, achieving up to a 46.2% improvement in K&K Logic Puzzle reasoning tasks. Our implementation is available at https://github.com/zhyang2226/AR-Lopti.

cs.CL

Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key

Hallucination remains a major challenge for Large Vision-Language Models (LVLMs). Direct Preference Optimization (DPO) has gained increasing attention as a simple solution to hallucination issues. It directly learns from constructed preference pairs that reflect the severity of hallucinations in responses to the same prompt and image. Nonetheless, different data construction methods in existing works bring notable performance variations. We identify a crucial factor here: outcomes are largely contingent on whether the constructed data aligns on-policy w.r.t the initial (reference) policy of DPO. Theoretical analysis suggests that learning from off-policy data is impeded by the presence of KL-divergence between the updated policy and the reference policy. From the perspective of dataset distribution, we systematically summarize the inherent flaws in existing algorithms that employ DPO to address hallucination issues. To alleviate the problems, we propose On-Policy Alignment (OPA)-DPO framework, which uniquely leverages expert feedback to correct hallucinated responses and aligns both the original and expert-revised responses in an on-policy manner. Notably, with only 4.8k data, OPA-DPO achieves an additional reduction in the hallucination rate of LLaVA-1.5-7B: 13.26% on the AMBER benchmark and 5.39% on the Object-Hal benchmark, compared to the previous SOTA algorithm trained with 16k samples. Our implementation is available at https://github.com/zhyang2226/OPA-DPO.

cs.CV

A Multi-timescale and Chance-Constrained Energy Dispatching Strategy of Integrated Heat-Power Community with Shared Hybrid Energy Storage

The community in the future may develop into an integrated heat-power system, which includes a high proportion of renewable energy, power generator units, heat generator units, and shared hybrid energy storage. In the integrated heat-power system with coupling heat-power generators and demands, the key challenges lie in the interaction between heat and power, the inherent uncertainty of renewable energy and consumers' demands, and the multi-timescale scheduling of heat and power. In this paper, we propose a game theoretic model of the integrated heat-power system. For the welfare-maximizing community operator, its energy dispatch strategy is under chance constraints, where the day-ahead scheduling determines the scheduled energy dispatching strategies, and the real-time dispatch considers the adjustment of generators. For utility-maximizing consumers, their demands are sensitive to the preference parameters. Taking into account the uncertainty in both renewable energy and consumer demand, we prove the existence and uniqueness of the Stackelberg game equilibrium and develop a fixed point algorithm to find the market equilibrium between the community operator and community consumers. Numerical simulations on integrated heat-power system validate the effectiveness of the proposed multi-timescale integrated heat and power model.

eess.SY

Efficient and Robust Equilibrium Strategies of Utilities in Day-ahead Market with Load Uncertainty

We consider the scenario where $N$ utilities strategically bid for electricity in the day-ahead market and balance the mismatch between the committed supply and actual demand in the real-time market, with uncertainty in demand and local renewable generation in consideration. We model the interactions among utilities as a non-cooperative game, in which each utility aims at minimizing its per-unit electricity cost. We investigate utilities' optimal bidding strategies and show that all utilities bidding according to (net load) prediction is a unique pure strategy Nash Equilibrium with two salient properties. First, it incurs no loss of efficiency; hence, competition among utilities does not increase the social cost. Second, it is robust and (0, $N-1$) fault immune. That is, fault behaviors of irrational utilities only help to reduce other rational utilities' costs. The expected market supply-demand mismatch is minimized simultaneously, which improves the planning and supply-and-demand matching efficiency of the electricity supply chain. We prove the results hold under the settings of correlated prediction errors and a general class of real-time spot pricing models, which capture the relationship between the spot price, the day-ahead clearing price, and the market-level mismatch. Simulations based on real-world traces corroborate our theoretical findings. Our study adds new insights to market mechanism design. In particular, we derive a set of fairly general sufficient conditions for the market operator to design real-time pricing schemes so that the interactions among utilities admit the desired equilibrium.

eess.SY

Deadline Scheduling as Restless Bandits

The problem of stochastic deadline scheduling is considered. A constrained Markov decision process model is introduced in which jobs arrive randomly at a service center with stochastic job sizes, rewards, and completion deadlines. The service provider faces random processing costs, convex non-completion penalties, and a capacity constraint that limits the simultaneous processing of jobs. Formulated as a restless multi-armed bandit problem, the stochastic deadline scheduling problem is shown to be indexable. A closed-form expression of the Whittle's index is obtained for the case when the processing costs are constant. An upper bound on the gap-to-optimality for the Whittle's index policy is obtained, and it is shown that the bound converges to zero as the job arrival rate and the number of available processors increase simultaneously to infinity.

math.OC

Policy Design for Controlling Set-Point Temperature of ACs in Shared Spaces of Buildings

Air conditioning systems are responsible for the major percentage of energy consumption in buildings. Shared spaces constitute considerable office space area, in which most office employees perform their meetings and daily tasks, and therefore the ACs in these areas have significant impact on the energy usage of the entire office building. The cost of this energy consumption, however, is not paid by the shared space users, and the AC's temperature set-point is not determined based on the users' preferences. This latter factor is compounded by the fact that different people may have different choices of temperature set-points and sensitivities to change of temperature. Therefore, it is a challenging task to design an office policy to decide on a particular set-point based on such a diverse preference set. As a result, users are not aware of the energy consumption in shared spaces, which may potentially increase the energy wastage and related cost of office buildings. In this context, this paper proposes an energy policy for an office shared space by exploiting an established temperature control mechanism. In particular, we choose meeting rooms in an office building as the test case and design a policy according to which each user of the room can give a preference on the temperature set-point and is paid for felt discomfort if the set-point is not fixed according to the given preference. On the other hand, users who enjoy the thermal comfort compensate the other users of the room. Thus, the policy enables the users to be cognizant and responsible for the payment on the energy consumption of the office space they are sharing, and at the same time ensures that the users are satisfied either via thermal comfort or through incentives. The policy is also shown to be beneficial for building management. Through experiment based case studies, we show the effectiveness of the proposed policy.

eess.SY

Deadline Differentiated Pricing of Deferrable Electric Loads

A large fraction of the total electric load is comprised of end-use devices whose demand for energy is inherently deferrable in time. Of interest is the potential to leverage on such latent flexibility in demand to absorb variability in power supplied from intermittent renewable generation. The challenge, however, lies in designing incentives to reliably induce the desired response in demand. With an eye to electric vehicle charging, we propose a novel forward market for differentiated electric power services, where consumers consent to deferred service of pre-specified loads in exchange for a reduced per-unit price for energy. The longer a consumer is willing to defer, the larger the reduction in price. The proposed forward contract provides a guarantee on the aggregate quantity of energy to be delivered by a consumer-specified deadline. Under the earliest-deadline-first (EDF) scheduling policy, which is shown to be optimal for the supplier, we explicitly characterize a non-discriminatory, deadline-differentiated pricing scheme that yields an efficient competitive equilibrium between the supplier and consumers. We further show that this efficient pricing scheme, in combination with EDF scheduling, is incentive compatible (IC) in that every consumer would like to reveal her true deadline to the supplier, regardless of the actions taken by other consumers.

math.OC

Large-scale Charging of Electric Vehicles: A Multi-Armed Bandit Approach

The successful launch of electric vehicles (EVs) depends critically on the availability of convenient and economic charging facilities. The problem of scheduling of large-scale charging of EVs by a service provider is considered. A Markov decision process model is introduced in which EVs arrive randomly at a charging facility with random demand and completion deadlines. The service provider faces random charging costs, convex non-completion penalties, and a peak power constraint that limits the maximum number of simultaneous activation of EV chargers. Formulated as a restless multi-armed bandit problem, the EV charging problem is shown to be indexable. A closed-form expression of the Whittle's index is obtained for the case when the charging costs are constant. The Whittle's index policy, however, is not optimal in general. An enhancement of the Whittle's index policy based on spatial interchange according to the less laxity and longer processing time principle is presented. The proposed policy outperforms existing charging algorithms, especially when the charging costs are time varying.

math.OC

Dynamic Scheduling for Charging Electric Vehicles: A Priority Rule

We consider the scheduling of multiple tasks with pre-determined deadlines under random processing cost. This problem is motivated by the potential of large scale adoption of plug-in (hybrid) electric vehicles (PHEVs) in the near future. The charging requests of PHEVs usually have deadline constraints, and the electricity cost associated with PHEV charging is usually random due to the uncertainty in both system load and renewable generation. We seek to properly schedule the battery charging of multiple PHEVs so as to minimize the overall cost, which is derived from the total charging cost and the penalty for not completing charging before requested deadlines. Through a dynamic programming formulation, we establish the Less Laxity and Longer remaining Processing time (LLLP) principle that improves any charging policy on a sample-path basis, when the non-completion penalty is a convex function of the additional time needed to fulfill the uncompleted request. Specifically, the LLLP principle states that priority should be given to vehicles that have less laxity and longer remaining processing times. Numerical results demonstrate that heuristic policies that violate the LLLP principle, for example, the earliest deadline first (EDF) policy, can result in significant performance loss.

math.OC

Pricing of Fluctuations in Electricity Markets

In an electric power system, demand fluctuations may result in significant ancillary cost to suppliers. Furthermore, in the near future, deep penetration of volatile renewable electricity generation is expected to exacerbate the variability of demand on conventional thermal generating units. We address this issue by explicitly modeling the ancillary cost associated with demand variability. We argue that a time-varying price equal to the suppliers' instantaneous marginal cost may not achieve social optimality, and that consumer demand fluctuations should be properly priced. We propose a dynamic pricing mechanism that explicitly encourages consumers to adapt their consumption so as to offset the variability of demand on conventional units. Through a dynamic game-theoretic formulation, we show that (under suitable convexity assumptions) the proposed pricing mechanism achieves social optimality asymptotically, as the number of consumers increases to infinity. Numerical results demonstrate that compared with marginal cost pricing, the proposed mechanism creates a stronger incentive for consumers to shift their peak load, and therefore has the potential to reduce the need for long-term investment in peaking plants.

math.OC

Efficiency Loss in a Cournot Oligopoly with Convex Market Demand

We consider a Cournot oligopoly model where multiple suppliers (oligopolists) compete by choosing quantities. We compare the social welfare achieved at a Cournot equilibrium to the maximum possible, for the case where the inverse market demand function is convex. We establish a lower bound on the efficiency of Cournot equilibria in terms of a scalar parameter derived from the inverse demand function, namely, the ratio of the slope of the inverse demand function at the Cournot equilibrium to the average slope of the inverse demand function between the Cournot equilibrium and a social optimum. Also, for the case of a single, monopolistic, profit maximizing supplier, or of multiple suppliers who collude to maximize their total profit, we establish a similar but tighter lower bound on the efficiency of the resulting output. Our results provide nontrivial quantitative bounds on the loss of social welfare for several convex inverse demand functions that appear in the economics literature.

math.OC