SearcharxivSearch

arXiv subjects

Haocheng Luo

Publications and source records attributed to Haocheng Luo.

11 recordsLinked to original sources

Learning Visual Spatial Planning from Symbolic State via Modality-Gap-Aware Self-Distillation

While Vision-Language Models excel at general multimodal understanding, they still struggle with visual spatial planning. We attribute this limitation to a perception--reasoning modality gap. Visual planning requires models to infer latent state structures from pixels and then reason over the recovered structure to produce valid actions, whereas symbolic planning directly leverages explicit representation. This discrepancy introduces two sequential bottlenecks: visual state recovery at the perception stage and multi-step planning at the reasoning stage. To address this, we propose MGSD, a two-stage modality-gap-aware self-distillation framework. First, a cold-start grounding stage establishes reliable visual state recovery before on-policy training. Second, a symbol-guided on-policy self-distillation stage transfers the privileged teacher's planning behavior to the student through token-level supervision on student-generated prefixes. Crucially, symbolic information is used only during training, while inference relies exclusively on visual inputs. Experiments on visual planning benchmarks show that MGSD consistently improves performance across different model scales, raising the macro average by 19.3% and 18.4%, respectively. The resulting models substantially reduce the gap to the upper bounds obtained with symbolic inputs. Ablation studies and diagnostic analyses further confirm that the gains arise from improvements in both visual state recovery and optimal-path reasoning. These results demonstrate that MGSD strengthens not only the recovery of actionable states from visual observations but also the ability to plan over the inferred structures. Code is available at https://github.com/Oranger-l/MGSD.

cs.AI

Sharpness-Aware Minimization in Logit Space Efficiently Enhances Direct Preference Optimization

Direct Preference Optimization (DPO) has emerged as a popular algorithm for aligning pretrained large language models with human preferences, owing to its simplicity and training stability. However, DPO suffers from the recently identified squeezing effect (also known as likelihood displacement), where the probability of preferred responses decreases unintentionally during training. To understand and mitigate this phenomenon, we develop a theoretical framework that models the coordinate-wise dynamics in logit space. Our analysis reveals that negative-gradient updates cause residuals to expand rapidly along high-curvature directions, which underlies the squeezing effect, whereas Sharpness-Aware Minimization (SAM) can suppress this behavior through its curvature-regularization effect. Building on this insight, we investigate logits-SAM, a computationally efficient variant that perturbs only the output layer with negligible overhead. Extensive experiments on Pythia-2.8B, Mistral-7B, and Gemma-2B-IT across multiple datasets and benchmarks demonstrate that logits-SAM consistently improves the effectiveness of DPO and integrates seamlessly with other DPO variants. Code is available at https://github.com/RitianLuo/logits-sam-dpo.

cs.LG

MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment

Scalable embodied intelligence is constrained by the scarcity of diverse, long-horizon robotic manipulation data. Existing video world models in this domain are limited to synthesizing short clips of simple actions and often rely on manually defined trajectories. To this end, we introduce MIND-V, a cognitive hierarchical world model designed to synthesize physically plausible and logically coherent videos of long-horizon robotic manipulation. Inspired by cognitive science, MIND-V bridges high-level reasoning with pixel-level synthesis through three core components: a Semantic Reasoning Hub (SRH) that leverages a pre-trained vision-language model for task planning; a Behavioral Semantic Bridge (BSB) that translates abstract instructions into domain-invariant representations; and a Motor Video Generator (MVG) for conditional video rendering. MIND-V employs Staged Visual Future Rollouts, a test-time optimization strategy to enhance long-horizon robustness. To enforce adherence to physical laws, we introduce a GRPO reinforcement learning post-training phase guided by a novel Physical Foresight Coherence (PFC) reward. PFC leverages the V-JEPA2 world model as a physics referee to penalize implausible dynamics in the latent feature space. Experiments confirm MIND-V's SOTA performance in long-horizon simulation and its significant value for policy learning, introducing a scalable and fully autonomous framework for embodied data synthesis.

cs.RO

Unveiling m-Sharpness Through the Structure of Stochastic Gradient Noise

Sharpness-aware minimization (SAM) has emerged as a highly effective technique to improve model generalization, but its underlying principles are not fully understood. We investigate m-sharpness, where SAM performance improves monotonically as the micro-batch size for computing perturbations decreases, a phenomenon critical for distributed training yet lacking rigorous explanation. We leverage an extended Stochastic Differential Equation (SDE) framework and analyze stochastic gradient noise (SGN) to characterize the dynamics of SAM variants, including n-SAM and m-SAM. Our analysis reveals that stochastic perturbations induce an implicit variance-based sharpness regularization whose strength increases as m decreases. Motivated by this insight, we propose Reweighted SAM (RW-SAM), which employs sharpness-weighted sampling to mimic the generalization benefits of m-SAM while remaining parallelizable. Comprehensive experiments validate our theory and method.Code is available at https://github.com/RitianLuo/RW-SAM.

cs.LG

Sharpness-Aware Teleportation on Riemannian Manifolds

Recent studies highlight the effectiveness of flat minima in enhancing generalization, with sharpness-aware minimization (SAM) achieving state-of-the-art performance. Additionally, insights into the intrinsic geometry of the loss landscape have shown promise for improving model generalization. Building on these advancements, we introduce a novel sharpness-aware, geometry-aware teleportation mechanism to further enhance robustness and generalization. The core innovation of our approach is to decompose each iteration into a teleportation step within a local orbit and a sharpness-aware step that transitions between different orbits, leveraging the Riemannian quotient manifold. Our approach is grounded in a theoretical framework that analyzes the generalization gap between population loss and worst-case empirical loss within the context of Riemannian manifolds. To demonstrate the effectiveness of our method, we evaluate and compare our algorithm on diverse vision benchmarks with various datasets and Riemannian manifolds.

cs.LG

Explicit Eigenvalue Regularization Improves Sharpness-Aware Minimization

Sharpness-Aware Minimization (SAM) has attracted significant attention for its effectiveness in improving generalization across various tasks. However, its underlying principles remain poorly understood. In this work, we analyze SAM's training dynamics using the maximum eigenvalue of the Hessian as a measure of sharpness, and propose a third-order stochastic differential equation (SDE), which reveals that the dynamics are driven by a complex mixture of second- and third-order terms. We show that alignment between the perturbation vector and the top eigenvector is crucial for SAM's effectiveness in regularizing sharpness, but find that this alignment is often inadequate in practice, limiting SAM's efficiency. Building on these insights, we introduce Eigen-SAM, an algorithm that explicitly aims to regularize the top Hessian eigenvalue by aligning the perturbation vector with the leading eigenvector. We validate the effectiveness of our theory and the practical advantages of our proposed approach through comprehensive experiments. Code is available at https://github.com/RitianLuo/EigenSAM.

cs.LG

Re-weighting Tokens: A Simple and Effective Active Learning Strategy for Named Entity Recognition

Active learning, a widely adopted technique for enhancing machine learning models in text and image classification tasks with limited annotation resources, has received relatively little attention in the domain of Named Entity Recognition (NER). The challenge of data imbalance in NER has hindered the effectiveness of active learning, as sequence labellers lack sufficient learning signals. To address these challenges, this paper presents a novel reweighting-based active learning strategy that assigns dynamic smoothed weights to individual tokens. This adaptable strategy is compatible with various token-level acquisition functions and contributes to the development of robust active learners. Experimental results on multiple corpora demonstrate the substantial performance improvement achieved by incorporating our re-weighting strategy into existing acquisition functions, validating its practical efficacy.

cs.CL

An Extreme Learning Machine-Based System Frequency Nadir Constraint Linearization Method

Large-scale integration of converter-based renewable energy sources (RESs) into the power system will lead to a higher risk of frequency nadir limit violation and even frequency instability after the large power disturbance. Therefore, it is essential to consider the frequency nadir constraint (FNC) in power system scheduling. Nevertheless, the FNC is highly nonlinear and non-convex. The state-of-the-art method to simplify the constraint is to construct a low-order frequency response model at first, and then linearize the frequency nadir equation. In this letter, an extreme learning machine (ELM)-based network is built to de-rive the linear formulation of FNC, where the two-step fitting process is integrated into one training process and more details about the physical model of the generator are considered to reduce the fitting error. Simulation results show the superiority of the proposed method on the fitting accuracy.

eess.SY

Plug-in Electric Vehicle Charging Congestion Analysis Using Taxi Travel Data in the Central Area of Beijing

Recharging a plug-in electric vehicle is more time-consuming than refueling an internal combustion engine vehicle. As a result, charging stations may face serious congestion problems during peak traffic hours in the near future with the rapid growth of plug-in electric vehicle population. Considering that drivers' time costs are usually expensive, charging congestion will be a dominant factor that affect a charging station's quality of service. Hence, it is indispensable to conduct adequate congestion analysis when designing charging stations in order to guarantee acceptable quality of service in the future. This paper proposes a data-driven approach for charging congestion analysis of plug-in electric vehicle charging stations. Based on a data-driven plug-in electric vehicle charging station planning model, we adopt the queuing theory to model and analyze the charging congestion phenomenon in these planning results. We simulate and analyze the proposed method for charging stations servicing shared-use electric taxis in the central area of Beijing leveraging real-world taxi travel data.

eess.SY

Coordinated Charging and Discharging Strategies for Plug-in Electric Bus Fast Charging Station with Energy Storage System

Plug-in electric bus (PEB) is an environmentally friendly mode of public transportation and plug-in electric bus fast charging stations (PEBFCSs) play an essential role in the operation of PEBs. Under effective control, deploying an energy storage system (ESS) within a PEBFCS can reduce the peak charging loads and the electricity purchase costs. To deal with the (integrated) scheduling problem of (PEBs charging and) ESS charging and discharging, in this study, we propose an optimal real-time coordinated charging and discharging strategy for a PEBFCS with ESS to achieve maximum economic benefits. According to whether the PEB charging loads are controllable, the corresponding mathematical models are respectively established under two scenarios, i.e., coordinated PEB charging scenario and uncoordinated PEB charging scenario. The price and lifespan of ESS, the capacity charge of PEBFCS and the electricity price arbitrage are considered in the models. Further, under the coordinated PEB charging scenario, a heuristics-based method is developed to get the approximately optimal strategy with computation efficiency dramatically enhanced. Finally, we validate the effectiveness of the proposed strategies, interpret the effect of ESS prices on the usage of ESS, and provide the sensitivity analysis of ESS capacity through the case studies.

eess.SY

Energy Storage Sharing Strategy in Distribution Networks Using Bi-level Optimization Approach

In this paper, we address the energy storage management problem in distribution networks from the perspective of an independent energy storage manager (IESM) who aims to realize optimal energy storage sharing with multi-objective optimization, i.e., optimizing the system peak loads and the electricity purchase costs of the distribution company (DisCo) and its customers. To achieve the goal of the IESM, an energy storage sharing strategy is therefore proposed, which allows DisCo and customers to control the assigned energy storage. The strategy is updated day by day according to the system information change. The problem is formulated as a bi-level mathematical model where the upper level model (ULM) seeks for optimal division of energy storage among Disco and customers, and the lower level models (LLMs) represent the minimizations of the electricity purchase costs of DisCo and customers. Further, in order to enhance the computation efficiency, we transform the bi-level model into a single-level mathematical program with equilibrium constraints (MPEC) model and linearize it. Finally, we validate the effectiveness of the strategy and complement our analysis through case studies.

math.OC