SearcharxivSearch

arXiv subjects

Mingzhong Wang

Publications and source records attributed to Mingzhong Wang.

11 recordsLinked to original sources

SRG: Score-based Relaxation-guided Generation for Mixed Integer Linear Programming

We propose Score-based Relaxation-guided Generation (SRG), a generative framework based on an approximate formulation of relaxation-guided stochastic differential equations (SDEs) for mixed-integer linear programming. SRG employs a Transformer-based score network that incorporates feasibility and optimality signals into score modeling, encouraging the learned generative model to place more probability mass on feasible, high-quality regions of the solution space. At inference time, SRG directly samples diverse candidate solutions from the learned score model without requiring any additional guidance module. These candidates are then used to construct compact trust-region subproblems for standard MILP solvers. Across multiple public benchmarks, SRG matches or improves upon the solution quality of the strongest learning-based baselines, with particularly strong gains in challenging candidate-generation settings. Moreover, SRG shows promising zero-shot transferability to unseen cross-scale and cross-problem instances, improving solver objectives and reducing search time in several cases through higher-quality initial candidates and compact trust-region search.

cs.LG

Offline Meta-Reinforcement Learning with Flow-Based Task Inference and Adaptive Correction of Feature Overgeneralization

Offline meta-reinforcement learning (OMRL) combines the strengths of learning from diverse datasets in offline RL with the adaptability to new tasks of meta-RL, promising safe and efficient knowledge acquisition by RL agents. However, OMRL still suffers extrapolation errors due to out-of-distribution (OOD) actions, compromised by broad task distributions and Markov Decision Process (MDP) ambiguity in meta-RL setups. Existing research indicates that the generalization of the $Q$ network affects the extrapolation error in offline RL. This paper investigates this relationship by decomposing the $Q$ value into feature and weight components, observing that while decomposition enhances adaptability and convergence in the case of high-quality data, it often leads to policy degeneration or collapse in complex tasks. We observe that decomposed $Q$ values introduce a large estimation bias when the feature encounters OOD samples, a phenomenon we term ''feature overgeneralization''. To address this issue, we propose FLORA, which identifies OOD samples by modeling feature distributions and estimating their uncertainties. FLORA integrates a return feedback mechanism to adaptively adjust feature components. Furthermore, to learn precise task representations, FLORA explicitly models the complex task distribution using a chain of invertible transformations. We theoretically and empirically demonstrate that FLORA achieves rapid adaptation and meta-policy improvement compared to baselines across various environments.

cs.LG

Wavelet Predictive Representations for Non-Stationary Reinforcement Learning

The real world is inherently non-stationary, with ever-changing factors, such as weather conditions and traffic flows, making it challenging for agents to adapt to varying environmental dynamics. Non-Stationary Reinforcement Learning (NSRL) addresses this challenge by training agents to adapt rapidly to sequences of distinct Markov Decision Processes (MDPs). However, existing NSRL approaches often focus on tasks with regularly evolving patterns, leading to limited adaptability in highly dynamic settings. Inspired by the success of Wavelet analysis in time series modeling, specifically its ability to capture signal trends at multiple scales, we propose WISDOM to leverage wavelet-domain predictive task representations to enhance NSRL. WISDOM captures these multi-scale features in evolving MDP sequences by transforming task representation sequences into the wavelet domain, where wavelet coefficients represent both global trends and fine-grained variations of non-stationary changes. In addition to the auto-regressive modeling commonly employed in time series forecasting, we devise a wavelet temporal difference (TD) update operator to enhance tracking and prediction of MDP evolution. We theoretically prove the convergence of this operator and demonstrate policy improvement with wavelet task representations. Experiments on diverse benchmarks show that WISDOM significantly outperforms existing baselines in both sample efficiency and asymptotic performance, demonstrating its remarkable adaptability in complex environments characterized by non-stationary and stochastically evolving tasks.

cs.LG

Robust Deep Signed Graph Clustering via Weak Balance Theory

Signed graph clustering is a critical technique for discovering community structures in graphs that exhibit both positive and negative relationships. We have identified two significant challenges in this domain: i) existing signed spectral methods are highly vulnerable to noise, which is prevalent in real-world scenarios; ii) the guiding principle ``an enemy of my enemy is my friend'', rooted in \textit{Social Balance Theory}, often narrows or disrupts cluster boundaries in mainstream signed graph neural networks. Addressing these challenges, we propose the \underline{D}eep \underline{S}igned \underline{G}raph \underline{C}lustering framework (DSGC), which leverages \textit{Weak Balance Theory} to enhance preprocessing and encoding for robust representation learning. First, DSGC introduces Violation Sign-Refine to denoise the signed network by correcting noisy edges with high-order neighbor information. Subsequently, Density-based Augmentation enhances semantic structures by adding positive edges within clusters and negative edges across clusters, following \textit{Weak Balance} principles. The framework then utilizes \textit{Weak Balance} principles to develop clustering-oriented signed neural networks to broaden cluster boundaries by emphasizing distinctions between negatively linked nodes. Finally, DSGC optimizes clustering assignments by minimizing a regularized clustering loss. Comprehensive experiments on synthetic and real-world datasets demonstrate DSGC consistently outperforms all baselines, establishing a new benchmark in signed graph clustering.

cs.SI

Towards Control-Centric Representations in Reinforcement Learning from Images

Image-based Reinforcement Learning is a practical yet challenging task. A major hurdle lies in extracting control-centric representations while disregarding irrelevant information. While approaches that follow the bisimulation principle exhibit the potential in learning state representations to address this issue, they still grapple with the limited expressive capacity of latent dynamics and the inadaptability to sparse reward environments. To address these limitations, we introduce ReBis, which aims to capture control-centric information by integrating reward-free control information alongside reward-specific knowledge. ReBis utilizes a transformer architecture to implicitly model the dynamics and incorporates block-wise masking to eliminate spatiotemporal redundancy. Moreover, ReBis combines bisimulation-based loss with asymmetric reconstruction loss to prevent feature collapse in environments with sparse rewards. Empirical studies on two large benchmarks, including Atari games and DeepMind Control Suit, demonstrate that ReBis has superior performance compared to existing methods, proving its effectiveness.

cs.LG

SimSR: Simple Distance-based State Representation for Deep Reinforcement Learning

This work explores how to learn robust and generalizable state representation from image-based observations with deep reinforcement learning methods. Addressing the computational complexity, stringent assumptions and representation collapse challenges in existing work of bisimulation metric, we devise Simple State Representation (SimSR) operator. SimSR enables us to design a stochastic approximation method that can practically learn the mapping functions (encoders) from observations to latent representation space. In addition to the theoretical analysis and comparison with the existing work, we experimented and compared our work with recent state-of-the-art solutions in visual MuJoCo tasks. The results shows that our model generally achieves better performance and has better robustness and good generalization.

cs.LG

Relationship between the TC of smart meta-superconductor Bi(Pb)SrCaCuO and inhomogeneous phase content

A smart meta-superconductor Bi(Pb)SrCaCuO (B(P)SCCO) may increase the critical transition temperature (TC) of B(P)SCCO by electroluminescence (EL) energy injection of inhomogeneous phases. However, the increase amplitude ΔTC (ΔTC=TC-T(C,pure)) of TC is relatively small. In this study, a smart meta-superconductor B(P)SCCO with different matrix sizes was designed. Three kinds of raw materials with different particle sizes were used, and different series of Y2O3:Sm3+, Y2O3, Y2O3:Eu3+, and Y2O3:Eu3++Ag doped samples and pure B(P)SCCO were prepared. Results indicated that the TC of the Y2O3 or Y2O3:Sm3+ non-luminescent dopant doping sample is lower than that of pure B(P)SCCO. However, the TC of the Y2O3:Eu3++Ag or Y2O3:Eu3+ luminescent inhomogeneous phase doping sample is higher than that of pure B(P)SCCO. With the decrease of the raw material particle size from 30 to 5 μm, the particle size of the B(P)SCCO superconducting matrix in the prepared samples decreases, and the doping content of the Y2O3:Eu3++Ag or Y2O3:Eu3+ increases from 0.2% to 0.4%. Meanwhile, the increase of the inhomogeneous phase content enhances the ΔTC. When the particle size of raw material is 5 μm, the doping concentration of the luminescent inhomogeneous phase can be increased to 0.4%. At this time, the zero-resistance temperature and onset transition temperature of the Y2O3:Eu3++Ag doped sample are 4 and 6.3 K higher than those of pure B(P)SCCO, respectively.

cond-mat.supr-con

Smart metastructure method for increasing TC of Bi(Pb)SrCaCuO high-temperature superconductors

Improving the critical transition temperature (TC) of Bi(Pb)SrCaCuO (B(P)SCCO) high-temperature superconductors is important, however, considerable challenges exist. In this study, on the basis of the metamaterial structure and the idea that the injecting energy will promote the formation of Cooper pairs, a smart meta-superconductor B(P)SCCO consisting of B(P)SCCO microparticles and Y2O3:Eu3++Ag or Y2O3:Eu3+ luminophor was designed. In the applied electric field, the Y2O3:Eu3++Ag or Y2O3:Eu3+ luminophor generates an electroluminescence (EL), thereby promoting the TC via EL energy injection. A series of Y2O3:Eu3++Ag topological luminophor-doped B(P)SCCO samples was prepared. Results showed that Y2O3:Eu3++Ag was dispersed around B(P)SCCO particles, forming a metastructure. Accordingly, the onset transition temperature (T_(C,on)) and zero resistance transition temperature (T_(C,0)) of B(P)SCCO increased. Meanwhile, the B(P)SCCO sample doped with 0.2 wt% Y2O3 or Y2O3:Sm3+ nonluminous inhomogeneous phase was also prepared to further prove the influence of EL on the T_C rather than the rare earth effect. Results indicated that the TC of the Y2O3 or Y2O3:Sm3+ doping sample decreased. However, the TC of the 0.2 wt% Y2O3:Eu3++Ag or Y2O3:Eu3+ luminophor-doped sample improved. This outcome further demonstrated that the smart metastructure method can improve the TC of B(P)SCCO.

cond-mat.supr-con

Smart meta-superconductor MgB2 constructed by inhomogeneous phase of luminescent nanocomposite

On the basis of the idea that the injecting energy will improve the conditions for the formation of Cooper pairs, a smart meta-superconductor (SMSC) was prepared by doping inhomogeneous phase of luminescent nanocomposite Y2O3:Eu3+/Ag, which has the strong luminescence characteristic, in MgB2 to improve the superconducting transition temperature (TC) of the MgB2-based superconductor. Two types of Y2O3:Eu3+/Ag with different sizes were prepared and marked as m-Y2O3:Eu3+/Ag and nY2O3:Eu3+/Ag. MgB2 SMSC was prepared through an ex situ process. Results show that when the inhomogeneous phase content was fixed at 2.0 wt.%, the TC of MgB2 SMSC increased initially then decreased with the increase in the Ag content in the dopant. When the Ag content accounted for 5 wt.% of the inhomogeneous phase weight, the TC of MgB2 SMSC was 37.2-38.0 K, which was similar to that of pure MgB2. Meanwhile, the TC of MgB2 SMSC doped with n-Y2O3:Eu3+/Ag increased initially then decreased basically with the increase in the content of n-Y2O3:Eu3+/Ag, in which Ag accounted for 5 wt.% of the inhomogeneous phase. The TC of MgB2 SMSC doped with 0.5 wt.% n-Y2O3:Eu3+/Ag was 37.6-38.4 K, which was 0.4 K higher than that of pure MgB2. It is thought that the doping inhomogeneous phase of luminescent nanocomposite into the superconductor is a new means to improve the TC of SMSC.

cond-mat.supr-con

Topological luminophor Y2O3:Eu3++Ag with high electroluminescence performance

Improving luminescent intensity is a significant technical requirement and scientific problem for the luminescent performance of fluorophor materials through the ages. The process control and luminescence performance still limit the developments of luminescent intensity even through it can be improved partly by covering or magnetron sputtering of precious metals on the surface of the fluorophore materials. On the basis of the improvement of luminescence center radiative transition rate by surface plasma resonance and Y2O3:Eu3+ microsheet phosphors, a fundamental model for topological luminophor Y2O3:Eu3++Ag was designed referencing the concepts of topological materials in order to enhance luminescent performance by composite-luminescence, which composed of Eu3+centric electroluminescence and surface plasma-enhanced photoluminescence by Ag. The topological luminophor Y2O3:Eu3++Ag was successfully synthesized with an asymmetric-discrete Ag nanocrystal topological structure on the surface just via illumination. Experiment results suggest that the luminescence performance of topological luminophor Y2O3:Eu3++Ag increased by about 300% compared with that of Y2O3: Eu3+ phosphors on the same conditions. The design of a topological luminophor provides a new approach to further improve the luminescent intensity of phosphors.

cond-mat.mes-hall

GANE: A Generative Adversarial Network Embedding

Network embedding has become a hot research topic recently which can provide low-dimensional feature representations for many machine learning applications. Current work focuses on either (1) whether the embedding is designed as an unsupervised learning task by explicitly preserving the structural connectivity in the network, or (2) whether the embedding is a by-product during the supervised learning of a specific discriminative task in a deep neural network. In this paper, we focus on bridging the gap of the two lines of the research. We propose to adapt the Generative Adversarial model to perform network embedding, in which the generator is trying to generate vertex pairs, while the discriminator tries to distinguish the generated vertex pairs from real connections (edges) in the network. Wasserstein-1 distance is adopted to train the generator to gain better stability. We develop three variations of models, including GANE which applies cosine similarity, GANE-O1 which preserves the first-order proximity, and GANE-O2 which tries to preserves the second-order proximity of the network in the low-dimensional embedded vector space. We later prove that GANE-O2 has the same objective function as GANE-O1 when negative sampling is applied to simplify the training process in GANE-O2. Experiments with real-world network datasets demonstrate that our models constantly outperform state-of-the-art solutions with significant improvements on precision in link prediction, as well as on visualizations and accuracy in clustering tasks.

cs.LG