SearcharxivSearch

arXiv subjects

Yuhong Zhao

Publications and source records attributed to Yuhong Zhao.

6 recordsLinked to original sources

Asymmetric Hierarchical Anchoring for Robust Audio-Visual Cross-Modal Generalization

Audio-visual joint representation learning under Cross-Modal Generalization (CMG) aims to transfer knowledge from a labeled source modality to an unlabeled target modality through a unified discrete representation space. Existing symmetric frameworks often suffer from information allocation ambiguity, where the absence of structural inductive bias leads to semantic-specific leakage across modalities. We propose Asymmetric Hierarchical Anchoring (AHA), which enforces directional information allocation by designating a structured semantic anchor within a shared hierarchy. In our instantiation, we exploit the hierarchical discrete representations induced by audio Residual Vector Quantization (RVQ) to guide video feature distillation into a shared semantic space. To ensure representational purity, we replace fragile mutual information estimators with a GRL-based adversarial decoupler that explicitly suppresses semantic leakage in modality-specific branches, and introduce Local Sliding Alignment (LSA) to encourage fine-grained temporal alignment across modalities. Extensive experiments on AVE and AVVP benchmarks demonstrate that AHA consistently outperforms symmetric baselines in cross-modal transfer. Additional analyses on talking-face disentanglement experiment further validate that the learned representations exhibit improved semantic consistency and disentanglement, indicating the broader applicability of the proposed framework.

cs.LG

How Does Urban Context Relate to Residential Building Health? A Vision-POI Fusion Framework for Building-Level Housing Inspection

Housing-level urban physical examination is essential for identifying residential building problems and supporting targeted urban renewal. Existing automated inspection studies primarily rely on individual images and rarely examine whether surrounding urban functional context can provide supplementary information for building-level assessment. This study proposes a vision-POI fusion framework that combines multi-view visual inspection with POI-derived neighborhood context for residential building health assessment. The empirical dataset covers 92 old residential communities, 3,237 residential buildings, and 25,608 field-acquired inspection images in Qingdao, China, encompassing seven categories of housing-related issues. First, multiple object detection models are evaluated to extract issue locations, categories, and confidence scores from individual images. The image-level outputs are subsequently aggregated across multiple views to construct interpretable building-level representations. Second, POI features are extracted within 500m, 1,000m, and 1,500m neighborhood buffers to characterize surrounding functional environments. Pearson and Spearman correlation analyses, combined with false discovery rate correction, are used to identify candidate contextual features. Finally, visual and POI features are integrated using a cost-sensitive Random Forest classifier under community-isolated spatial cross-validation. The results show that multi-view aggregation provides the main performance improvement, increasing the building-level Macro-F1 from 60.84% under Direct Detection to 74.95%. Incorporating POI context further increases Macro-F1 to 76.79%, although the additional gain is modest and category-dependent. POI information therefore functions as a supplementary contextual prior rather than a substitute for direct visual evidence or a causal determinant of building condition.

cs.CV

Bridging Data-Driven and Model-Based Methods: A Learn-to-Optimize Architecture for Distributed Optimal Power Flow

This letter proposes a learn-to-optimize (LTO) architecture for distributed optimal power flow (D-OPF) as the nexus between data-driven and model-based methods. By unfolding alternating direction method of multipliers (ADMM) into a deep neural network (NN) and embedding differentiable optimization layers, our architecture realizes near-instantaneous interpretable distributed decision-making. For mainstream relaxed formulations of D-OPF, the decisions from our architecture achieve comparable optimality with that of state-of-the-art solvers and excelled feasibility compared with existing data-driven approaches. Comparative case studies underpin the effectiveness of our architecture regarding the optimality and feasibility.

eess.SY

Power System Robust State Estimation As a Layer: An Optimization-embedded End-to-end Learning Approach

Serving as an essential prerequisite for modern power system operation, robust state estimation (RSE) could effectively resist noises and outliers in measurements. The emerging neural network (NN) based end-to-end (E2E) learning framework enables real-time application of RSE but potentially yields solutions that are statistically accurate yet physically inconsistent. To bridge this gap, this work proposes a novel E2E learning based RSE framework, where the convex-relaxed RSE problem is innovatively constructed as an explicit differentiable layer into an NN as the first trial. This optimization-embedded layer (termed as `Opt-Layer` in our work) serves as a solver of the RSE problem. Then, the relaxed solutions are recovered through post-processing layers. Through seamlessly embedding the underlying KKT conditions into the gradients during backward propagation, the physical consistency in the estimated states could be significantly enhanced, realizing lower measurement residuals. Also, the measurement weights are treated as learnable parameters of NN to enhance estimation robustness, enabling the Opt-Layer to actively denoise. A hybrid loss function is formulated to pursue accurate and physically consistent solutions. Extensive simulations have been carried out to demonstrate that the proposed framework can significantly improve the SE performance especially in terms of physical consistency on eight test systems, in comparison to classical E2E learning models, physics-informed NN (PINN) models, graph-based learning models, and conventional optimization-based approaches. The estimation performances under partial observability, severe noise contamination are systematically evaluated. Computational complexity and runtime analysis are also comprehensively demonstrated.

eess.SY

Preparation and magnetic properties of (Ln0.2La0.2Nd0.2Sm0.2 Eu0.2)MnO3 (Ln = Dy, Ho, Er) high-entropy perovskite ceramics containing heavy rare earth elements

Equimolar ratio high-entropy perovskite ceramics (HEPCs) have attracted much attention due to their excellent magnetization intensity. To further enhance their magnetization intensities, (Ln0.2La0.2Nd0.2Sm0.2Eu0.2)MnO3 (Ln = Dy, Ho and Er, labeled as Ln-LNSEMO) HEPCs are designed based on the configuration entropy Sconfig, tolerance factor t, and mismatch degree. Single-phase HEPCs are synthesized by the solid-phase method in this work, in which the effects of the heavy rare-earth elements Dy, Ho and Er on the structure and magnetic properties of Ln-LNSEMO are systematically studied. The results show that all Ln-LNSEMO HEPCs exhibit high crystallinity and maintain excellent structural stability after sintering at 1250 degree centigrade for 16 h. Ln-LNSEMO HEPCs exhibit significant lattice distortion effects, with smooth surface morphology, clearly distinguishable grain boundaries, and irregular polygonal shapes. The three high-entropy ceramic samples exhibit hysteresis behavior at T = 5 K, with the Curie temperature TC decreasing as the radius of the introduced rare-earth ions decreases, while the saturation magnetization and coercivity increase accordingly. When the average ionic radius of A-site decreases, the interaction between their valence electrons and local electrons in the crystal increases, thereby enhancing the conversion of electrons to oriented magnetic moments under an external magnetic field. Thus, Er-LNSEMO HEPC shows a higher saturation magnetization strength (42.8 emu/g) and coercivity (2.09 kOe) than the other samples, which is attributed to the strong magnetic crystal anisotropy, larger lattice distortion (0.00652), smaller average grain size (440.49 plus or minus 22.02 nm), unit cell volume (229.432 A3) and A-site average ion radius (1.24 A) of its magnet. The Er-LNSEMO HEPC has potential applications in magnetic recording materials.

cond-mat.mtrl-sci

Learning Explicit User Interest Boundary for Recommendation

The core objective of modelling recommender systems from implicit feedback is to maximize the positive sample score $s_p$ and minimize the negative sample score $s_n$, which can usually be summarized into two paradigms: the pointwise and the pairwise. The pointwise approaches fit each sample with its label individually, which is flexible in weighting and sampling on instance-level but ignores the inherent ranking property. By qualitatively minimizing the relative score $s_n - s_p$, the pairwise approaches capture the ranking of samples naturally but suffer from training efficiency. Additionally, both approaches are hard to explicitly provide a personalized decision boundary to determine if users are interested in items unseen. To address those issues, we innovatively introduce an auxiliary score $b_u$ for each user to represent the User Interest Boundary(UIB) and individually penalize samples that cross the boundary with pairwise paradigms, i.e., the positive samples whose score is lower than $b_u$ and the negative samples whose score is higher than $b_u$. In this way, our approach successfully achieves a hybrid loss of the pointwise and the pairwise to combine the advantages of both. Analytically, we show that our approach can provide a personalized decision boundary and significantly improve the training efficiency without any special sampling strategy. Extensive results show that our approach achieves significant improvements on not only the classical pointwise or pairwise models but also state-of-the-art models with complex loss function and complicated feature encoding.

cs.IR