SearcharxivSearch

arXiv subjects

Can Hu

Publications and source records attributed to Can Hu.

3 recordsLinked to original sources

Local Laws and Edge Universality for Noncentral Sample Covariance Matrices

We consider the real noncentral sample covariance matrices $\mathcal{W}=YY^\top$ with $Y=A+\Sigma^{1/2}X$. Here $A\in\mathbb{R}^{M\times N}$ is deterministic, $\Sigma$ is a deterministic positive definite population covariance matrix and $X\in\mathbb{R}^{M\times N}$ has independent centered entries with variance $N^{-1}$. We prove local laws near regular right edges down to optimal spectral scales without requiring the commutativity of $AA^\top$ and $\Sigma$. As a consequence, we obtain optimal eigenvalue rigidity at the rightmost regular edge and delocalization of the corresponding left and right singular vectors. We also show that, with high probability, there are no eigenvalues in the adjacent spectral gap beyond the optimal $N^{-2/3}$ edge scale, up to an arbitrarily small $N^\varepsilon$ loss. Finally, we establish edge universality at the rightmost regular edge: after centering and scaling, the largest eigenvalue converges to the Tracy--Widom distribution. The main technical ingredient is a stability analysis of the matrix Dyson equation (MDE) associated with the linearization of $Y$, whose self-energy operator does not satisfy the flatness condition of the general MDE theory. Exploiting the special block structure, we reduce the stability analysis exactly to a two-dimensional operator. This reduction yields regularity of the spectral density and square-root behavior at regular right edges, together with sharp stability bounds near such edges.

math.PR

Survey on Remote Sensing Scene Classification: From Traditional Methods to Large Generative AI Models

Remote sensing scene classification has experienced a paradigmatic transformation from traditional handcrafted feature methods to sophisticated artificial intelligence systems that now form the backbone of modern Earth observation applications. This comprehensive survey examines the complete methodological evolution, systematically tracing development from classical texture descriptors and machine learning classifiers through the deep learning revolution to current state-of-the-art foundation models and generative AI approaches. We chronicle the pivotal shift from manual feature engineering to automated hierarchical representation learning via convolutional neural networks, followed by advanced architectures including Vision Transformers, graph neural networks, and hybrid frameworks. The survey provides in-depth coverage of breakthrough developments in self-supervised foundation models and vision-language systems, highlighting exceptional performance in zero-shot and few-shot learning scenarios. Special emphasis is placed on generative AI innovations that tackle persistent challenges through synthetic data generation and advanced feature learning strategies. We analyze contemporary obstacles including annotation costs, multimodal data fusion complexities, interpretability demands, and ethical considerations, alongside current trends in edge computing deployment, federated learning frameworks, and sustainable AI practices. Based on comprehensive analysis of recent advances and gaps, we identify key future research priorities: advancing hyperspectral and multi-temporal analysis capabilities, developing robust cross-domain generalization methods, and establishing standardized evaluation protocols to accelerate scientific progress in remote sensing scene classification systems.

cs.CV

MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook

This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We hope it better allows researchers to follow the state-of-the-art in this very dynamic area. Meanwhile, a growing number of testbeds have boosted the evolution of general-purpose large language models. Thus, this year's MARS2 focuses on real-world and specialized scenarios to broaden the multimodal reasoning applications of MLLMs. Our organizing team released two tailored datasets Lens and AdsQA as test sets, which support general reasoning in 12 daily scenarios and domain-specific reasoning in advertisement videos, respectively. We evaluated 40+ baselines that include both generalist MLLMs and task-specific models, and opened up three competition tracks, i.e., Visual Grounding in Real-world Scenarios (VG-RS), Visual Question Answering with Spatial Awareness (VQA-SA), and Visual Reasoning in Creative Advertisement Videos (VR-Ads). Finally, 76 teams from the renowned academic and industrial institutions have registered and 40+ valid submissions (out of 1200+) have been included in our ranking lists. Our datasets, code sets (40+ baselines and 15+ participants' methods), and rankings are publicly available on the MARS2 workshop website and our GitHub organization page https://github.com/mars2workshop/, where our updates and announcements of upcoming events will be continuously provided.

cs.CV