SearcharxivSearch

arXiv subjects

Zhiming Xu

Publications and source records attributed to Zhiming Xu.

At least 19 recordsLinked to original sources

Task-Anchored Representation Shaping for Pre-Trained Model-Based Continual Learning

Pre-trained models (PTMs) provide a strong foundation for continual learning by offering stable representations that facilitate lightweight adaptation to new tasks. However, adapting well to each task does not ensure reliable inference over all learned tasks. Since task boundaries are often artificial and semantically entangled, an input from an unknown task can remain ambiguous even with strong PTM features, making cross-task prediction a key bottleneck. We propose Task-Anchored Inference Latent Shaping (TAILS), a lightweight post-PTM module that can be integrated into diverse continual learners and optimized through a decoupled step. TAILS uses fixed task anchors as persistent references to accumulated knowledge. It interprets each sample's feature representation relative to these references, then composes relevant evidence across tasks into latent recall. Rather than selecting a task-specific path or adjusting classifier outputs, TAILS uses latent recall to directly correct the feature representation before prediction. It therefore resolves cross-task ambiguity at the representation level, while leaving the original PTM, method-specific modules, and classifier unchanged. Extensive experiments across multiple PTM-based continual learning paradigms show that TAILS can improve classification and task-inference performance with modest parameter overhead and negligible inference cost.

cs.LG

SAMBA: A Scatter-Guided Masked Bidirectional Mamba Foundation Model for SAR Target Recognition

Synthetic aperture radar automatic target recognition (SAR ATR) is critical for Earth observation and defense, but its practical deployment is constrained by scarce annotated training data. Self-supervised pre-training alleviates this label bottleneck, yet prevailing Transformer architectures incur prohibitive quadratic computational complexity, and conventional universal masking neglects the unique electromagnetic scattering properties intrinsic to SAR imagery. To address these limitations, we propose SAMBA (Scattering-Guided Bidirectional Mamba), an efficient self-supervised pre-training foundation model for SAR target interpretation. Our framework features three core innovations: (i) a linear-complexity Mamba encoder with a mid-sequence class token to mitigate computational bottlenecks; (ii) a three-level hierarchical Scattering-Guided Masked Autoencoder (SG-MAE) masking strategy guided by SAR physical priors, aligning the pretext task with SAR's intrinsic imaging mechanism; (iii) a lightweight SpatialMix feature interaction module to enhance cross-region feature fusion. We also design a two-stage cross-domain pre-training pipeline to optimize the overall pre-training process. Extensive evaluations demonstrate that SAMBA consistently delivers superior performance across all pre-training configurations, with substantially fewer parameters than both CNN and Transformer baselines. Compared with the default masking strategy in standard MAE, the proposed SG-MAE strategy further boosts the model's few-shot transfer capability. Benchmarking on seven downstream datasets covering classification and detection tasks shows SAMBA achieves state-of-the-art (SOTA) performance on most metrics, fully validating its robust generalizability across diverse SAR interpretation tasks. Source code and pre-trained weights are publicly available at https://github.com/mynswkk/SAMBA.

cs.CV

Thickness-Independent Quantum Geometric Responses Driven by Interlayer Antiferroic Coupling

Two-dimensional ferroic materials exhibit rich and intriguing physical phenomena, but their response properties generally depend sensitively on thickness, requiring precise layer-number control and thereby limiting practical applications. Here, we propose a general strategy for realizing thickness-independent quantum geometric responses through symmetry engineering induced by interlayer antiferroic coupling. Using spatial-dependent symmetry analysis, we show that thickness-independent behavior emerges when the symmetry breaking required for a given response is generated by interlayer antiferromagnetic (AFM) or antiferroelectric (AFE) coupling, without invoking topological mechanisms. Our first-principles calculations predict that multilayer MnS in the G-type AFM configuration exhibits a surface-dominated anomalous Hall effect, whose thickness-independent behavior can be significantly influenced by the stacking order. We further propose design principles for achieving thickness-independent anomalous and nonlinear Hall effects driven by interlayer AFE coupling, and suggest potential applications in distinguishing magnetic structures. Our findings open a new route towards robust functional devices based on antiferroic materials.

cond-mat.mtrl-sci

KinematicRL: A Sim-to-Real Reinforcement Learning Framework For Social Navigation With Kinodynamic Feasibility

Deep Reinforcement Learning (DRL) has shown promise for social navigation, yet its real-world deployment remains hindered by a persistent sim-to-real gap arising from simplified first-order dynamics and context-specific human state estimation pipelines. This work presents a unified framework that addresses these limitations to produce dynamically feasible navigation policies suitable for real-world deployment. First, theoretical analysis reveals that tracking error between simulated and actual robot position decays exponentially with increased control order, motivating the use of higher-order control inputs as DRL action space. A second-order control formulation tailored to differential drive robots is developed, complemented by a stochastic iterative Linear Quadratic Regulator (iLQR) that pretrains the policy via a divergence minimization objective. Second, to avoid the added system complexity of camera-LiDAR fusion, a cluster-based human tracking pipeline using only 2D LiDAR is introduced. Human detections are associated according to both spatial proximity and velocity similarity, enabling reliable differentiation of nearby pedestrians and yielding stable velocity estimates through temporal aggregation. Third, we introduce an unbiased residual gating block to balance reaction- and memory-based behaviors while handling time-varying crowd sizes, both critical for social navigation. The resulting policy, KinematicRL, consistently improves kinematic performance and adapts to varying number of detected humans. Experiments in real-world environments demonstrate that, when combined with the proposed tracking pipeline, KinematicRL can be deployed on a real differential drive robot with minimal modifications.

cs.RO

Dynamics Are Learned, Not Told: Semi-Supervised Discovery of Latent Dynamics Geometries For Zero-Shot Policy Adaptation

Real-world dynamics shifts pose a critical challenge for reinforcement learning in robotics, as policies tightly coupled to nominal environments often fail catastrophically when physical conditions change. Most existing methods rely on encoding explicitly identified physical parameters into a latent context, a parameter-centric paradigm that depends on pre-specified axes of variation and becomes brittle under unmodeled or compound dynamics changes. We revisit dynamics adaptation from an outcome-centric perspective: rather than telling policies what the dynamics are, we enable them to learn how dynamics affect interaction outcomes. Theoretically, this is grounded in a monotonic relationship between target-domain regret and the Lipschitz constant of a trajectory dynamics encoder. Practically, this constant can be upper-bounded through contrastive learning, yielding a smooth, task-relevant latent topology without privileged dynamics information. On MuJoCo benchmarks, our method consistently outperforms parameter-centric baselines under severe dynamics shifts, including unmodeled and time-varying parameters, while also improving in-distribution stability and latent interpretability. Overall, these results validate that controlling latent geometry is a principled mechanism for robust adaptation.

cs.RO

Beyond Point-wise Neural Collapse: A Topology-Aware Hierarchical Classifier for Class-Incremental Learning

The Nearest Class Mean (NCM) classifier is widely favored in Class-Incremental Learning (CIL) for its superior resistance to catastrophic forgetting compared to Fully Connected layers. While Neural Collapse (NC) theory supports NCM's optimality by assuming features collapse into single points, non-linear feature drift and insufficient training in CIL often prevent this ideal state. Consequently, classes manifest as complex manifolds rather than collapsed points, rendering the single-point NCM suboptimal. To address this, we propose Hierarchical-Cluster SOINN (HC-SOINN), a novel classifier that captures the topological structure of these manifolds via a ``local-to-global'' representation. Furthermore, we introduce Structure-Topology Alignment via Residuals (STAR) method, which employs a fine-grained pointwise trajectory tracking mechanism to actively deform the learned topology, allowing it to adapt precisely to complex non-linear feature drift. Theoretical analysis and Procrustes distance experiments validate our framework's resilience to manifold deformations. We integrated HC-SOINN into seven state-of-the-art methods by replacing their original classifiers, achieving consistent improvements that highlight the effectiveness and robustness of our approach. Code is available at https://github.com/yhyet/HC_SOINN.

cs.CV

Free-Flow Class-Incremental Learning: Towards Robust CIL under Variable Class Arrivals

Class-incremental learning (CIL) is commonly evaluated under predefined schedules with fixed or nearly equal class increments, leaving irregular class-arrival scenarios underexplored. However, practical CIL systems may need to update whenever new categories emerge, without forcing them into balanced task partitions. We formalize this setting as Free-Flow Class-Incremental Learning (FFCIL), where the number of newly arriving classes can vary substantially across learning stages. We show that variable class arrivals alter the class composition of incremental training, the reliability of new-class classifier statistics, and the consistency of representations learned across increments, causing clear performance degradation in both conventional and pre-trained model (PTM)-based CIL methods. To improve robustness under FFCIL, we introduce a general framework consisting of Class-Wise Mean (CWM), which replaces instance-wise loss aggregation with class-wise averaging; Dynamic Intervention Weight Alignment (DIWA), which adjusts new-class weight calibration according to the current increment size; and Head-Agnostic Alignment (HA), which performs feature-level correction for PTM-based methods using current and previous-class feature supervision. Extensive experiments across diverse methods, datasets, and class-arrival schedules demonstrate the general impact of FFCIL and the consistent effectiveness of our framework.

cs.LG

ELIQ: A Label-Free Framework for Quality Assessment of Evolving AI-Generated Images

Generative text-to-image models are advancing at an unprecedented pace, continuously shifting the perceptual quality ceiling and rendering previously collected labels unreliable for newer generations. To address this, we present ELIQ, a Label-free Framework for Quality Assessment of Evolving AI-generated Images. Specifically, ELIQ focuses on visual quality and prompt-image alignment, automatically constructs positive and aspect-specific negative pairs to cover both conventional distortions and AIGC-specific distortion modes, enabling transferable supervision without human annotations. Building on these pairs, ELIQ adapts a pre-trained multimodal model into a quality-aware critic via instruction tuning and predicts two-dimensional quality using lightweight gated fusion and a Quality Query Transformer. Experiments across multiple benchmarks demonstrate that ELIQ consistently outperforms existing label-free methods, generalizes from AI-generated content (AIGC) to user-generated content (UGC) scenarios without modification, and paves the way for scalable and label-free quality assessment under continuously evolving generative models. The code will be released upon publication.

cs.CV

Decoupling Perception and Calibration: Label-Efficient Image Quality Assessment Framework

Recent multimodal large language models (MLLMs) have demonstrated strong capabilities in image quality assessment (IQA) tasks. However, adapting such large-scale models is computationally expensive and still relies on substantial Mean Opinion Score (MOS) annotations. We argue that for MLLM-based IQA, the core bottleneck lies not in the quality perception capacity of MLLMs, but in MOS scale calibration. Therefore, we propose LEAF, a Label-Efficient Image Quality Assessment Framework that distills perceptual quality priors from an MLLM teacher into a lightweight student regressor, enabling MOS calibration with minimal human supervision. Specifically, the teacher conducts dense supervision through point-wise judgments and pair-wise preferences, with an estimate of decision reliability. Guided by these signals, the student learns the teacher's quality perception patterns through joint distillation and is calibrated on a small MOS subset to align with human annotations. Experiments on both user-generated and AI-generated IQA benchmarks demonstrate that our method significantly reduces the need for human annotations while maintaining strong MOS-aligned correlations, making lightweight IQA practical under limited annotation budgets.

cs.CV

Pushing the Limits of Distillation-Based Continual Learning via Classifier-Proximal Lightweight Plugins

Continual learning requires models to learn continuously while preserving prior knowledge under evolving data streams. Distillation-based methods are appealing for retaining past knowledge in a shared single-model framework with low storage overhead. However, they remain constrained by the stability-plasticity dilemma: knowledge acquisition and preservation are still optimized through coupled objectives, and existing enhancement methods do not alter this underlying bottleneck. To address this issue, we propose a plugin extension paradigm termed Distillation-aware Lightweight Components (DLC) for distillation-based CL. DLC deploys lightweight residual plugins into the base feature extractor's classifier-proximal layer, enabling semantic-level residual correction for better classification accuracy while minimizing disruption to the overall feature extraction process. During inference, plugin-enhanced representations are aggregated to produce classification predictions. To mitigate interference from non-target plugins, we further introduce a lightweight weighting unit that learns to assign importance scores to different plugin-enhanced representations. DLC could deliver a significant 8% accuracy gain on large-scale benchmarks while introducing only a 4% increase in backbone parameters, highlighting its exceptional efficiency. Moreover, DLC is compatible with other plug-and-play CL enhancements and delivers additional gains when combined with them.

cs.LG

Tunable Chern Insulator States with Coexisting Magnonic and Electronic Topology in 2D Honeycomb Kitaev Ferromagnets

The coexistence of topological magnons and electrons in magnetic materials presents a compelling route toward developing low-dissipation, multifunctional spintronic devices. However, material systems enabling their simultaneous realization and control remain largely unexplored. Here, we propose the coexistence and concurrent tunability of magnonic and electronic Chern insulator phases in Kitaev magnets and use MnBr$_{3}$ monolayer as a prototype. We find the significant Kitaev interaction in MnBr$_{3}$ induces the magnonic Chern insulator phase, manifesting as the magnon thermal Hall effect. Concurrently, MnBr$_{3}$ exhibits the quantum anomalous Hall effect driven by its electronic Chern insulator phase. Crucially, we demonstrate that these dual topological phases can be simultaneously controlled by reorienting the in-plane spins with an external magnetic field. Our findings not only deepen the fundamental understanding of spin excitations in Kitaev magnets but also provide a promising platform for exploring the interplay between electronic and magnonic topology.

cond-mat.mes-hall

Magnetic-Field Tunable M\"{o}bius and Higher-Order Topological Insulators in Three-Dimensional Layered Octagonal Quasicrystals

We propose that three-dimensional layered octagonal quasicrystals can host magnetic-field-tunable M\"{o}bius insulators and various higher-order topological insulators (HOTIs), enabled by the interplay of quasicrystalline symmetry and magnetic order. By constructing a minimal model based on stacked Ammann-Beenker tilings with magnetic exchange coupling and octagonal warping, we demonstrate that an A-type antiferromagnetic (AFM) configuration yields a topological phase protected by an effective time-reversal symmetry $\mathcal{S}=\mathcal{T}\tau_{1/2}$. Breaking $\mathcal{S}$ via an in-plane magnetic field induced canting of the AFM order while preserving a nonsymmorphic glide symmetry $\mathcal{G}_n=\tau_{1/2}\mathcal{M}_n$ leads to M\"{o}bius-twisted surface states, realizing a M\"{o}bius insulator in an aperiodic 3D system. Furthermore, we show that the quasicrystal with a general magnetic configuration supports multiple HOTI phases characterized by distinct hinge mode configurations that can be switched by rotating the magnetic field. A low-energy effective theory reveals that these transitions are driven by mass kinks between adjacent surfaces. Our work establishes a platform for realizing symmetry-protected topological phases unique to quasicrystals and highlights the tunability of hinge and surface states via magnetic control.

cond-mat.mes-hall

Multimodal-Guided Dynamic Dataset Pruning for Robust and Efficient Data-Centric Learning

Modern deep models are trained on large real-world datasets, where data quality varies and redundancy is common. Data-centric approaches such as dataset pruning have shown promise in improving training efficiency and model performance. However, most existing methods rely on static heuristics or task-specific metrics, limiting their robustness and generalizability across domains. In this work, we introduce a dynamic dataset pruning framework that adaptively selects training samples based on both task-driven difficulty and cross-modality semantic consistency. By incorporating supervision from pretrained multimodal foundation models, our approach captures training dynamics while effectively filtering out uninformative samples. Our work highlights the potential of integrating cross-modality alignment for robust sample selection, advancing data-centric learning toward more efficient and robust practices across application domains.

cs.LG

Long-Range Spin-Orbit-Coupled Magnetoelectricity in Type-II Multiferroic NiI$_2$

Type-II multiferroics, where spin order induces ferroelectricity, exhibit strong magnetoelectric coupling. However, for the typical 2D type-II multiferroic NiI$_2$, the underlying magnetoelectric mechanism remains unclear. Here, applying generalized spin-current model, together with first-principles calculations and a tight-binding approach, we build a comprehensive magnetoelectric model for spin-induced polarization. Such model reveals that the spin-orbit coupling extends its influence to the third-nearest neighbors, whose contribution to polarization rivals that of the first-nearest neighbors. By analyzing the orbital-resolved contributions to polarization, our tight-binding model reveals that the long-range magnetoelectric coupling is enabled by the strong $e_g$-$p$ hopping of NiI$_2$. Monte Carlo simulations further predict a Bloch-type magnetic skyrmion lattice at moderate magnetic fields, accompanied by polar vortex arrays. These findings can guide the discovery and design of strongly magnetoelectric multiferroics.

cond-mat.mtrl-sci

Profiling Apple Silicon Performance for ML Training

Apple Silicon has attracted much attention for its performance and role in machine learning (ML) training. Unlike NVIDIA GPUs, which have traditionally dominated ML training, Apple Silicon has a significant difference in memory architecture. It uses Unified Memory, which integrates CPU and GPU memory instead of separate CPU memory and GPU VRAM. However, it is difficult to tell whether Unified Memory means more performance benefits. This paper investigates the performance differences by training several large language model (LLM) workloads end-to-end under different memory scenarios. The results show a significant performance gap between Apple Silicon and NVIDIA GPUs. This paper attributes this gap to system-level factors such as page faults, power consumption, and kernel launch time. In addition, the performance difference of basic linear algebra subprograms (BLAS) on the NVIDIA GPUs and Apple Silicon chips is analyzed to further explain the observed gap.

cs.PF

WhisperFlow: speech foundation models in real time

Speech foundation models, such as OpenAI's Whisper, become the state of the art in speech understanding due to their strong accuracy and generalizability. Yet, their applications are mostly limited to processing pre-recorded speech, whereas processing of streaming speech, in particular doing it efficiently, remains rudimentary. Behind this inefficiency are multiple fundamental reasons: (1) speech foundation models are trained to process long, fixed-length voice inputs (often 30 seconds); (2) encoding each voice input requires encoding as many as 1,500 tokens with tens of transformer layers; (3) decoding each output entails an irregular, complex beam search. As such, streaming speech processing on resource-constrained client devices is more expensive than other AI tasks, e.g., text generation. To this end, we present a novel framework, WhisperFlow, which embodies both model and system optimizations. (1) Hush word as a short, learnable audio segment; appended to a voice input, a hush word gracefully stops the speech model from processing more input without hallucination; (2) Beam pruning, which aligns streaming audio buffers over time and reuses results from earlier decoding rounds, therefore significantly accelerating decoding; and (3) CPU/GPU pipelining, which not only maps to the encoding/decoding stages dynamically, but also tunes to an optimal resource ratio, respecting the encoding/decoding speed that varies across voice inputs, models, and hardware. We test WhisperFlow on commodity ARM platforms with 4-12 CPU cores and 10-30 GPU cores. It reduces per-word latency by 1.6x-4.7x to as low as 0.5 second, while seeing negligible accuracy degradation. On an entry-level MacBook Air, WhisperFlow can keep the per-word latency around 1 second, with the whole device drawing only 7 Watts in total.

cs.SD

Dual Prototypes for Adaptive Pre-Trained Model in Class-Incremental Learning

Class-incremental learning (CIL) aims to learn new classes while retaining previous knowledge. Although pre-trained model (PTM) based approaches show strong performance, directly fine-tuning PTMs on incremental task streams often causes renewed catastrophic forgetting. This paper proposes a Dual-Prototype Network with Task-wise Adaptation (DPTA) for PTM-based CIL. For each incremental learning task, an adapter module is built to fine-tune the PTM, where the center-adapt loss forces the representation to be more centrally clustered and class separable. The dual prototype network improves the prediction process by enabling test-time adapter selection, where the raw prototypes deduce several possible task indexes of test samples to select suitable adapter modules for PTM, and the augmented prototypes that could separate confusable classes are utilized to determine the final result. Experiments on multiple benchmarks show that DPTA consistently surpasses recent methods by 1\% - 5\%. Notably, on the VTAB dataset, it achieves approximately 3\% improvement over state-of-the-art methods. The code is open-sourced in https://github.com/Yorkxzm/DPTA}

cs.LG

Deep learning density functional theory Hamiltonian in real space

Deep learning electronic structures from ab initio calculations holds great potential to revolutionize computational materials studies. While existing methods proved success in deep-learning density functional theory (DFT) Hamiltonian matrices, they are limited to DFT programs using localized atomic-like bases and heavily depend on the form of the bases. Here, we propose the DeepH-r method for deep-learning DFT Hamiltonians in real space, facilitating the prediction of DFT Hamiltonian in a basis-independent manner. An equivariant neural network architecture for modeling the real-space DFT potential is developed, targeting a more fundamental quantity in DFT. The real-space potential exhibits simplified principles of equivariance and enhanced nearsightedness, further boosting the performance of deep learning. When applied to evaluate the Hamiltonian matrix, this method significantly improved in accuracy, as exemplified in multiple case studies. Given the abundance of data in the real-space potential, this work may pave a novel pathway for establishing a ``large materials model" with increased accuracy.

physics.comp-ph