SearcharxivSearch

arXiv subjects

Junming Zhang

Publications and source records attributed to Junming Zhang.

At least 19 recordsLinked to original sources

EStream: Fast and Memory-Efficient MoE Prefill through Expert Virtualization on Mobile NPUs

Mobile vendors and application developers increasingly deploy LLMs on smartphones for diverse prefill-only services. Yet current systems rely mainly on dense models whose regular computation maps efficiently to mobile NPUs, leaving more capable MoEs underused. MoE prefill does not fit mobile NPUs: NPU graphs are fixed at compile time, yet MoE picks experts at runtime; and one request touches most experts, more than a phone can hold in memory. We present EStream, which resolves both by separating what the NPU must fix from what MoE decides at runtime. A single compiled expert graph serves every expert, with each expert's routed tokens and weight address bound at call time, so dynamic MoE execution runs entirely on the NPU without padding or CPU/GPU fallback. Expert virtualization keeps the expert pool in UFS flash storage and pages it through a fixed-size NPU-addressable arena, group by group, with loading hidden behind computation, so memory is bounded by the arena rather than by the model. It further introduces a hardware-aware configuration algorithm that automatically configures the UFS--NPU pipeline and maximizes loading--computation overlap. Across 18 comparative settings covering three 7B--16B MoEs and 256--4,096-token prompts, we evaluate EStream on a commercial Snapdragon smartphone. Compared to the fastest baseline at each setting, EStream achieves a 2.25--27.57X pure-prefill TTFT speedup and reduces peak physical memory by 1.19--12.29X. EStream further scales to MoE models with up to 46.7B parameters.

cs.DC

An Ultrathin Laterally Conductive Mesh Interphase Enables Spatially Extended Zinc Deposition for Aqueous Zinc Batteries

Zn metal anodes often suffer from nonuniform interfacial reactions during cycling, resulting in uneven deposition and dendrite growth. Existing artificial interphases can mitigate side reactions or regulate nucleation, but rarely achieve regulation of the interfacial electron/field distribution to sustain uniform deposition at the evolving Zn/electrolyte interface. Here, we develop an Au mesh interphase (AuMI), an ultrathin two-dimensional Au aerogel network that couples lateral electron redistribution with open pathways for ions. During Zn plating/stripping, the conductive AuMI distributes electron transport across the Zn surface, while its porous mesh preserves Zn$^{2+}$ access, enabling more uniform interfacial reactions. Experiments and simulations show that AuMI homogenizes the interfacial electric field and current distribution, promotes more uniform Zn plating/stripping, and limits dendrite growth. As a result, this regulated interfacial reaction mode enables AuMI Zn symmetric cells to operate stably for 3000 h at 1 mA cm$^{-2}$/1 mAh cm$^{-2}$ and for 1100 h at 10 mA cm$^{-2}$/10 mAh cm$^{-2}$, while AuMI Zn||NVO (NaV$_3$O$_8$\cdot$1.5H$_2$O) full cells retain 80.8 % capacity after over 5000 cycles at 1 A g$^{-1}$. These findings highlight the importance of combining ultrathin architecture, lateral electron transport, and open Zn2+ access in artificial interphases for stable aqueous Zn metal anodes.

cond-mat.mtrl-sci

Curvature-Aware Zeroth-Order Optimization for Memory-Efficient Test-Time Adaptation

Test-time adaptation (TTA) aims to enhance the cross-domain performance of pre-trained models by adapting to unlabeled test data. While most existing TTA methods rely on backpropagation (BP) for finetuning, BP-free methods such as zeroth-order (ZO) methods are more desired in practical on-device scenarios. ZO methods rely only on forward computation, which can largely reduce the complexity and memory overhead of on-device deployment. However, ZO methods suffer from much higher variance compared with first-order methods in estimating the gradient. To address this, we propose an improved ZO method to substantially boost the performance of ZO optimization based TTA. First, we provide an observation to reveal the persistent low-rank Hessian structure of the loss during the adaptation process. Based on this insight, we then propose a loss-landscape curvature-aware zeroth-order (CAZO) method, which leverages a sliding-average estimation of the diagonal Hessian to construct a covariance matrix for anisotropic perturbation sampling. CAZO operates by freezing pretrained weights and optimizing minimal adapter parameters via forward-only passes based gradient estimation, which can substantially reduce the memory overhead compared to BP-based methods. Extensive experiments demonstrate that CAZO significantly outperforms existing TTA methods, achieving state-of-the-art performance while maintaining an excellent balance between accuracy and memory efficiency. Code is available at https://github.com/Hollyming/CAZO.

cs.CV

Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independently. Memory Decoder introduces a parametric long-term memory module but only studies it at a relatively small scale. In this work, we present Memory Decoder at Scale, scaling memory models up to 6.9B parameters and pretraining them on 300B tokens. At this data scale, the combined cost of indexing and search makes a standard Faiss pipeline infeasible. We address this bottleneck with a distributed pipeline for Faiss indexing and retrieval, together with sparse, batch-wise loading of kNN distributions. Across model scales, we find that allocating more parameters to memory yields a better parameter-performance tradeoff than scaling the base model alone. On 17 benchmarks, pairing a 6.9B general memory with Pythia-410M raises its average score from 29.86 to 37.34, surpassing Pythia-12B (37.24) with 39% fewer total parameters. For Qwen3 Base models ranging from 0.6B to 14B, 1.7B domain memories improve the average score across the three domains by more than 9 points at every scale. Overall, our results demonstrate that independently scaling pretrained memory offers a more parameter efficient path to improving language model performance.

cs.CL

Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards

Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. Existing external rewards provide either sparse outcome supervision or richer feedback from process annotations and LLM judges. Outcome rewards scale readily but cannot distinguish grounded retrieval from redundant search, whereas richer signals require costly annotation or inference during training. Internal rewards based on policy-side signals such as entropy, likelihood, or information gain are graded and inexpensive to evaluate, yet mainly reflect model confidence rather than evidence grounding. We propose Search-G1, a representation-based intrinsic reward framework that measures the operational grounding of an agent's answers through two intervention-calibrated readouts. A prompt-state readout predicts closed-book sufficiency, whose complement defines policy-relative retrieval necessity; an answer-commit readout estimates evidence reliance from answer-stage sensitivity to evidence deletion. Together, they provide additional credit to correct searched trajectories when retrieval is estimated necessary and the answer is evidence-sensitive, favor correct direct answers when closed-book knowledge suffices, and penalize repeated search. After calibration, reward scoring requires neither process annotations nor LLM-as-judge inference during policy optimization. Because reinforcement learning changes policy representations, Search-G1 periodically refits both readouts on trajectories from the latest checkpoint, allowing the reward to co-evolve with the policy. Experiments across multiple search-based question-answering benchmarks and two model scales show that Search-G1 improves the grounding--search-cost trade-off, producing shorter response-side trajectories at competitive task accuracy. Code is available at https://github.com/Rosy0912/Search-G1.

cs.CL

Real-space identification of distinct magnetic configurations in a candidate d-wave altermagnet

Altermagnetism is an emerging class of magnetic order characterized by momentum-dependent spin-split electronic structures despite vanishing net magnetization. Although momentum-space signatures consistent with altermagnetism have been reported in a growing number of materials, their relationship to the underlying real-space magnetic configurations remains incompletely understood, because similar spin-split electronic structures can arise from distinct magnetic orders. In the candidate d-wave altermagnet KV2Se2O, the magnetic origin of the observed momentum-dependent spin splitting has remained controversial. Here, we employ spin-polarized scanning tunnelling microscopy combined with magnetic-field-dependent quasiparticle interference imaging to determine the magnetic configuration of KV2Se2O at the atomic scale. Spin-resolved quasiparticle interference reveals a checkerboard-like antiparallel spin texture within the V2O layer and determines its interlayer spin arrangement across unit-cell step edges. Remarkably, we identify both C-type and G-type magnetic configurations, both of which generate similar spin-split electronic structures at the single-layer level but correspond to d-wave altermagnetic and conventional antiferromagnetic orders, respectively. These observations reveal a complex magnetic landscape arising from nearly degenerate magnetic states. Our results establish a direct connection between momentum-space spin splitting and real-space magnetic order, providing a framework for identifying the microscopic origin of spin-split electronic structures in altermagnetic materials.

cond-mat.mtrl-sci

PepALD: Macrocyclic Peptide Generation via Autoregressive Latent Diffusion

Macrocyclic peptides are promising therapeutic candidates for intracellular targets, but their design requires simultaneous control over non-natural monomer chemistry, ring topology, membrane permeability, and target binding. Existing SMILES- or HELM-string generative models either operate in long atom-level sequence spaces or treat monomers as symbolic tokens with limited chemical grounding. We introduce PepALD, an Autoregressive Latent Diffusion (ALD) foundation model for \textit{de novo} macrocyclic peptide generation. The model represents HELM monomers with structured chemical embeddings, generates each residue through context-conditioned diffusion in chemically informed latent space, predicts R-group-aware ring closures during autoregressive generation, and aligns the denoiser to affinity rewards using winner-protected diffusion-adapted preference optimization. In silico experiments demonstrate PepALD's generation quality and reward-optimization performance against representative peptide generation baselines.

cs.LG

Infinitesimal Rigidity of Cyclic Surfaces and Alternating Surfaces

We study the infinitesimal rigidity of equivariant minimal maps from the universal cover of a smooth oriented surface (possibly non-compact) into a Riemannian symmetric space, focusing on representations arising from cyclic harmonic bundles. By developing a unified Lie-theoretic framework that connects cyclic surfaces and cyclic harmonic bundles over Riemann surfaces, we prove the infinitesimal rigidity for irreducible cyclic surfaces under admissible smooth variations, including both compactly supported deformations and $L^p$-integrable variations on non-compact surfaces. As a geometric application, we introduce $n$-alternating surfaces in $\mathbb H^{p,q}$ and establish their correspondence with a special class of cyclic surfaces. This yields an infinitesimal rigidity theorem that conceptually unifies and extends known rigidity results for maximal space-like surfaces, alternating holomorphic curves, and $A$-surfaces in certain $\mathbb H^{p,q}$.

math.DG

M2I2HA: Multi-modal Object Detection Based on Intra- and Inter-Modal Hypergraph Attention

Recent advances in multi-modal detection have significantly improved detection accuracy in challenging environments (e.g., low light, overexposure). By integrating RGB with modalities such as thermal and depth, multi-modal fusion increases data redundancy and system robustness. However, significant challenges remain in effectively extracting task-relevant information both within and across modalities, as well as in achieving precise cross-modal alignment. While CNNs excel at feature extraction, they are limited by constrained receptive fields, strong inductive biases, and difficulty in capturing long-range dependencies. Transformer-based models offer global context but suffer from quadratic computational complexity and are confined to pairwise correlation modeling. Mamba and other State Space Models (SSMs), on the other hand, are hindered by their sequential scanning mechanism, which flattens 2D spatial structures into 1D sequences, disrupting topological relationships and limiting the modeling of complex higher-order dependencies. To address these issues, we propose a multi-modal perception network based on hypergraph theory called M2I2HA. Our architecture includes an Intra-Hypergraph Enhancement module to capture global many-to-many high-order relationships within each modality, and an Inter-Hypergraph Fusion module to align, enhance, and fuse cross-modal features by bridging configuration and spatial gaps between data sources. We further introduce a M2-FullPAD module to enable adaptive multi-level fusion of multi-modal enhanced features within the network, meanwhile enhancing data distribution and flow across the architecture. Extensive object detection experiments on multiple public datasets against baselines demonstrate that M2I2HA achieves state-of-the-art performance in multi-modal object detection tasks.

cs.CV

SwiftKV: An Edge-Oriented Attention Algorithm and Multi-Head Accelerator for Fast, Efficient LLM Decoding

Edge acceleration for large language models is crucial for their widespread application; however, achieving fast attention inference and efficient decoding on resource-constrained edge accelerators remains challenging. This paper presents SwiftKV Attention, a per-token pipelined, low-latency single-pass attention inference algorithm, where every (kt, vt) in the KV cache is processed exactly once in a uniform per-token pipeline without score materialization, blockwise softmax, or a second pass, thereby enabling fast execution on edge accelerators with a single hardware set and no resource-intensive parallelism. Furthermore, to address the limited support for multi-head LLM decoding in existing accelerators, we design the SwiftKV-MHA accelerator, which enables high precision attention and low precision GEMV on the same processor array, achieving fast and efficient multi-head parallel decoding. Experimental results show that, on the edge accelerator, the SwiftKV Attention algorithm achieves a 7.16* speedup over native attention and significantly outperforms other attention algorithms. SwiftKV-MHA further reduces attention latency by 13.48*; under the same settings, it improves generation speed by 17.4% and increases token efficiency by 1.98* compared with state-of-the-art works.

cs.AR

TAO-Net: Two-stage Adaptive OOD Classification Network for Fine-grained Encrypted Traffic Classification

Encrypted traffic classification aims to identify applications or services by analyzing network traffic data. One of the critical challenges is the continuous emergence of new applications, which generates Out-of-Distribution (OOD) traffic patterns that deviate from known categories and are not well represented by predefined models. Current approaches rely on predefined categories, which limits their effectiveness in handling unknown traffic types. Although some methods mitigate this limitation by simply classifying unknown traffic into a single "Other" category, they fail to make a fine-grained classification. In this paper, we propose a Two-stage Adaptive OOD classification Network (TAO-Net) that achieves accurate classification for both In-Distribution (ID) and OOD encrypted traffic. The method incorporates an innovative two-stage design: the first stage employs a hybrid OOD detection mechanism that integrates transformer-based inter-layer transformation smoothness and feature analysis to effectively distinguish between ID and OOD traffic, while the second stage leverages large language models with a novel semantic-enhanced prompt strategy to transform OOD traffic classification into a generation task, enabling flexible fine-grained classification without relying on predefined labels. Experiments on three datasets demonstrate that TAO-Net achieves 96.81-97.70% macro-precision and 96.77-97.68% macro-F1, outperforming previous methods that only reach 44.73-86.30% macro-precision, particularly in identifying emerging network applications.

cs.LG

Design and Control of a Perching Drone Inspired by the Prey-Capturing Mechanism of Venus Flytrap

The endurance and energy efficiency of drones remain critical challenges in their design and operation. To extend mission duration, numerous studies explored perching mechanisms that enable drones to conserve energy by temporarily suspending flight. This paper presents a new perching drone that utilizes an active flexible perching mechanism inspired by the rapid predation mechanism of the Venus flytrap, achieving perching in less than 100 ms. The proposed system is designed for high-speed adaptability to the perching targets. The overall drone design is outlined, followed by the development and validation of the biomimetic perching structure. To enhance the system stability, a cascade extended high-gain observer (EHGO) based control method is developed, which can estimate and compensate for the external disturbance in real time. The experimental results demonstrate the adaptability of the perching structure and the superiority of the cascaded EHGO in resisting wind and perching disturbances.

cs.RO

From Bits to Boardrooms: A Cutting-Edge Multi-Agent LLM Framework for Business Excellence

Large Language Models (LLMs) have shown promising potential in business applications, particularly in enterprise decision support and strategic planning, yet current approaches often struggle to reconcile intricate operational analyses with overarching strategic goals across diverse market environments, leading to fragmented workflows and reduced collaboration across organizational levels. This paper introduces BusiAgent, a novel multi-agent framework leveraging LLMs for advanced decision-making in complex corporate environments. BusiAgent integrates three core innovations: an extended Continuous Time Markov Decision Process (CTMDP) for dynamic agent modeling, a generalized entropy measure to optimize collaborative efficiency, and a multi-level Stackelberg game to handle hierarchical decision processes. Additionally, contextual Thompson sampling is employed for prompt optimization, supported by a comprehensive quality assurance system to mitigate errors. Extensive empirical evaluations across diverse business scenarios validate BusiAgent's efficacy, demonstrating its capacity to generate coherent, client-focused solutions that smoothly integrate granular insights with high-level strategy, significantly outperforming established approaches in both solution quality and user satisfaction. By fusing cutting-edge AI technologies with deep business insights, BusiAgent marks a substantial step forward in AI-driven enterprise decision-making, empowering organizations to navigate complex business landscapes more effectively.

cs.AI

ADSEL: Adaptive dual self-expression learning for EEG feature selection via incomplete multi-dimensional emotional tagging

EEG based multi-dimension emotion recognition has attracted substantial research interest in human computer interfaces. However, the high dimensionality of EEG features, coupled with limited sample sizes, frequently leads to classifier overfitting and high computational complexity. Feature selection constitutes a critical strategy for mitigating these challenges. Most existing EEG feature selection methods assume complete multi-dimensional emotion labels. In practice, open acquisition environment, and the inherent subjectivity of emotion perception often result in incomplete label data, which can compromise model generalization. Additionally, existing feature selection methods for handling incomplete multi-dimensional labels primarily focus on correlations among various dimensions during label recovery, neglecting the correlation between samples in the label space and their interaction with various dimensions. To address these issues, we propose a novel incomplete multi-dimensional feature selection algorithm for EEG-based emotion recognition. The proposed method integrates an adaptive dual self-expression learning (ADSEL) with least squares regression. ADSEL establishes a bidirectional pathway between sample-level and dimension-level self-expression learning processes within the label space. It could facilitate the cross-sharing of learned information between these processes, enabling the simultaneous exploitation of effective information across both samples and dimensions for label reconstruction. Consequently, ADSEL could enhances label recovery accuracy and effectively identifies the optimal EEG feature subset for multi-dimensional emotion recognition.

cs.HC

FDC-Net: Rethinking the association between EEG artifact removal and multi-dimensional affective computing

Electroencephalogram (EEG)-based emotion recognition holds significant value in affective computing and brain-computer interfaces. However, in practical applications, EEG recordings are susceptible to the effects of various physiological artifacts. Current approaches typically treat denoising and emotion recognition as independent tasks using cascaded architectures, which not only leads to error accumulation, but also fails to exploit potential synergies between these tasks. Moreover, conventional EEG-based emotion recognition models often rely on the idealized assumption of "perfectly denoised data", lacking a systematic design for noise robustness. To address these challenges, a novel framework that deeply couples denoising and emotion recognition tasks is proposed for end-to-end noise-robust emotion recognition, termed as Feedback-Driven Collaborative Network for Denoising-Classification Nexus (FDC-Net). Our primary innovation lies in establishing a dynamic collaborative mechanism between artifact removal and emotion recognition through: (1) bidirectional gradient propagation with joint optimization strategies; (2) a gated attention mechanism integrated with frequency-adaptive Transformer using learnable band-position encoding. Two most popular EEG-based emotion datasets (DEAP and DREAMER) with multi-dimensional emotional labels were employed to compare the artifact removal and emotion recognition performance between FDC-Net and nine state-of-the-art methods. In terms of the denoising task, FDC-Net obtains a maximum correlation coefficient (CC) value of 96.30% on DEAP and a maximum CC value of 90.31% on DREAMER. In terms of the emotion recognition task under physiological artifact interference, FDC-Net achieves emotion recognition accuracies of 82.3+7.1% on DEAP and 88.1+0.8% on DREAMER.

cs.HC

Charge-polarized superconducting state emerging in a superatomic antipolar metal

The simultaneous presence of polarity and metallicity or superconductivity in a material signifies the exotic polar metallic or superconducting (SC) state, while such materials are extremely rare due to their exclusive nature. Recently, the interweaved CDW and antipolar charge orders have been discovered in a metallic superatomic crystal of Au6Te12Se8 (ATS), while their interplay and competition with the following emergent SC state remains elusive. Here, we report a further experimental investigation of the SC state emerged from the preformed CDW and antipolar order states using scanning tunneling microscopy/spectroscopy in combination with transport and Raman measurements. The temperature-dependent pre-formation and condensation of Cooper pairs are experimentally identified. The pre-existent CDW is gradually suppressed by the preformed Cooper pairs, and then the antipolar charge order is spatially suppressed into a ferrielectric-like polar order by the condensed Cooper pairs of SC state. The exotic charge-polarized superconducting state is discovered in the polar metal of ATS, suggesting a valuable platform for the exploration of intriguing polar superconducting properties.

cond-mat.supr-con

A simplified method for full-wave simulation of metamaterials: utilizing near-field decoupling technology

Simulating the electromagnetic properties of large-scale, complex metamaterial structures demands significant time and memory resources. If these large-scale structures can be divided into smaller, simpler components, the overall cost of studying all the smaller structures could be much lower than directly simulating the entire structure. Unfortunately, decoupling complex structures has been challenging due to the unclear mechanisms of near-field coupling in metamaterials. In this paper, we identify that the key to understanding near-field coupling in metamaterials lies in evanescent wave interactions, which can be captured through full-wave simulations. Our findings suggest that by accounting for the influence of evanescent waves, it becomes possible to analytically decouple and then recouple structures, even when the types of metamaterial structures vary. Building on this insight, we successfully decomposed complex structures into multiple groups of simpler components. By studying these simpler components, the electromagnetic properties of the entire structure can be calculated analytically. This decoupling method dramatically reduces the computation time or memory required for research into the electromagnetic properties of metamaterials.

physics.optics

Non-maximal Anosov representations from surface groups to $\mathrm{SO}_0(2,3)$

We prove the representation given by a stable $\alpha_1$-cyclic parabolic $\mathrm{SO}_0(2,3)$-Higgs bundle through the non-Abelian Hodge correspondence is $\{\alpha_2\}$-almost dominated. This is a generalization of Filip's result on weight $3$ variation of Hodge structures and answers a question asked by Collier, Tholozan and Toulisse.

math.DG