Searcharxiv⌕ Search

arXiv subjects

Xin Fu

Publications and source records attributed to Xin Fu.

At least 37 records · Page 2Linked to original sources

Fed MobiLLM: Efficient Federated LLM Fine-Tuning over Heterogeneous Mobile Devices via Server Assisted Side-Tuning

Collaboratively fine-tuning (FT) large language models (LLMs) over heterogeneous mobile devices fosters immense potential applications of personalized intelligence. However, such a vision faces critical system challenges. Conventional federated LLM FT approaches place prohibitive computational and memory burdens on mobile hardware, and their synchronous model aggregation protocols stall for slower devices. In this paper, we propose Fed MobiLLM, a novel design to facilitate efficient federated LLM FT across mobile devices with diverse computing/communication speeds and local model architectures. In particular, Fed MobiLLM implements a pioneering server-assisted federated side-tuning paradigm. Briefly, mobile devices perform lightweight forward propagation computations on local data using their frozen pre-scaled backbone LLMs, and then upload selected intermediate activations. The server trains a shared side-network independently, eliminating client-side backpropagation and enabling asynchronous updates. To bridge model heterogeneity across different devices, we introduce an adaptive layer-wise feature alignment method, which ensures consistent representations for collaboratively tuning a shared side network. Extensive experimental results demonstrate that Fed MobiLLM can maintain robust fine-tuning performance while achieving extremely low on-device memory, with at least 95.2% reduction in computation overhead, 93.2% reduction in communication costs and 5.1x faster convergence compared to existing methods, validating its efficacy for practical LLM adaptation over heterogeneous mobile devices.

cs.LG↗

NeuroMoE: A Transformer-Based Mixture-of-Experts Framework for Multi-Modal Neurological Disorder Classification

The integration of multi-modal Magnetic Resonance Imaging (MRI) and clinical data holds great promise for enhancing the diagnosis of neurological disorders (NDs) in real-world clinical settings. Deep Learning (DL) has recently emerged as a powerful tool for extracting meaningful patterns from medical data to aid in diagnosis. However, existing DL approaches struggle to effectively leverage multi-modal MRI and clinical data, leading to suboptimal performance. To address this challenge, we utilize a unique, proprietary multi-modal clinical dataset curated for ND research. Based on this dataset, we propose a novel transformer-based Mixture-of-Experts (MoE) framework for ND classification, leveraging multiple MRI modalities-anatomical (aMRI), Diffusion Tensor Imaging (DTI), and functional (fMRI)-alongside clinical assessments. Our framework employs transformer encoders to capture spatial relationships within volumetric MRI data while utilizing modality-specific experts for targeted feature extraction. A gating mechanism with adaptive fusion dynamically integrates expert outputs, ensuring optimal predictive performance. Comprehensive experiments and comparisons with multiple baselines demonstrate that our multi-modal approach significantly enhances diagnostic accuracy, particularly in distinguishing overlapping disease states. Our framework achieves a validation accuracy of 82.47\%, outperforming baseline methods by over 10\%, highlighting its potential to improve ND diagnosis by applying multi-modal learning to real-world clinical data.

eess.IV↗

RCD structures on singular Kahler spaces of complex dimension three

Let X be a projective variety of complex dimension 3 with log terminal singularities. We prove that every singular Kahler metric on X with bounded Nash entropy and Ricci curvature bounded below induces a compact RCD space homeomorphic to the projective variety X itself. In particular, singular Kahler-Einstein spaces of complex dimension 3 with bounded Nash entropy are compact RCD spaces topologically and holomorphically equivalent to the underlying projective variety. Various compactness theorems are also obtained for 3-dimensional projective varieties with bounded Ricci curvature. Such results establish connections among algebraic, geometric and analytic structures of klt singularities from birational geometry and provide abundant examples of RCD spaces from algebraic geometry via complex Monge-Ampere equations.

math.DG↗

WHALE-FL: Wireless and Heterogeneity Aware Latency Efficient Federated Learning over Mobile Devices via Adaptive Subnetwork Scheduling

As a popular distributed learning paradigm, federated learning (FL) over mobile devices fosters numerous applications, while their practical deployment is hindered by participating devices' computing and communication heterogeneity. Some pioneering research efforts proposed to extract subnetworks from the global model, and assign as large a subnetwork as possible to the device for local training based on its full computing and communications capacity. Although such fixed size subnetwork assignment enables FL training over heterogeneous mobile devices, it is unaware of (i) the dynamic changes of devices' communication and computing conditions and (ii) FL training progress and its dynamic requirements of local training contributions, both of which may cause very long FL training delay. Motivated by those dynamics, in this paper, we develop a wireless and heterogeneity aware latency efficient FL (WHALE-FL) approach to accelerate FL training through adaptive subnetwork scheduling. Instead of sticking to the fixed size subnetwork, WHALE-FL introduces a novel subnetwork selection utility function to capture device and FL training dynamics, and guides the mobile device to adaptively select the subnetwork size for local training based on (a) its computing and communication capacity, (b) its dynamic computing and/or communication conditions, and (c) FL training status and its corresponding requirements for local training contributions. Our evaluation shows that, compared with peer designs, WHALE-FL effectively accelerates FL training without sacrificing learning accuracy.

cs.LG↗

MobiLLM: Enabling LLM Fine-Tuning on the Mobile Device via Server Assisted Side Tuning

Large Language Model (LLM) at mobile devices and its potential applications never fail to fascinate. However, on-device LLM fine-tuning poses great challenges due to extremely high memory requirements and slow training speeds. Even with parameter-efficient fine-tuning (PEFT) methods that update only a small subset of parameters, resource-constrained mobile devices cannot afford them. In this paper, we propose MobiLLM to enable memory-efficient transformer LLM fine-tuning on a mobile device via server-assisted side-tuning. Particularly, MobiLLM allows the resource-constrained mobile device to retain merely a frozen backbone model, while offloading the memory and computation-intensive backpropagation of a trainable side-network to a high-performance server. Unlike existing fine-tuning methods that keep trainable parameters inside the frozen backbone, MobiLLM separates a set of parallel adapters from the backbone to create a backpropagation bypass, involving only one-way activation transfers from the mobile device to the server with low-width quantization during forward propagation. In this way, the data never leaves the mobile device while the device can remove backpropagation through the local backbone model and its forward propagation can be paralyzed with the server-side execution. Thus, MobiLLM preserves data privacy while significantly reducing the memory and computational burdens for LLM fine-tuning. Through extensive experiments, we demonstrate that MobiLLM can enable a resource-constrained mobile device, even a CPU-only one, to fine-tune LLMs and significantly reduce convergence time and memory usage.

cs.LG↗

A continuous cusp closing process for negative Kähler-Einstein metrics

We give an example of a family of smooth complex algebraic surfaces of degree $6$ in $\mathbb{CP}^3$ developing an isolated elliptic singularity. We show via a gluing construction that the unique Kähler-Einstein metrics of Ricci curvature $-1$ on these sextics develop a complex hyperbolic cusp in the limit, and that near the tip of the forming cusp a Tian-Yau gravitational instanton bubbles off.

math.DG↗

Multi-Grained Preference Enhanced Transformer for Multi-Behavior Sequential Recommendation

Sequential recommendation (SR) aims to predict the next purchasing item according to users' dynamic preference learned from their historical user-item interactions. To improve the performance of recommendation, learning dynamic heterogeneous cross-type behavior dependencies is indispensable for recommender system. However, there still exists some challenges in Multi-Behavior Sequential Recommendation (MBSR). On the one hand, existing methods only model heterogeneous multi-behavior dependencies at behavior-level or item-level, and modelling interaction-level dependencies is still a challenge. On the other hand, the dynamic multi-grained behavior-aware preference is hard to capture in interaction sequences, which reflects interaction-aware sequential pattern. To tackle these challenges, we propose a Multi-Grained Preference enhanced Transformer framework (M-GPT). First, M-GPT constructs a interaction-level graph of historical cross-typed interactions in a sequence. Then graph convolution is performed to derive interaction-level multi-behavior dependency representation repeatedly, in which the complex correlation between historical cross-typed interactions at specific orders can be well learned. Secondly, a novel multi-scale transformer architecture equipped with multi-grained user preference extraction is proposed to encode the interaction-aware sequential pattern enhanced by capturing temporal behavior-aware multi-grained preference . Experiments on the real-world datasets indicate that our method M-GPT consistently outperforms various state-of-the-art recommendation methods.

cs.IR↗

Cohomology bases of toric surfaces

Given a compact toric surface, the multiplication of its rational cohomology can be described in terms of the intersection products of Weil divisors, or in terms of the cup products of cohomology classes representing specific cells. In this paper, we aim to compare these two descriptions. More precisely, we define two different cohomology bases, the \emph{Poincaré dual basis} and the \emph{cellular basis}, which give rise to matrices representing the intersection product and the cup product. We prove that these representing matrices are inverse of each other.

math.AT↗

The generalized Kähler Calabi-Yau problem

We formulate an extension of the Calabi conjecture to the setting of generalized Kähler geometry. We show a transgression formula for the Bismut Ricci curvature in this setting, which requires a new local Goto/Kodaira-Spencer deformation result, and use it to show that solutions of the generalized Calabi-Yau equation on compact manifolds are classically Kähler, Calabi-Yau, and furthermore unique in their generalized Kähler class. We show that the generalized Kähler-Ricci flow is naturally adapted to this conjecture, and exhibit a number of a priori estimates and monotonicity formulas which suggest global existence and convergence. For initial data in the generalized Kähler class of a Kähler Calabi-Yau structure we prove the flow exists globally and converges to this unique fixed point. This has applications to understanding the space of generalized Kähler structures, and as a special case yields the topological structure of natural classes of Hamiltonian symplectomorphisms on hyperKähler manifolds. In the case of commuting-type generalized Kähler structures we establish global existence and convergence with arbitrary initial data to a Kähler, Calabi-Yau metric, which yields a new $d d^c$-lemma for these structures.

math.DG↗

Wave packets propagation in the subwavelength regime near the Dirac point

In [Ammari et al., SIAM J Math Anal., 52 (2020), pp. 5441--5466], the first author with collaborators proved the existence of Dirac dispersion cones at subwavelength scales in bubbly honeycomb phononic crystals. In this paper, we study the time-evolution of wave packets that are spectrally concentrated near such conical points. We prove that the wave packets dynamics is governed by a time-dependent effective Dirac system, which still depends, but in a simple way, on the subwavelength scale.

math.AP↗

Path homology of digraphs without multisquares and its comparison with homology of spaces

For a digraph $G$ without multisquares and a field $\mathbb{F}$, we construct a basis of the vector space of path $n$-chains $Ω_n(G;\mathbb{F})$ for $n\geq 0$, generalising the basis of $Ω_3(G;\mathbb{F})$ constructed by Grigory'an. For a field $\mathbb{F},$ we consider the $\mathbb{F}$-path Euler characteristic $χ^\mathbb{F}(G)$ of a digraph $G$ defined as the alternating sum of dimensions of path homology groups with coefficients in $\mathbb{F}.$ If $Ω_\bullet(G;\mathbb{F})$ is a bounded chain complex, the constructed bases can be applied to compute $χ^\mathbb{F}(G)$. We provide an explicit example of a digraph $\mathcal{G}$ whose $\mathbb{F}$-path Euler characteristic depends on whether the characteristic of $\mathbb{F}$ is two, revealing the differences between GLMY theory and the homology theory of spaces. This allows us to prove that there is no topological space $X$ whose homology is isomorphic to path homology of the digraph $H_*(X;\mathbb{K})\cong {\rm PH}_*(\mathcal{G};\mathbb{K})$ simultaneously for $\mathbb{K}=\mathbb{Z}$ and $\mathbb{K}=\mathbb{Z}/2\mathbb{Z}.$

math.AT↗

Convergence rate and uniform Lipschitz estimate in periodic homogenization of high-contrast elliptic systems

We consider the Dirichlet problem for elliptic systems with periodically distributed inclusions whose conduction parameter exhibits a significant contrast compared to the background media. We develop a unified method to quantify the convergence rates both as the periodicity of inclusions tends to zero and as the parameter approaches either zero or infinity. Based on the obtained convergence rates and a Campanato-type scheme, we also derive the regularity estimates that are uniform both in the periodicity and the contrast.

math.AP↗

Tuning Quantum Computing Privacy through Quantum Error Correction

Quantum computing is a promising paradigm for efficiently solving large and high-complexity problems. To protect quantum computing privacy, pioneering research efforts proposed to redefine differential privacy (DP) in quantum computing, i.e., quantum differential privacy (QDP), and harvest inherent noises generated by quantum computing to implement QDP. However, such an implementation approach is limited by the amount of inherent noises, which makes the privacy budget of the QDP mechanism fixed and uncontrollable. To address this issue, in this paper, we propose to leverage quantum error correction (QEC) techniques to reduce quantum computing errors, while tuning the privacy protection levels in QDP. In short, we gradually decrease the quantum noise error rate by deciding whether to apply QEC operations on the gate in a multiple single qubit gates circuit. We have derived a new calculation formula for the general error rate and corresponding privacy budgets after QEC operation. Then, we expand to achieve further noise reduction using multi-level concatenated QEC operation. Through extensive numerical simulations, we demonstrate that QEC is a feasible way to regulate the degree of privacy protection in quantum computing.

quant-ph↗

Ultra-broadband and compact 2$\times$2 3-dB silicon adiabatic coupler based on supermode-injected adjoint shape optimization

The 2$\times$2 3-dB couplers are one of the most widely used and important components in silicon photonics. We propose an ultra-broadband and compact 2$\times$2 3-dB adiabatic coupler defined by b-splines and optimized with an efficient supermode-injected adjoint shape optimization. By employing mode adiabatic evolution and mode coupling at two different wavelength ranges, respectively, we achieve an ultra-broad bandwidth of 530 nm from 1150nm to1680nm with a power imbalance below $\pm$0.76 dB in a compact coupling length of 30 $μm$ according to our simulation results. The supermode-injected adjoint shape optimization can also be applied to the design of other photonic devices based on supermode manipulation.

physics.optics↗

Cohomology of smooth toric varieties: naturality

Building on the recent computation of the cohomology rings of smooth toric varieties and partial quotients of moment-angle complexes, we investigate the naturality properties of the resulting isomorphism between the cohomology of such a space and the torsion product involving the Stanley-Reisner ring. If 2 is invertible in the chosen coefficient ring, then the isomorphism is natural with respect to toric morphisms, which for partial quotients are defined in analogy with toric varieties. In general there are deformation terms that we describe explicitly.

math.AT↗

Accelerating convolutional neural network by exploiting sparsity on GPUs

Convolutional neural network (CNN) is an important deep learning method. The convolution operation takes a large proportion of the total execution time for CNN. Feature maps for convolution operation are usually sparse. Multiplications and additions for zero values in the feature map are useless for convolution results. In addition, the convolution layer and pooling layer are computed separately in traditional methods, which leads to frequent data transfer between CPU and GPU. Based on these observations, we propose two new methods to accelerate CNN on GPUs. The first method focuses on accelerating convolution operation and reducing the calculation of zero values. The second method combines the operations of one convolution layer with the following pooling layer to effectively reduce traffic between CPU and GPU. For the first method, we extract some convolution layers from LeNet, AlexNet, and GoogLeNet, and can achieve up to 3.6X speedup over cuDNN for the single-layer convolution on GPU. Experiment on VGG-19 achieves 3.5X speedup over cuDNN for convolution operation on average. For the second method, the experiment on VGG-19 achieves 4.3X speedup over cuDNN on average.

cs.DC↗

Homogenization of eigenvalues for problems with high-contrast inclusions

We study quantitative homogenization of the eigenvalues for elliptic systems with periodically distributed inclusions, where the conductivity of inclusions are strongly contrast to that of the matrix. We propose a quantitative version of periodic unfolding method, based on this and the recent results concerned on high-contrast homogenization, the convergence rates of eigenvalues are studied for any contrast $δ\in (0,\infty)$.

math.AP↗