Searcharxiv⌕ Search

arXiv subjects

Yu Feng

Publications and source records attributed to Yu Feng.

At least 109 records · Page 6Linked to original sources

Stable parabolic Higgs bundles of rank two and singular hyperbolic metrics

In this paper, we construct a stable parabolic Higgs bundle of rank two, which corresponds to the uniformization associated with a conformal hyperbolic metric on a compact Riemann surface $\overline{X}$ with prescribed singularities. This provides an alternative proof of the classical existence theorem for singular hyperbolic metrics, originally established by Heins ({\it Nagoya Math. J.} 21 (1962), 1-60). We also introduce a family of stable parabolic Higgs bundles of rank two on $\overline{X}$, parametrized by a nonempty open subset of a complex vector space. These bundles correspond to singular hyperbolic metrics with the same type of singularity as the original, but are defined on deformed Riemann surfaces of $\overline{X}$. Thus, we extend partially the final section of Hitchin's celebrated work ({\it Proc. London Math. Soc.} 55(3) (1987), 59-125) to the context of hyperbolic metrics with singularities.

math.DG↗

M-ANT: Efficient Low-bit Group Quantization for LLMs via Mathematically Adaptive Numerical Type

Large language models (LLMs) are one of the most important killer computer applications. The recent algorithmic advancement proposes a fine-grained group-wise quantization for LLMs, which treats a small set (e.g., 64) of values in a tensor as a compression unit. It effectively preserves the model accuracy without retraining, and has become the standard approach to efficiently deploy LLMs. On the other hand, there are works that propose various adaptive data types to better adapt to different distributions and further reduce the required bit length for LLMs. In this work, our detailed analysis unveils a key finding that while different tensors exhibit similar distributions, small groups can have markedly different distributions. As such, the group-level diversity requires a new level of adaptivity for which existing adaptive data types fail to provide. In this paper, we propose MANT, a mathematically adaptive numeric type, featuring a more flexible encoding paradigm with a wider range of data distribution and more efficient decodingcomputation fusion mechanism to address these challenges. Based on MANT, we develop a supporting framework to assign the appropriate data type for each group adaptively. Meanwhile, the dynamically generated Key-Value (KV) caches in LLMs introduce further complexity for real-time quantization. To tackle this, we propose an efficient real-time quantization mechanism. Besides, we implement a specific processing element (PE) to efficiently support MANT and incorporate a real-time quantization unit. By integrating these components into a systolic array, MANT unifies the group-wise weight and KV cache quantization and addresses the associated challenges. Our evaluation shows achieving, on average, 2.99x (up to 4.46x) speedup and 2.81x (up to 4.10x) energy reduction to the state-of-the-art LLM accelerator.

cs.AR↗

Interpreting core forms of urban morphology linked to urban functions with explainable graph neural network

Understanding the high-order relationship between urban form and function is essential for modeling the underlying mechanisms of sustainable urban systems. Nevertheless, it is challenging to establish an accurate data representation for complex urban forms that are readily explicable in human terms. This study proposed the concept of core urban morphology representation and developed an explainable deep learning framework for explicably symbolizing complex urban forms into the novel representation, which we call CoMo. By interpretating the well-trained deep learning model with a stable weighted F1-score of 89.14%, CoMo presents a promising approach for revealing links between urban function and urban form in terms of core urban morphology representation. Using Boston as a study area, we analyzed the core urban forms at the individual-building, block, and neighborhood level that are important to corresponding urban functions. The residential core forms follow a gradual morphological pattern along the urban spine, which is consistent with a center-urban-suburban transition. Furthermore, we prove that urban morphology directly affects land use efficiency, which has a significantly strong correlation with the location (R2=0.721, p<0.001). Overall, CoMo can explicably symbolize urban forms, provide evidence for the classic urban location theory, and offer mechanistic insights for digital twins.

cs.CE↗

PM-MOE: Mixture of Experts on Private Model Parameters for Personalized Federated Learning

Federated learning (FL) has gained widespread attention for its privacy-preserving and collaborative learning capabilities. Due to significant statistical heterogeneity, traditional FL struggles to generalize a shared model across diverse data domains. Personalized federated learning addresses this issue by dividing the model into a globally shared part and a locally private part, with the local model correcting representation biases introduced by the global model. Nevertheless, locally converged parameters more accurately capture domain-specific knowledge, and current methods overlook the potential benefits of these parameters. To address these limitations, we propose PM-MoE architecture. This architecture integrates a mixture of personalized modules and an energy-based personalized modules denoising, enabling each client to select beneficial personalized parameters from other clients. We applied the PM-MoE architecture to nine recent model-split-based personalized federated learning algorithms, achieving performance improvements with minimal additional training. Extensive experiments on six widely adopted datasets and two heterogeneity settings validate the effectiveness of our approach. The source code is available at \url{https://github.com/dannis97500/PM-MOE}.

cs.LG↗

A note On the existence of solutions to Hitchin's self-duality equations

In 1987, Hitchin introduced the self-duality equations on rank-2 complex vector bundles over compact Riemann surfaces with genus greater than one as a reduction of the Yang-Mills equation and established the existence of solutions to these equations starting from a Higgs stable bundle. In this paper, we fill in some technical details in Hitchin's original proof by the following three steps. First, we reduce the existence of a solution of class $L_1^2$ to minimizing the energy functional within a Higgs stable orbit of the $L_2^2$ complex gauge group action. Second, using this transformation, we obtain a solution of class $L_1^2$ in this orbit. These two steps primarily follow Hitchin's original approach. Finally, using the Coulomb gauge, we construct a smooth solution by applying an $L_2^2$ unitary gauge transformation to the $L_1^2$ solution constructed previously. This last step provides additional technical details to Hitchin's original proof.

math.DG↗

Spin correlations in the parent phase of Li$_{1-x}$Fe$_x$ODFeSe

Elucidating spin correlations in the parent compounds of high-temperature superconductors is crucial for understanding superconductivity. We used neutron scattering to study spin correlations in Li$_{1-x}$Fe$_x$ODFeSe, an insulating material with reduced electron carriers compared to its superconducting counterpart ($T_c$ = 41 K), serving as the undoped parent compound. Our findings show a reduced total fluctuating moment in this insulator relative to FeSe and 122 iron pnictides, likely due to increased interlayer distances from intercalation, which enhance fluctuations and reduce the intensity of spin excitations. Moreover, we observed a V-shaped spin wave-like excitation dispersion, contrasting with the twisted hourglass pattern in the superconducting counterpart. Electron doping shifts spin excitation from ($π$, 0) point to an incommensurate position towards ($π$, $π$) direction below 65 meV. This transition from V-shaped to hourglass-like dispersion, akin to behaviors in hole-doped cuprates, suggests a potential shared mechanism in magnetism and superconductivity across these diverse systems.

cond-mat.supr-con↗

Toward Ethical Spatial Analysis: Addressing Endogenous Bias Through Visual Analytics

Spatial analysis can generate both exogenous and endogenous biases, which will lead to ethics issues. Exogenous biases arise from external factors or environments and are unrelated to internal operating mechanisms, while endogenous biases stem from internal processes or technologies. Although much attention has been given to exogenous biases, endogenous biases in spatial analysis have been largely overlooked, and a comprehensive methodology for addressing them is yet to be developed. To tackle this challenge, we propose that visual analytics can play a key role in understanding geographic data and improving the interpretation of analytical results. In this study, we conducted a preliminary investigation using various visualization techniques to explore endogenous biases. Our findings demonstrate the potentials of visual analytics to uncover hidden biases and identify associated issues. Additionally, we synthesized these visualization strategies into a framework that approximates a method for detecting endogenous biases. Through this work, we advocate for the integration of visualization at three critical stages of spatial analysis in order to minimize errors, address ethical concerns, and reduce misinterpretations associated with endogenous biases.

cs.HC↗

Neural Modulation Alteration to Positive and Negative Emotions in Depressed Patients: Insights from fMRI Using Positive/Negative Emotion Atlas

Background: Although it has been noticed that depressed patients show differences in processing emotions, the precise neural modulation mechanisms of positive and negative emotions remain elusive. FMRI is a cutting-edge medical imaging technology renowned for its high spatial resolution and dynamic temporal information, making it particularly suitable for the neural dynamics of depression research. Methods: To address this gap, our study firstly leveraged fMRI to delineate activated regions associated with positive and negative emotions in healthy individuals, resulting in the creation of positive emotion atlas (PEA) and negative emotion atlas (NEA). Subsequently, we examined neuroimaging changes in depression patients using these atlases and evaluated their diagnostic performance based on machine learning. Results: Our findings demonstrate that the classification accuracy of depressed patients based on PEA and NEA exceeded 0.70, a notable improvement compared to the whole-brain atlases. Furthermore, ALFF analysis unveiled significant differences between depressed patients and healthy controls in eight functional clusters during the NEA, focusing on the left cuneus, cingulate gyrus, and superior parietal lobule. In contrast, the PEA revealed more pronounced differences across fifteen clusters, involving the right fusiform gyrus, parahippocampal gyrus, and inferior parietal lobule. Limitations: Due to the limited sample size and subtypes of depressed patients, the efficacy may need further validation in future. Conclusions: These findings emphasize the complex interplay between emotion modulation and depression, showcasing significant alterations in both PEA and NEA among depression patients. This research enhances our understanding of emotion modulation in depression, with implications for diagnosis and treatment evaluation.

cs.CV↗

MetaSapiens: Real-Time Neural Rendering with Efficiency-Aware Pruning and Accelerated Foveated Rendering

Point-Based Neural Rendering (PBNR) is emerging as a promising class of rendering techniques, which are permeating all aspects of society, driven by a growing demand for real-time, photorealistic rendering in AR/VR and digital twins. Achieving real-time PBNR on mobile devices is challenging. This paper proposes MetaSapiens, a PBNR system that for the first time delivers real-time neural rendering on mobile devices while maintaining human visual quality. MetaSapiens combines three techniques. First, we present an efficiency-aware pruning technique to optimize rendering speed. Second, we introduce a Foveated Rendering (FR) method for PBNR, leveraging humans' low visual acuity in peripheral regions to relax rendering quality and improve rendering speed. Finally, we propose an accelerator design for FR, addressing the load imbalance issue in (FR-based) PBNR. Our evaluation shows that our system achieves an order of magnitude speedup over existing PBNR models without sacrificing subjective visual quality, as confirmed by a user study. The code and demo are available at: https://horizon-lab.org/metasapiens/.

cs.GR↗

Human Multi-View Synthesis from a Single-View Model:Transferred Body and Face Representations

Generating multi-view human images from a single view is a complex and significant challenge. Although recent advancements in multi-view object generation have shown impressive results with diffusion models, novel view synthesis for humans remains constrained by the limited availability of 3D human datasets. Consequently, many existing models struggle to produce realistic human body shapes or capture fine-grained facial details accurately. To address these issues, we propose an innovative framework that leverages transferred body and facial representations for multi-view human synthesis. Specifically, we use a single-view model pretrained on a large-scale human dataset to develop a multi-view body representation, aiming to extend the 2D knowledge of the single-view model to a multi-view diffusion model. Additionally, to enhance the model's detail restoration capability, we integrate transferred multimodal facial features into our trained human diffusion model. Experimental evaluations on benchmark datasets demonstrate that our approach outperforms the current state-of-the-art methods, achieving superior performance in multi-view human synthesis.

cs.CV↗

FORAY: Towards Effective Attack Synthesis against Deep Logical Vulnerabilities in DeFi Protocols

Blockchain adoption has surged with the rise of Decentralized Finance (DeFi) applications. However, the significant value of digital assets managed by DeFi protocols makes them prime targets for attacks. Current smart contract vulnerability detection tools struggle with DeFi protocols due to deep logical bugs arising from complex financial interactions between multiple smart contracts. These tools primarily analyze individual contracts and resort to brute-force methods for DeFi protocols crossing numerous smart contracts, leading to inefficiency. We introduce Foray, a highly effective attack synthesis framework against deep logical bugs in DeFi protocols. Foray proposes a novel attack sketch generation and completion framework. Specifically, instead of treating DeFis as regular programs, we design a domain-specific language (DSL) to lift the low-level smart contracts into their high-level financial operations. Based on our DSL, we first compile a given DeFi protocol into a token flow graph, our graphical representation of DeFi protocols. Then, we design an efficient sketch generation method to synthesize attack sketches for a certain attack goal (e.g., price manipulation, arbitrage, etc.). This algorithm strategically identifies candidate sketches by finding reachable paths in TFG, which is much more efficient than random enumeration. For each candidate sketch written in our DSL, Foray designs a domain-specific symbolic compilation to compile it into SMT constraints. Our compilation simplifies the constraints by removing redundant smart contract semantics. It maintains the usability of symbolic compilation, yet scales to problems orders of magnitude larger. Finally, the candidates are completed via existing solvers and are transformed into concrete attacks via direct syntax transformation.

cs.CR↗

Existence and non-uniqueness of cone spherical metrics with prescribed singularities on a compact Riemann surface with positive genus

Cone spherical metrics, defined on compact Riemann surfaces, are conformal metrics with constant curvature one and finitely many cone singularities. Such a metric is termed \textit{reducible} if a developing map of the metric has monodromy in ${\rm U(1)}$, and \textit{irreducible} otherwise. Utilizing the polystable extensions of two line bundles on a compact Riemann surface $X$ with genus $g_X>0$, we establish the following three primary results concerning these metrics with cone angles in $2π{\mathbb Z}_{>1}$: \begin{itemize} \item[(1)] Given an effective divisor $D$ with an odd degree surpassing $2g_X$ on $X$, we find the existence of an effective divisor $D'$ in the complete linear system $|D|$ that can be represented by at least two distinct irreducible cone spherical metrics on $X$. \item[(2)] For a generic effective divisor $D$ with an even degree and $°D\geq 6g_X-2$ on $X$, we can identify an arcwise connected Borel subset in $|D|$ that demonstrates a Hausdorff dimension of no less than $\big(°D-4g_{X}+2\big)$. Within this subset, each divisor $D'$ can be distinctly represented by a family of reducible metrics, defined by a single real parameter. \item[(3)] For an effective divisor $D$ with $°D=2$ on an elliptic curve, we can identify a Borel subset in $|D|$ that is arcwise connected, showcasing a Hausdorff dimension of one. Within this subset, each divisor $D'$ can be distinctly represented by a family of reducible metrics, defined by a single real parameter.

math.DG↗

Potamoi: Accelerating Neural Rendering via a Unified Streaming Architecture

Neural Radiance Field (NeRF) has emerged as a promising alternative for photorealistic rendering. Despite recent algorithmic advancements, achieving real-time performance on today's resource-constrained devices remains challenging. In this paper, we identify the primary bottlenecks in current NeRF algorithms and introduce a unified algorithm-architecture co-design, Potamoi, designed to accommodate various NeRF algorithms. Specifically, we introduce a runtime system featuring a plug-and-play algorithm, SpaRW, which significantly reduces the per-frame computational workload and alleviates compute inefficiencies. Furthermore, our unified streaming pipeline coupled with customized hardware support effectively tames both SRAM and DRAM inefficiencies by minimizing repetitive DRAM access and completely eliminating SRAM bank conflicts. When evaluated against a baseline utilizing a dedicated DNN accelerator, our framework demonstrates a speed-up and energy reduction of 53.1$\times$ and 67.7$\times$, respectively, all while maintaining high visual quality with less than a 1.0 dB reduction in peak signal-to-noise ratio.

cs.AR↗

CP-Prompt: Composition-Based Cross-modal Prompting for Domain-Incremental Continual Learning

The key challenge of cross-modal domain-incremental learning (DIL) is to enable the learning model to continuously learn from novel data with different feature distributions under the same task without forgetting old ones. However, existing top-performing methods still cause high forgetting rates, by lacking intra-domain knowledge extraction and inter-domain common prompting strategy. In this paper, we propose a simple yet effective framework, CP-Prompt, by training limited parameters to instruct a pre-trained model to learn new domains and avoid forgetting existing feature distributions. CP-Prompt captures intra-domain knowledge by compositionally inserting personalized prompts on multi-head self-attention layers and then learns the inter-domain knowledge with a common prompting strategy. CP-Prompt shows superiority compared with state-of-the-art baselines among three widely evaluated DIL tasks. The source code is available at https://github.com/dannis97500/CP_Prompt.

cs.CL↗

vTensor: Flexible Virtual Tensor Management for Efficient LLM Serving

Large Language Models (LLMs) are widely used across various domains, processing millions of daily requests. This surge in demand poses significant challenges in optimizing throughput and latency while keeping costs manageable. The Key-Value (KV) cache, a standard method for retaining previous computations, makes LLM inference highly bounded by memory. While batching strategies can enhance performance, they frequently lead to significant memory fragmentation. Even though cutting-edge systems like vLLM mitigate KV cache fragmentation using paged Attention mechanisms, they still suffer from inefficient memory and computational operations due to the tightly coupled page management and computation kernels. This study introduces the vTensor, an innovative tensor structure for LLM inference based on GPU virtual memory management (VMM). vTensor addresses existing limitations by decoupling computation from memory defragmentation and offering dynamic extensibility. Our framework employs a CPU-GPU heterogeneous approach, ensuring efficient, fragmentation-free memory management while accommodating various computation kernels across different LLM architectures. Experimental results indicate that vTensor achieves an average speedup of 1.86x across different models, with up to 2.42x in multi-turn chat scenarios. Additionally, vTensor provides average speedups of 2.12x and 3.15x in kernel evaluation, reaching up to 3.92x and 3.27x compared to SGLang Triton prefix-prefilling kernels and vLLM paged Attention kernel, respectively. Furthermore, it frees approximately 71.25% (57GB) of memory on the NVIDIA A100 GPU compared to vLLM, enabling more memory-intensive workloads.

cs.DC↗

AutoVCoder: A Systematic Framework for Automated Verilog Code Generation using LLMs

Recently, the use of large language models (LLMs) for software code generation, e.g., C/C++ and Python, has proven a great success. However, LLMs still suffer from low syntactic and functional correctness when it comes to the generation of register-transfer level (RTL) code, such as Verilog. To address this issue, in this paper, we develop AutoVCoder, a systematic open-source framework that significantly improves the LLMs' correctness of generating Verilog code and enhances the quality of its output at the same time. Our framework integrates three novel techniques, including a high-quality hardware dataset generation approach, a two-round LLM fine-tuning method and a domain-specific retrieval-augmented generation (RAG) mechanism. Experimental results demonstrate that AutoVCoder outperforms both industrial and academic LLMs in Verilog code generation. Specifically, AutoVCoder shows a 0.5% and 2.2% improvement in functional correctness on the EvalMachine and EvalHuman benchmarks compared with BetterV, and also achieves a 3.4% increase in syntax correctness and a 3.4% increase in functional correctness on the RTLLM benchmark compared with RTLCoder.

cs.AR↗

Logarithmic correlation functions in 2D critical percolation

It is believed that the large-scale geometric properties of two-dimensional critical percolation are described by a logarithmic conformal field theory, but it has been challenging to exhibit concrete examples of logarithmic singularities and to find an explanation and a physical interpretation, in terms of lattice observables, for their appearance. We show that certain percolation correlation functions receive independent contributions from a large number of similar connectivity events happening at different scales. Combined with scale invariance, this leads to logarithmic divergences. We study several logarithmic correlation functions for critical percolation in the bulk and in the presence of a boundary, including the four-point function of the density (spin) field. Our analysis confirms previous findings, provides new explicit calculations and explains, in terms of lattice observables, the physical mechanism that leads to the logarithmic singularities we discover. Although we adopt conformal field theory (CFT) terminology to present our results, the core of our analysis relies on probabilistic arguments and recent rigorous results on the scaling limit of critical percolation and does not assume a priori the existence of a percolation CFT. As a consequence, our results provide strong support for the validity of a CFT description of critical percolation and a step in the direction of a mathematically rigorous formulation of a logarithmic CFT of two-dimensional critical percolation.

math-ph↗