SearcharxivSearch

arXiv subjects

Zhe Su

Publications and source records attributed to Zhe Su.

At least 19 recordsLinked to original sources

Small-World Communication Fabrics for Neuromorphic Multicore-SoCs

As neuromorphic systems scale beyond a single core, inter-core event communication can become a dominant contributor to memory footprint, latency, and energy consumption. Biological neural systems address a similar scaling challenge through small-world organization, combining dense local connectivity with sparse long-range projections. In this work, we compare two recent multicore neuromorphic systems implemented in the same 22-nm FDSOI technology and explicitly optimized for such connectivity. The first, NeoCorAl, uses an asynchronous packet-switched tree with hierarchical multicast, whereas the second, MOSAIC, employs an RRAM-based, circuit-switched two-dimensional mesh that performs routing in memory. We examine the resulting trade-offs in routing flexibility, hop count, memory requirements, multicast efficiency, and scalability. We further study how the relative efficiency of tree- and mesh-based routing depends on communication locality in spatially-embedded, random, and layered networks. Finally, we discuss routing-aware training as a means of jointly optimizing neural connectivity, task performance, and hardware mappability.

cs.ET

Weighted Hodge Laplacians on Manifolds with Boundary

The spectrum of the Hodge Laplacian on differential manifolds encodes rich topological and geometric information and thus provides a powerful tool for analyzing data on manifolds. However, the classical unweighted formulation is restricted in its ability to study data with varying local features. To address this limitation, we propose a weighted Hodge Laplacian framework for manifolds with boundary, both in theory and in computation, by incorporating a weight function on the manifold. Under appropriate boundary conditions, we formulate the corresponding weighted de Rham-Hodge theory, in which the kernel of the weighted Hodge Laplacian coincides with the weighted harmonic space, and remains isomorphic to the de Rham cohomology of the underlying manifold. The harmonic spectrum of the weighted Hodge Laplacian captures the global topological information, while its non-harmonic spectrum encodes the local geometric property induced by the weight. The proposed framework therefore enables the study of topological and geometric features of data on manifolds across varying weights, and in addition, allows local structure to be highlighted by choosing weights that emphasize regions of interest. We demonstrate the effectiveness of the proposed method through proof-of-principle experiments in protein flexibility analysis, and the results show its promise.

math.DG

Persistent Manifold Learning of Protein Properties

Predicting how tightly two biomolecules bind remains a major challenge, in part because different interaction classes present dissimilar interfaces, from compact metal-coordinated pockets to broad, featureless protein surfaces. We introduce persistent manifold learning (PML), a novel computational framework that describes a binding interface as a family of multiscale manifolds. Boundary-Induced Graph Laplacian, a discrete realization of de Rham-Hodge theory, then extracts topological invariants together with nonharmonic spectral information, capturing the geometry of an interface as well as its topology. These manifold embeddings are combined with protein and molecular language model representations and paired with gradient boosting decision trees. Our PML outperforms state-of-the-art methods on metalloprotein-ligand and protein-protein benchmarks.

q-bio.BM

SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation

Uncertainty estimation is essential not only for the trustworthy deployment of large language models (LLMs) but also as a foundation for self-refinement in LLM generation. However, existing approaches operate at suboptimal granularities: token-level scores lack semantic coherence, while sequence-level scores fail to localize errors. We formalize Span-Level Uncertainty Estimation (SLUE), a new task that targets the natural granularity for uncertainty: semantically coherent text spans, each conveying a single assessable unit of meaning. To address this task, we introduce SPANUQ, a lightweight probe that distills the uncertainty knowledge from expensive multi-sample inference into a single forward pass over LLM hidden states. SPANUQ employs a DETR-style span decoder to simultaneously detect spans and estimate their uncertainty via a Mixture of Beta distribution, trained with a principled combination of Beta NLL regression and contrastive ranking objectives. We construct SPANUQ-BENCH, the first span-level uncertainty benchmark comprising 20K prompts, 293K annotated spans, and continuous soft labels derived from multi-sample claim verification. Experiments on five LLM backbones show that SPANUQ consistently achieves the best span-level uncertainty quality, outperforming the strongest probe baseline and all sampling-based methods while being 10-20x faster. Its DETR-based span detector attains 0.910 F1, surpassing the best heuristic by 39.4%, enabling precise error localization that sequence-level methods cannot provide. The framework generalizes across five LLMs spanning two model families.

cs.CL

A vector field induced de Rham-Hodge theory on manifolds

We introduce a de Rham-Hodge framework induced by a vector field on a compact, oriented smooth manifold. Using a vector field induced bundle isomorphism on differential forms, we define a vector field induced Hodge $L^2$-inner product, codifferential, and Hodge Laplacian. Unlike classical deformations, such as the drifting and Witten-type Hodge Laplacians, the induced Laplacian modifies the principal symbol and gives rise to an anisotropic Laplace-Beltrami type operator on functions. We establish the resulting de Rham-Hodge theory for closed manifolds, including the ellipticity of the induced Hodge Laplacian and the corresponding Hodge decomposition and isomorphism results. We further extend the framework to manifolds with boundary by imposing certain vector field induced boundary conditions, which are necessary to restore the adjointness between the differential and induced codifferential, and to obtain a well-posed boundary value problem. Under these boundary conditions, we establish analogues of the Hodge-Morrey and Friedrichs decompositions. We also discuss several structural properties of the framework, including its relation to anisotropic Laplace-Beltrami operators, its spectral behavior in several explicit examples, and its invariance under isometries.

math.DG

Unsupervised Graph Modeling for Anomaly Detection in Accounting Subject Relationships

This paper addresses the problem of anomaly detection in accounting subject association structures, proposing a structured modeling and unsupervised discriminant framework based on graph neural networks. This framework is used to mine stable correspondences between subjects and identify structural deviations from general ledger details and voucher entries. The method first abstracts accounting subjects as graph nodes, and the co-occurrence and debit/credit correspondence of subjects in the same business record are abstracted as weighted edges. The edge weights are characterized by statistical measures such as co-occurrence frequency or amount aggregation, thus forming a period-level accounting subject association graph. In the representation learning stage, a message passing mechanism is used to fuse the node's own attributes and neighborhood context to obtain node embeddings containing structural information. In the anomaly detection stage, the rationality of subject pair connections is estimated through a relation reconstruction decoder, and edge-level anomaly scores are defined based on the degree of deviation in reconstruction probabilities. These scores are then aggregated to obtain node-level risk ranking and local anomaly localization. This framework can simultaneously capture local substructure anomalies and cross-community anomaly connections without relying on anomaly labeling, outputting traceable subject pair risk clues. Comparative experiments demonstrate more stable comprehensive discriminant capabilities and higher top-ranking accuracy.

cs.LG

GLM-OCR Technical Report

GLM-OCR is an efficient 0.9B-parameter compact multimodal model designed for real-world document understanding. It combines a 0.4B-parameter CogViT visual encoder with a 0.5B-parameter GLM language decoder, achieving a strong balance between computational efficiency and recognition performance. To address the inefficiency of standard autoregressive decoding in deterministic OCR tasks, GLM-OCR introduces a Multi-Token Prediction (MTP) mechanism that predicts multiple tokens per step, significantly improving decoding throughput while keeping memory overhead low through shared parameters. At the system level, a two-stage pipeline is adopted: PP-DocLayout-V3 first performs layout analysis, followed by parallel region-level recognition. Extensive evaluations on public benchmarks and industrial scenarios show that GLM-OCR achieves competitive or state-of-the-art performance in document parsing, text and formula transcription, table structure recovery, and key information extraction. Its compact architecture and structured generation make it suitable for both resource-constrained edge deployment and large-scale production systems.

cs.CL

DendroNN: Dendrocentric Neural Networks for Energy-Efficient Classification of Event-Based Data

Spatiotemporal information is at the core of diverse sensory processing and computational tasks. Feed-forward spiking neural networks can be used to solve these tasks while offering potential benefits in terms of energy efficiency by computing event-based. However, they have trouble decoding temporal information with high accuracy. Thus, they commonly resort to recurrence or delays to enhance their temporal computing ability which, however, bring downsides in terms of hardware-efficiency. In the brain, dendrites are computational powerhouses that just recently started to be acknowledged in such machine learning systems. In this work, we focus on a sequence detection mechanism present in branches of dendrites and translate it into a novel type of neural network by introducing a dendrocentric neural network, DendroNN. DendroNNs identify unique incoming spike sequences as spatiotemporal features. This work further introduces a rewiring phase to train the non-differentiable spike sequences without the use of gradients. During the rewiring, the network memorizes frequently occurring sequences and additionally discards those that do not contribute any discriminative information. The networks display competitive accuracies across various event-based time series datasets. We also propose an asynchronous digital hardware architecture using a time-wheel mechanism that builds on the event-driven design of DendroNNs, eliminating per-step global updates typical of delay- or recurrence-based models. By leveraging a DendroNN's dynamic and static sparsity along with intrinsic quantization, it achieves up to 4x higher efficiency than state-of-the-art neuromorphic hardware at comparable accuracy on the same audio classification task, demonstrating its suitability for spatiotemporal event-based computing. This work offers a novel approach to low-power spatiotemporal processing on event-driven hardware.

cs.LG

ElfCore: A 28nm Neural Processor Enabling Dynamic Structured Sparse Training and Online Self-Supervised Learning with Activity-Dependent Weight Update

In this paper, we present ElfCore, a 28nm digital spiking neural network processor tailored for event-driven sensory signal processing. ElfCore is the first to efficiently integrate: (1) a local online self-supervised learning engine that enables multi-layer temporal learning without labeled inputs; (2) a dynamic structured sparse training engine that supports high-accuracy sparse-to-sparse learning; and (3) an activity-dependent sparse weight update mechanism that selectively updates weights based solely on input activity and network dynamics. Demonstrated on tasks including gesture recognition, speech, and biomedical signal processing, ElfCore outperforms state-of-the-art solutions with up to 16X lower power consumption, 3.8X reduced on-chip memory requirements, and 5.9X greater network capacity efficiency.

cs.AR

Topological surface phonons modulate thermal transport in semiconductor thin films

While phonon topology in crystalline solids has been extensively studied, its influence on thermal transport-especially in nanostructures-remains elusive. Here, by combining first-principles-based machine learning potentials with the phonon Boltzmann transport equation and molecular dynamics simulations, we systematically investigate the role of topological surface phonons in the in-plane thermal transport of semiconductor thin films (Si, 4H -SiC, and c-BN). These topological surface phonons, originating from nontrivial acoustic phonon nodal lines, not only serve as key scattering channels for dominant acoustic phonons but also contribute substantially to the overall thermal conductivity. Remarkably, for these thin semiconductor films below 10 nm this contribution can be as large as over 30% of the in-plane thermal conductivity at 300 K, and the largest absolute contribution can reach 82 W/m-K, highlighting their significant role in nanoscale thermal transport in semiconductors. Furthermore, we demonstrate that both temperature and biaxial strain provide effective means to modulate this contribution. Our work establishes a direct link between topological surface phonons and nanoscale thermal transport, offering the first quantitative assessment of their role and paving the way for topology-enabled thermal management in semiconductors.

cond-mat.mtrl-sci

Algorithm-hardware co-design of neuromorphic networks with dual memory pathways

Spiking neural networks excel at event-driven sensing. Yet, maintaining task-relevant context over long timescales both algorithmically and in hardware, while respecting both tight energy and memory budgets, remains a core challenge in the field. We address this challenge through an algorithm-hardware co-design effort. At the algorithm level, inspired by the cortical fast-slow organization in the brain, we introduce a neural network with an explicit slow memory pathway that, combined with fast spiking activity, enables a dual memory pathway (DMP) architecture in which each layer maintains a compact low-dimensional state that summarizes recent activity and modulates spiking dynamics. This explicit memory stabilizes learning while preserving event-driven sparsity, achieving competitive accuracy on long-sequence benchmarks with 40-60% fewer parameters than equivalent state-of-the-art spiking neural networks. At the hardware level, we introduce a near-memory-compute architecture that fully leverages the advantages of the DMP architecture by retaining its compact shared state while optimizing dataflow, across heterogeneous sparse-spike and dense-memory pathways. We show experimental results that demonstrate more than a 4X increase in throughput and over a 5X improvement in energy efficiency compared with state-of-the-art implementations. Together, these contributions demonstrate that biological principles can guide functional abstractions that are both algorithmically effective and hardware-efficient, establishing a scalable co-design framework for real-time neuromorphic computation and learning.

cs.NE

PlotGen-Bench: Evaluating VLMs on Generating Visualization Code from Diverse Plots across Multiple Libraries

Recent advances in vision-language models (VLMs) have expanded their multimodal code generation capabilities, yet their ability to generate executable visualization code from plots, especially for complex 3D, animated, plot-to-plot transformations, or multi-library scenarios, remains underexplored. To address this gap, we introduce PlotGen-Bench, a comprehensive benchmark for evaluating plot-to-code generation under realistic and complex visualization scenarios. The benchmark spans 9 major categories, 30 subcategories, and 3 core tasks-plot replication, plot transformation, and multi-library generation, covering both 2D, 3D and animated plots across 5 widely used visualization libraries. Through systematic evaluation of state-of-the-art open- and closed-source VLMs, we find that open-source models still lag considerably behind in visual fidelity and semantic consistency, despite achieving comparable code executability. Moreover, all models exhibit substantial degradation on reasoning-intensive tasks such as chart type conversion and animation generation. PlotGen-Bench establishes a rigorous foundation for advancing research toward more capable and reliable VLMs for visualization authoring and code synthesis, with all data and code available at https://plotgen.github.io.

cs.HC

Exploiting heterogeneous delays for efficient computation in low-bit neural networks

Neural networks rely on learning synaptic weights. However, this overlooks other neural parameters that can also be learned and may be utilized by the brain. One such parameter is the delay: the brain exhibits complex temporal dynamics with heterogeneous delays, where signals are transmitted asynchronously between neurons. It has been theorized that this delay heterogeneity, rather than a cost to be minimized, can be exploited in embodied contexts where task-relevant information naturally sits contextually in the time domain. We test this hypothesis by training spiking neural networks to modify not only their weights but also their delays at different levels of precision. We find that delay heterogeneity enables state-of-the-art performance on temporally complex neuromorphic problems and can be achieved even when weights are extremely imprecise (1.58-bit ternary precision: just positive, negative, or absent). By enabling high performance with extremely low-precision weights, delay heterogeneity allows memory-efficient solutions that maintain state-of-the-art accuracy even when weights are compressed over an order of magnitude more aggressively than typically studied weight-only networks. We show how delays and time-constants adaptively trade-off, and reveal through ablation that task performance depends on task-appropriate delay distributions, with temporally-complex tasks requiring longer delays. Our results suggest temporal heterogeneity is an important principle for efficient computation, particularly when task-relevant information is temporal - as in the physical world - with implications for embodied intelligent systems and neuromorphic hardware.

cs.NE

Glyph: Scaling Context Windows via Visual-Text Compression

Large language models (LLMs) increasingly rely on long-context modeling for tasks such as document understanding, code analysis, and multi-step reasoning. However, scaling context windows to the million-token level brings prohibitive computational and memory costs, limiting the practicality of long-context LLMs. In this work, we take a different perspective-visual context scaling-to tackle this challenge. Instead of extending token-based sequences, we propose Glyph, a framework that renders long texts into images and processes them with vision-language models (VLMs). This approach substantially compresses textual input while preserving semantic information, and we further design an LLM-driven genetic search to identify optimal visual rendering configurations for balancing accuracy and compression. Through extensive experiments, we demonstrate that our method achieves 3-4x token compression while maintaining accuracy comparable to leading LLMs such as Qwen3-8B on various long-context benchmarks. This compression also leads to around 4x faster prefilling and decoding, and approximately 2x faster SFT training. Furthermore, under extreme compression, a 128K-context VLM could scale to handle 1M-token-level text tasks. In addition, the rendered text data benefits real-world multimodal tasks, such as document understanding. Our code and model are released at https://github.com/thu-coai/Glyph.

cs.CV

Topological Data Analysis and Topological Deep Learning Beyond Persistent Homology -- A Review

Topological data analysis (TDA) is a rapidly evolving field in applied mathematics and data science that leverages tools from topology to uncover robust, shape-driven insights in complex datasets. The main workhorse is persistent homology, a technique rooted in algebraic topology. Paired with topological deep learning (TDL) or topological machine learning, persistent homology has achieved tremendous success in a wide variety of applications in science, engineering, medicine, and industry. However, persistent homology has many limitations due to its high-level abstraction, insensitivity to non-topological changes, and reliance on point cloud data. This paper presents a comprehensive review of TDA and TDL beyond persistent homology. It analyzes how persistent topological Laplacians and Dirac operators provide spectral representations to capture both topological invariants and homotopic evolution. Other formulations are presented in terms of sheaf theory, Mayer topology, and interaction topology. For data on differentiable manifolds, techniques rooted in differential topology, such as persistent de Rham cohomology, persistent Hodge Laplacian, and Hodge decomposition, are reviewed. For one-dimensional (1D) curves embedded in 3-space, approaches from geometric topology are discussed, including multiscale Gauss-link integrals, persistent Jones polynomials, and persistent Khovanov homology. This paper further discusses the appropriate selection of topological tools for different input data, such as point clouds, sequential data, data on manifolds, curves embedded in 3-space, and data with additional non-geometric information. A review is also given of various topological representations, software packages, and machine learning vectorizations. Finally, this review ends with concluding remarks.

math.HO

GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

We present GLM-4.1V-Thinking, GLM-4.5V, and GLM-4.6V, a family of vision-language models (VLMs) designed to advance general-purpose multimodal understanding and reasoning. In this report, we share our key findings in the development of the reasoning-centric training framework. We first develop a capable vision foundation model with significant potential through large-scale pre-training, which arguably sets the upper bound for the final performance. We then propose Reinforcement Learning with Curriculum Sampling (RLCS) to unlock the full potential of the model, leading to comprehensive capability enhancement across a diverse range of tasks, including STEM problem solving, video understanding, content recognition, coding, grounding, GUI-based agents, and long document interpretation. In a comprehensive evaluation across 42 public benchmarks, GLM-4.5V achieves state-of-the-art performance on nearly all tasks among open-source models of similar size, and demonstrates competitive or even superior results compared to closed-source models such as Gemini-2.5-Flash on challenging tasks including Coding and GUI Agents. Meanwhile, the smaller GLM-4.1V-9B-Thinking remains highly competitive-achieving superior results to the much larger Qwen2.5-VL-72B on 29 benchmarks. We open-source both GLM-4.1V-9B-Thinking and GLM-4.5V. We further introduce the GLM-4.6V series, open-source multimodal models with native tool use and a 128K context window. A brief overview is available at https://z.ai/blog/glm-4.6v. Code, models and more information are released at https://github.com/zai-org/GLM-V.

cs.CV

Exploring Big Five Personality and AI Capability Effects in LLM-Simulated Negotiation Dialogues

This paper presents an evaluation framework for agentic AI systems in mission-critical negotiation contexts, addressing the need for AI agents that can adapt to diverse human operators and stakeholders. Using Sotopia as a simulation testbed, we present two experiments that systematically evaluated how personality traits and AI agent characteristics influence LLM-simulated social negotiation outcomes--a capability essential for a variety of applications involving cross-team coordination and civil-military interactions. Experiment 1 employs causal discovery methods to measure how personality traits impact price bargaining negotiations, through which we found that Agreeableness and Extraversion significantly affect believability, goal achievement, and knowledge acquisition outcomes. Sociocognitive lexical measures extracted from team communications detected fine-grained differences in agents' empathic communication, moral foundations, and opinion patterns, providing actionable insights for agentic AI systems that must operate reliably in high-stakes operational scenarios. Experiment 2 evaluates human-AI job negotiations by manipulating both simulated human personality and AI system characteristics, specifically transparency, competence, adaptability, demonstrating how AI agent trustworthiness impact mission effectiveness. These findings establish a repeatable evaluation methodology for experimenting with AI agent reliability across diverse operator personalities and human-agent team dynamics, directly supporting operational requirements for reliable AI systems. Our work advances the evaluation of agentic AI workflows by moving beyond standard performance metrics to incorporate social dynamics essential for mission success in complex operations.

cs.AI

SOTOPIA-S4: a user-friendly system for flexible, customizable, and large-scale social simulation

Social simulation through large language model (LLM) agents is a promising approach to explore and validate hypotheses related to social science questions and LLM agents behavior. We present SOTOPIA-S4, a fast, flexible, and scalable social simulation system that addresses the technical barriers of current frameworks while enabling practitioners to generate multi-turn and multi-party LLM-based interactions with customizable evaluation metrics for hypothesis testing. SOTOPIA-S4 comes as a pip package that contains a simulation engine, an API server with flexible RESTful APIs for simulation management, and a web interface that enables both technical and non-technical users to design, run, and analyze simulations without programming. We demonstrate the usefulness of SOTOPIA-S4 with two use cases involving dyadic hiring negotiation and multi-party planning scenarios.

cs.CY