Searcharxiv⌕ Search

arXiv subjects

Runze Wang

Publications and source records attributed to Runze Wang.

At least 37 records · Page 2Linked to original sources

Pollard's theorem in general abelian groups

We make further progress towards a Kneser-type generalization of Pollard's Theorem to general abelian groups. For two sets $A$ and $B$ in an abelian group $G$, the \emph{$t$-popular sumset} of $A$ and $B$, denoted by $A+_t B$, is the set of elements in $G$ each with at least $t$ representations of the form $a+b$, where $a\in A$ and $b\in B$. For $|A|,\, |B|\ge t\geq 2$, we prove that if \begin{align*} \sum_{i=1}^t |A+_i B|< t|A|+t|B|-\frac{4}{3}t^2+\frac{2}{3}t, \end{align*} then there exist $A'\subseteq A$ and $B'\subseteq B$ with $|A\setminus A'|+|B\setminus B'|\le t-1$, $A'+_t B'=A'+B'=A+_t B$, and $ \sum_{i=1}^t |A+_i B|\ge t|A|+t|B|-t|H|,$ where $H$ is the stabilizer of $A'+B'=A+_t B$. Our result improves the main quadratic term in the previous best bound from $-2t^2$ to $-\frac{4}{3}t^2$.

math.NT↗

Geometric wavefront sets of genuine Iwahori-spherical representations

For Iwahori-spherical genuine representations of central covers with positive real Satake parameters, we prove the upper bound inequality for their geometric wavefront sets, formulated for general genuine representations in an earlier work by Gao--Liu--Lo--Shahidi. Meanwhile, we show the equality is attained for covers of type A groups and for some representations of covers of the exceptional groups. We also verify the equality for certain Iwahori-spherical representations occurring in regular unramified principal series; this uses and generalizes the earlier work of Karasiewicz--Okada--Wang on theta representations. Lastly, we determine the leading coefficients in the Harish-Chandra character expansion of a theta representation when its geometric wavefront set is of a special type.

math.RT↗

Online-PVLM: Advancing Personalized VLMs with Online Concept Learning

Personalized Visual Language Models (VLMs) are gaining increasing attention for their formidable ability in user-specific concepts aligned interactions (e.g., identifying a user's bike). Existing methods typically require the learning of separate embeddings for each new concept, which fails to support real-time adaptation during testing. This limitation becomes particularly pronounced in large-scale scenarios, where efficient retrieval of concept embeddings is not achievable. To alleviate this gap, we propose Online-PVLM, a framework for online concept learning by leveraging hyperbolic representations. Our approach makes a train-free paradigm for concept embeddings generation at test time, making the use of personalized VLMs both scalable and efficient. In addition, we develop OP-Eval, a comprehensive and large-scale benchmark comprising 1,292 concepts and over 30K high-quality instances with diverse question types, designed to rigorously assess online concept learning in realistic scenarios. Extensive experiments demonstrate the state-of-the-art performance of our proposed framework. Our source code and dataset will be made available.

cs.CL↗

Strong hub cover pebbling number

In a graph $G$, we define a set of vertices to be a \emph{strong hub set} if for any two vertices in $G$, we can find a path between them whose internal vertices are all in this set. We define the \emph{strong hub cover pebbling number} of $G$, denoted by $h_s^*(G)$, to be the smallest $t$ such that for any initial configuration with $t$ pebbles on $G$, we can make some pebbling moves (a pebbling move consists of removing two pebbles from a vertex $v$ and adding one pebble to another vertex adjacent to $v$) so that there is a strong hub set with every vertex in it having a pebble. We determine the strong hub cover pebbling numbers of paths, stars, and books.

math.CO↗

Discrete isoperimetric inequalities on the strong products of paths

For a graph $G=(V,\ E)$ and a nonempty set $S\subseteq V$, the \emph{vertex boundary} of $S$, denoted by $\partial_G(S)$, is defined to be the set of vertices that are not in $S$ but have at least one neighbor in $S$. In this paper, for $G$ being a strong product of two paths, we determine the cases in which $|\partial_G(S)|$ is minimized.

math.CO↗

On Relative Ordered Turán Density

For an ordered graph $F$, denote the Turán density by $\vecπ(F)$. The relative Turán density, denoted by $ρ(F)$, is the supremum over $α\in [0,1]$ such that every ordered graph $G$ contains an $F$-free subgraph $G'$ with $e(G') \geq αe(G)$. Reiher, Rödl, Sales and Schacht showed that $ρ(P) = \vecπ(P)/2$ and $ρ(K) = \vecπ(K)$ for any ascending path $P$ or clique $K$. They asked if there are any ordered graphs $F$ with $\vecπ(F)/2 < ρ(F) < \vecπ(F)$. We answer this question in the affirmative by describing a family of such $F$. We also show that the relative Turán densities of a large family of ordered matchings (including $\{\{1,6\}, \{2,3\}, \{4,5\}\}$ and $\{\{1,3\}, \{2,5\}, \{4,6\}\}$) are $0$.

math.CO↗

Strong edge-coloring of graphs with maximum edge weight seven

A strong edge-coloring of a graph $G$ is an edge-coloring such that any two edges of distance at most two receive distinct colors. The minimum number of colors we need in order to give $G$ a strong edge-coloring is called the strong chromatic index of $G$, denoted by $χ_s'(G)$. The maximum edge weight of $G$ is defined to be $\max\{d(u)+d(v):\ uv\in E(G)\}$. In this paper, using the discharging method, we prove that if $G$ is a graph with maximum edge weight $7$ and maximum average degree less than $\frac{40}{13}$, then $χ_s'(G)\le 13$. Also, we determine the largest possible maximum average degree of a graph with given maximum edge weight.

math.CO↗

Anatomy-Aware Low-Dose CT Denoising via Pretrained Vision Models and Semantic-Guided Contrastive Learning

To reduce radiation exposure and improve the diagnostic efficacy of low-dose computed tomography (LDCT), numerous deep learning-based denoising methods have been developed to mitigate noise and artifacts. However, most of these approaches ignore the anatomical semantics of human tissues, which may potentially result in suboptimal denoising outcomes. To address this problem, we propose ALDEN, an anatomy-aware LDCT denoising method that integrates semantic features of pretrained vision models (PVMs) with adversarial and contrastive learning. Specifically, we introduce an anatomy-aware discriminator that dynamically fuses hierarchical semantic features from reference normal-dose CT (NDCT) via cross-attention mechanisms, enabling tissue-specific realism evaluation in the discriminator. In addition, we propose a semantic-guided contrastive learning module that enforces anatomical consistency by contrasting PVM-derived features from LDCT, denoised CT and NDCT, preserving tissue-specific patterns through positive pairs and suppressing artifacts via dual negative pairs. Extensive experiments conducted on two LDCT denoising datasets reveal that ALDEN achieves the state-of-the-art performance, offering superior anatomy preservation and substantially reducing over-smoothing issue of previous work. Further validation on a downstream multi-organ segmentation task (encompassing 117 anatomical structures) affirms the model's ability to maintain anatomical awareness.

eess.IV↗

New developments on graph sum index

In a graph, we assign distinct integers to the vertices, and take the sum of two integers if they are on two adjacent vertices. The minimum possible number of different sums is the \emph{sum index} of this graph. In this paper, we present some new developments on graph sum index. First, we explain the connections between graph sum index and results in additive combinatorics. Then, we determine the sum indices of the complete multipartite graphs, hypercubes, and some cluster graphs. Also, we study the maximum number of edges in a graph with a fixed sum index, which is related to the forbidden subgraph problem.

math.CO↗

ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives

Bridging the gap between ego-centric and exo-centric views has been a long-standing question in computer vision. In this paper, we focus on the emerging Ego-Exo object correspondence task, which aims to understand object relations across ego-exo perspectives through segmentation. While numerous segmentation models have been proposed, most operate on a single image (view), making them impractical for cross-view scenarios. PSALM, a recently proposed segmentation method, stands out as a notable exception with its demonstrated zero-shot ability on this task. However, due to the drastic viewpoint change between ego and exo, PSALM fails to accurately locate and segment objects, especially in complex backgrounds or when object appearances change significantly. To address these issues, we propose ObjectRelator, a novel approach featuring two key modules: Multimodal Condition Fusion (MCFuse) and SSL-based Cross-View Object Alignment (XObjAlign). MCFuse introduces language as an additional cue, integrating both visual masks and textual descriptions to improve object localization and prevent incorrect associations. XObjAlign enforces cross-view consistency through self-supervised alignment, enhancing robustness to object appearance variations. Extensive experiments demonstrate ObjectRelator's effectiveness on the large-scale Ego-Exo4D benchmark and HANDAL-X (an adapted dataset for cross-view segmentation) with state-of-the-art performance. Code is made available at: http://yuqianfu.com/ObjectRelator.

cs.CV↗

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025

In this report, we present a cross-view multi-modal object segmentation approach for the object correspondence task in the Ego-Exo4D Correspondence Challenges 2025. Given object queries from one perspective (e.g., ego view), the goal is to predict the corresponding object masks in another perspective (e.g., exo view). To tackle this task, we propose a multimodal condition fusion module that enhances object localization by leveraging both visual masks and textual descriptions as segmentation conditions. Furthermore, to address the visual domain gap between ego and exo views, we introduce a cross-view object alignment module that enforces object-level consistency across perspectives, thereby improving the model's robustness to viewpoint changes. Our proposed method ranked second on the leaderboard of the large-scale Ego-Exo4D object correspondence benchmark. Code will be made available at https://github.com/lovelyqian/ObjectRelator.

cs.CV↗

PGPO: Enhancing Agent Reasoning via Pseudocode-style Planning Guided Preference Optimization

Large Language Model (LLM) agents have demonstrated impressive capabilities in handling complex interactive problems. Existing LLM agents mainly generate natural language plans to guide reasoning, which is verbose and inefficient. NL plans are also tailored to specific tasks and restrict agents' ability to generalize across similar tasks. To this end, we explore pseudocode-style plans (P-code Plan) to capture the structural logic of reasoning. We find that P-code Plan empowers LLM agents with stronger generalization ability and more efficiency. Inspired by this finding, we propose a pseudocode-style Planning Guided Preference Optimization method called PGPO for effective agent learning. With two planning-oriented rewards, PGPO further enhances LLM agents' ability to generate high-quality P-code Plans and subsequent reasoning. Experiments show that PGPO achieves superior performance on representative agent benchmarks and outperforms the current leading baselines. Analyses reveal the advantage of PGPO in reducing action errors and omissions during reasoning.

cs.AI↗

Application based Evaluation of an Efficient Spike-Encoder, "Spiketrum"

Spike-based encoders represent information as sequences of spikes or pulses, which are transmitted between neurons. A prevailing consensus suggests that spike-based approaches demonstrate exceptional capabilities in capturing the temporal dynamics of neural activity and have the potential to provide energy-efficient solutions for low-power applications. The Spiketrum encoder efficiently compresses input data using spike trains or code sets (for non-spiking applications) and is adaptable to both hardware and software implementations, with lossless signal reconstruction capability. The paper proposes and assesses Spiketrum's hardware, evaluating its output under varying spike rates and its classification performance with popular spiking and non-spiking classifiers, and also assessing the quality of information compression and hardware resource utilization. The paper extensively benchmarks both Spiketrum hardware and its software counterpart against state-of-the-art, biologically-plausible encoders. The evaluations encompass benchmarking criteria, including classification accuracy, training speed, and sparsity when using encoder outputs in pattern recognition and classification with both spiking and non-spiking classifiers. Additionally, they consider encoded output entropy and hardware resource utilization and power consumption of the hardware version of the encoders. Results demonstrate Spiketrum's superiority in most benchmarking criteria, making it a promising choice for various applications. It efficiently utilizes hardware resources with low power consumption, achieving high classification accuracy. This work also emphasizes the potential of encoders in spike-based processing to improve the efficiency and performance of neural computing systems.

eess.SP↗

A General-Purpose Neuromorphic Sensor based on Spiketrum Algorithm: Hardware Details and Real-life Applications

Spiking Neural Networks (SNNs) offer a biologically inspired computational paradigm, enabling energy-efficient data processing through spike-based information transmission. Despite notable advancements in hardware for SNNs, spike encoding has largely remained software-dependent, limiting efficiency. This paper addresses the need for adaptable and resource-efficient spike encoding hardware by presenting an area-optimized hardware implementation of the Spiketrum algorithm, which encodes time-varying analogue signals into spatiotemporal spike patterns. Unlike earlier performance-optimized designs, which prioritize speed, our approach focuses on reducing hardware footprint, achieving a 52% reduction in Block RAMs (BRAMs), 31% fewer Digital Signal Processing (DSP) slices, and a 6% decrease in Look-Up Tables (LUTs). The proposed implementation has been verified on an FPGA and successfully integrated into an IC using TSMC180 technology. Experimental results demonstrate the system's effectiveness in real-world applications, including sound and ECG classification. This work highlights the trade-offs between performance and resource efficiency, offering a flexible, scalable solution for neuromorphic systems in power-sensitive applications like cochlear implants and neural devices.

eess.SP↗

Bridging Molecular Graphs and Large Language Models

While Large Language Models (LLMs) have shown exceptional generalization capabilities, their ability to process graph data, such as molecular structures, remains limited. To bridge this gap, this paper proposes Graph2Token, an efficient solution that aligns graph tokens to LLM tokens. The key idea is to represent a graph token with the LLM token vocabulary, without fine-tuning the LLM backbone. To achieve this goal, we first construct a molecule-text paired dataset from multisources, including CHEBI and HMDB, to train a graph structure encoder, which reduces the distance between graphs and texts representations in the feature space. Then, we propose a novel alignment strategy that associates a graph token with LLM tokens. To further unleash the potential of LLMs, we collect molecular IUPAC name identifiers, which are incorporated into the LLM prompts. By aligning molecular graphs as special tokens, we can activate LLM generalization ability to molecular few-shot learning. Extensive experiments on molecular classification and regression tasks demonstrate the effectiveness of our proposed Graph2Token.

cs.LG↗

A biased edge coloring game

We combine the ideas of edge coloring games and asymmetric graph coloring games and define the \emph{$(m,1)$-edge coloring game}, which is alternatively played by two players Maker and Breaker on a finite simple graph $G$ with a set of colors $X$. Maker plays first and colors $m$ uncolored edges on each turn. Breaker colors only one uncolored edge on each turn. They make sure that adjacent edges get distinct colors. Maker wins if eventually every edge is colored; Breaker wins if at some point, the player who is playing cannot color any edge. We define the \emph{$(m,1)$-game chromatic index} of $G$ to be the smallest nonnegative integer $k$ such that Maker has a winning strategy with $|X|=k$. We give some general upper bounds on the $(m,1)$-game chromatic indices of trees, determine the exact $(m,1)$-game chromatic indices of some caterpillars and all wheels, and show that larger $m$ does not necessarily give us smaller $(m,1)$-game chromatic index.

math.CO↗

The stable wave front set of theta representations

We compute the stable wave front set of theta representations for certain tame Brylinski-Deligne covers of a connected reductive $p$-adic group. The computation involves two main inputs. First we use a theorem of Okada, adapted to covering groups, to reduce the computation of the wave front set to computing the Kawanaka wave front set of certain representations of finite groups of Lie type. Second, to compute the Kawanaka wave front sets we use Lusztig's formula. This requires a careful analysis of the action of the pro-$p$ Iwahori-Hecke algebra on the theta representation, using the structural results about Hecke algebras developed by Gao-Gurevich-Karasiewicz and Wang.

math.RT↗

Threshold numbers of some graphs

A graph $G=(V,E)$ is called a \emph{$k$-threshold graph} with \emph{thresholds} $θ_1<θ_2<...<θ_k$ if we can assign a real number $r(v)$ to each vertex $v\in V$, such that for any $u,v\in V$, we have $uv\in E$ if and only if $r(u)+r(v)\ge θ_i$ holds true for an odd number of elements in $\{θ_1,θ_2,...,θ_k\}$. The smallest integer $k$ such that $G$ is a $k$-threshold graph is called the \emph{threshold number} of $G$. For the complete multipartite graphs and the cluster graphs, Kittipassorn and Sumalroj determined the exact threshold numbers of $K_{n\times 3}$ and $nK_3$. In this paper, first we determine the threshold numbers of some path-related graphs, including linear forests, ladders, and tents. Then, on the basis of Kittipassorn and Sumalroj's results, we determine the exact threshold numbers of $K_{n_1\times 1, n_2\times 2, n_3\times 3}$ and $n_1 K_1\cup n_2 K_2\cup n_3 K_3$, which solve a problem proposed by Sumalroj.

math.CO↗