SearcharxivSearch

arXiv subjects

Fan Wei

Publications and source records attributed to Fan Wei.

At least 19 recordsLinked to original sources

Ground, Cover, and Refine: Evidence-Centric Frame Selection for Long-Video Question Answering

Long-video question answering requires identifying sparse yet critical evidence from videos containing thousands of frames under a constrained visual-token budget. Existing methods either select query-aware frames in a single pass or rely on timestamped text solely as retrieval guidance, leading to two key limitations. First, selected frames tend to cluster around local relevance peaks, and once the budget is exhausted, omitted evidence cannot be recovered. Second, textual and visual evidence remain weakly aligned. We propose GCR, a training-free framework that casts fixed-budget frame selection as a joint evidence curation problem. Ground converts timestamped text into temporal events, selects query-relevant real frame anchors, and renders each event text onto its temporally aligned frame. Cover supplements grounded events with direct visual anchors for complementary visual evidence and applies global maximal marginal relevance to preserve diverse context. Refine revisits omitted temporal regions and replaces the weakest revisable context frame with a real-frame medoid---but only when the medoid offers greater evidence value. GCR maintains a fixed number of chronologically ordered frames and requires no VLM training or architectural modification. Experiments on LongVideoBench and Video-MME, across three 7B backbones and frame budgets of 8, 32, and 64, demonstrate consistent improvements in long-video QA. With the 7B LLaVA-OV backbone and 32 frames, GCR achieves 64.25% and 62.15% on the two benchmarks, outperforming the strongest reproduced baselines by 2.54 and 1.93 percentage points, respectively.

cs.CV

Hallucination Rates in Language Generation

Language generation in the limit is an elegant model introduced by Kleinberg and Mullainathan [KM24] to formally study language generation by an algorithm that learns solely based on example strings. In this model, an algorithm is said to correctly generate from a language if it never makes an error after some finite time. In contrast, even sophisticated language models are known to regularly hallucinate in practice. In this paper, we initiate the study of language generation in the limit with (infinite) hallucination, i.e., the algorithm may generate incorrect strings infinitely often, but the errors occur at a limited rate (possibly even with 0-measure). We first show that hallucination, even at rate 0, makes generation in the limit strictly more powerful: there are language collections that cannot be generated with finite error but can be generated with infinite error, even when errors occur on a 0-measure set of time-steps. Furthermore, while all countable collections are generatable with finite error, we show a strict hierarchy of (uncountable) language collections characterized by the hallucination rate. This hierarchy extends to breadth, the fraction of the target language generated. While all countable collections can attain the optimal breadth of 1/2 [KW26b], we show strict separation at every breadth and hallucination rate. Finally, we study generation in the limit without repetition, where the algorithm may not repeat strings. This lets us compare the sets of correct and incorrect strings generated, rather than the fractions of correct and incorrect time-steps. Once again, we demonstrate a strict hierarchy at every hallucination rate and breadth. Taken together, these results reveal rich structure in language collections generatable in the limit with hallucination and establish hallucination rate as an important parameter in the theoretical study of language generation.

cs.DS

Attend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM Inference

Multimodal large language models (MLLMs) increasingly process long visual-token sequences, increasing the overall inference computation. Existing acceleration methods usually remove visual tokens or skip visual-token updates in entire layers, but these coarse strategies may discard fine-grained evidence or suppress useful operators together with redundant ones. In this paper, we study visual-token computation from an answer-observable perspective and find that late visual-token updates can remain large while having little effect on answer-token representations. Motivated by this answer-silent redundancy, we decompose each Transformer layer into attention and FFN operators and show that useful visual computation is often operator-dominant and layer-dependent. We propose an operator-level visual-token skipping framework that preserves the full visual-token sequence while selectively bypassing redundant attention, FFN, or both. Experiments across three MLLM architectures and 10 VQA benchmarks show that our method achieves strong efficiency-accuracy trade-offs, reducing \textbf{33.7\%} TFLOPs on Qwen3-VL while retaining \textbf{99.5\%} of the vanilla model performance.

cs.CV

Learning to Balance: Decoupled Siamese Diffusion Transformer for Reference-Based Remote Sensing Image Super-Resolution

Diffusion-based methods demonstrate significant potential for remote sensing image super-resolution at large scaling factors, particularly in reference-based super-resolution (RefSR), where high-resolution reference images provide critical fine-grained texture priors. However, existing methods often suffer from a trade-off between over-reliance on reference information, which leads to texture artifacts, and under-utilization of such information, which results in insufficient detail recovery. To address these issues, we propose DS-DiT, a Decoupled Siamese Diffusion Transformer that decouples the interaction between low-resolution (LR) and reference (Ref) conditions within the attention mechanism. By allowing LR structural priors and Ref texture information to independently interact with the noisy latent, the framework effectively mitigates competition between the two conditional sources. To further compensate for the limited local modeling ability of global attention, we introduce a Patch-Level Weighting (PLW) module that adaptively modulates the fusion of conditional sources. In addition, the siamese architecture enables an inference-time autoguidance strategy that exploits the prediction discrepancy between strong and weak Ref conditions to improve generation quality without additional training. Experimental results across multiple datasets and scaling factors show that DS-DiT outperforms existing methods in both quantitative metrics and visual fidelity.

cs.CV

Transforming the Use of Earth Observation Data: Exascale Training of a Generative Compression Model with Historical Priors for up to 10,000x Data Reduction

Earth observation is becoming one of the largest data-producing activities in science, yet current pipelines still treat compression as a storage and transmission tool rather than a new way to use data. We present a generative compression framework that learns from historical Earth observation archives and enables on-demand 100x to 10,000x data reduction across downstream tasks. Unlike general visual data, Earth observation repeatedly measures the same evolving planet, making historical-prior learning feasible for extreme compression. To realize this paradigm, we train large generative compression models at exascale on the LineShine Armv9 CPU supercomputer, with co-optimization across model design, kernels, memory hierarchy, runtime, and parallelism. Our implementation sustains 1.54 EFLOP/s and peaks at 2.16 EFLOP/s in end-to-end training. This work shows that historical-prior generative compression can turn Earth observation data into an active, task-adaptive foundation for acquisition, delivery, storage, and scientific use.

cs.DC

A Characterization for Spectral and Trace Positivity of Matrix Words

In this paper, we provide a complete structural characterization of which words in noncommuting matrix variables and their formal transposes universally evaluate to matrices with nonnegative eigenvalues or nonnegative trace, regardless of the dimension or entries of the substituted matrices. Crucially, we establish that in this pure monomial setting, both spectral and trace positivity exhibit an exact rigidity phenomenon analogous to Hermitian-square representations. This contrasts sharply with the broader polynomial setting, where tracial rigidity is known to fail even for approximation. Furthermore, we establish that this structural rigidity extends seamlessly to the complex and infinite-dimensional domains. We show that the exact same purely combinatorial conditions completely characterize matrix words evaluated over complex matrices, bounded operators on arbitrary Hilbert spaces, and positive tracial states on $C^*$-algebras and von Neumann algebras. As one application of our main theorems, we fully resolve open questions initially posed by Lieb and Pierce, and later conjectured by Hillar and Johnson, concerning the underlying algebraic structure of the summands in Bessis-Moussa-Villani (BMV) trace polynomials from quantum statistical mechanics. We achieve these results by establishing a surprising connection to graph theory, leading to a matrix analogue of the positive graph conjecture posed by Camarena, Cs\'oka, Hubai, Lippner, and Lov\'asz.

math.CO

Validity, Sparse Holes, and Breadth in Language Generation: Banach Density, Topology, and Geometry

Language generation in the limit, rooted in work of Gold and Angluin and revived by Kleinberg and Mullainathan, studies generation under minimal assumptions: an adversary enumerates strings from an unknown target language $K$, and an algorithm must eventually generate unseen strings from $K$. A central question is the tension between validity and breadth: can a generator avoid hallucination while covering the hidden language broadly? In our previous work, we proved that every countable language class admits the optimal 1/2 lower-asymptotic-density guarantee. This prefix-based measure makes validity and breadth universally compatible. Here we replace prefix averaging by the local requirement of lower Banach density, which examines every long interval or box. We prove that for some countable language collections, every eventually valid generator must leave arbitrarily large sparse holes. Thus a worst-case local form of mode collapse is unavoidable: sparse holes are forced by the requirement of valid generation itself. This local notion reveals structure hidden by prefix averaging. In one dimension, the problem is intrinsically topological and combinatorial: finite Cantor--Bendixson rank guarantees the optimal lower Banach density 1/2, whereas infinite-rank classes can force density zero. In dimensions $d \geq 2$, Ramsey/discrepancy geometry adds in to force ordinary Banach density zero even for a singleton language class. We introduce filtered lower Banach density to remove this geometric obstruction and prove a $1/2-\epsilon$ guarantee for finite-rank classes. We also establish a dichotomy for $f$-window densities, interpolating between lower asymptotic and lower Banach density.

cs.DM

RXNRECer Enables Fine-grained Enzymatic Function Annotation through Active Learning and Protein Language Models

A key challenge in enzyme annotation is identifying the biochemical reactions catalyzed by proteins. Most existing methods rely on Enzyme Commission (EC) numbers as intermediaries: they first predict an EC number and then retrieve the associated reactions. This indirect strategy introduces ambiguity due to the complex many-to-many mappings among proteins, EC numbers, and reactions, and is further complicated by frequent updates to EC numbers and inconsistencies across databases. To address these challenges, we present RXNRECer, a transformer-based ensemble framework that directly predicts enzyme-catalyzed reactions without relying on EC numbers. It integrates protein language modeling and active learning to capture both high-level sequence semantics and fine-grained transformation patterns. Evaluations on curated cross-validation and temporal test sets demonstrate consistent improvements over six EC-based baselines, with gains of 16.54% in F1 score and 15.43% in accuracy. Beyond accuracy gains, the framework offers clear advantages for downstream applications, including scalable proteome-wide reaction annotation, enhanced specificity in refining generic reaction schemas, systematic annotation of previously uncurated proteins, and reliable identification of enzyme promiscuity. By incorporating large language models, it also provides interpretable rationales for predictions. These capabilities make RXNRECer a robust and versatile solution for EC-free, fine-grained enzyme function prediction, with potential applications across multiple areas of enzyme research and industrial applications.

cs.LG

RS-Prune: Training-Free Data Pruning at High Ratios for Efficient Remote Sensing Diffusion Foundation Models

Diffusion-based remote sensing (RS) generative foundation models are cruial for downstream tasks. However, these models rely on large amounts of globally representative data, which often contain redundancy, noise, and class imbalance, reducing training efficiency and preventing convergence. Existing RS diffusion foundation models typically aggregate multiple classification datasets or apply simplistic deduplication, overlooking the distributional requirements of generation modeling and the heterogeneity of RS imagery. To address these limitations, we propose a training-free, two-stage data pruning approach that quickly select a high-quality subset under high pruning ratios, enabling a preliminary foundation model to converge rapidly and serve as a versatile backbone for generation, downstream fine-tuning, and other applications. Our method jointly considers local information content with global scene-level diversity and representativeness. First, an entropy-based criterion efficiently removes low-information samples. Next, leveraging RS scene classification datasets as reference benchmarks, we perform scene-aware clustering with stratified sampling to improve clustering effectiveness while reducing computational costs on large-scale unlabeled data. Finally, by balancing cluster-level uniformity and sample representativeness, the method enables fine-grained selection under high pruning ratios while preserving overall diversity and representativeness. Experiments show that, even after pruning 85\% of the training data, our method significantly improves convergence and generation quality. Furthermore, diffusion foundation models trained with our method consistently achieve state-of-the-art performance across downstream tasks, including super-resolution and semantic image synthesis. This data pruning paradigm offers practical guidance for developing RS generative foundation models.

cs.CV

On the growth rate of the Stanley-Wilf limit of blockable permutations

Given a permutation $\pi$, let $\text{Av}_n(\pi)$ be the number of permutations of length $n$ that avoid $\pi$ as a subpermutation. The celebrated resolution of the Stanley-Wilf conjecture by Marcus and Tardos confirmed that the limit $L(\pi) = \lim_{n \to \infty} |\text{Av}_n(\pi)|^{1/n}$ exists. A central and challenging question concerns the behavior of $L(\pi)$ as a function of the pattern length $|\pi|$. While Fox proved that $L(\pi)$ is exponential in $|\pi|$ for almost all permutations, it is known that $L(\pi)$ grows polynomially for specific structural classes. For instance, $L(\pi)$ is known to be quadratic in $|\pi|$ when $\pi$ is a monotone or a layered permutation. In this paper, we address this question for {\it blockable} permutations $\pi$.

math.CO

New Sidorenko-type inequalities in tournaments

As a directed analog of Sidorenko's conjecture in extremal graph theory, Fox, Himwich, Zhou, and the second author defined an oriented graph $H$ to be tournament Sidorenko (anti-Sidorenko) if the random tournament asymptotically minimizes (maximizes) the number of copies of $H$ among all tournaments. We prove new inequalities of this form for oriented trees and cycles, considering both local and global notions of the Sidorenko property. We make progress on a conjecture of the aforementioned authors that every tree has an anti-Sidorenko direction, and give a characterization of short paths. For long paths we show that orientations are split symmetrically between being locally Sidorenko and anti-Sidorenko, yet almost all orientations are not globally Sidorenko. Finally, we give algorithms characterizing the local Sidorenko status of paths and cycles when the number of vertices is not divisible by four.

math.CO

Language Generation and Identification From Partial Enumeration: Tight Density Bounds and Topological Characterizations

The success of large language models (LLMs) has motivated formal theories of language generation and learning. We study the framework of \emph{language generation in the limit}, where an adversary enumerates strings from an unknown language $K$ drawn from a countable class, and an algorithm must generate unseen strings from $K$. Prior work showed that generation is always possible, and that some algorithms achieve positive lower density, revealing a \emph{validity--breadth} trade-off between correctness and coverage. We resolve a main open question in this line, proving a tight bound of $1/2$ on the best achievable lower density. We then strengthen the model to allow \emph{partial enumeration}, where the adversary reveals only an infinite subset $C \subseteq K$. We show that generation in the limit remains achievable, and if $C$ has lower density $\alpha$ in $K$, the algorithm's output achieves density at least $\alpha/2$, matching the upper bound. This generalizes the $1/2$ bound to the partial-information setting, where the generator must recover within a factor $1/2$ of the revealed subset's density. We further revisit the classical Gold--Angluin model of \emph{language identification} under partial enumeration. We characterize when identification in the limit is possible -- when hypotheses $M_t$ eventually satisfy $C \subseteq M \subseteq K$ -- and in the process give a new topological formulation of Angluin's characterization, showing that her condition is precisely equivalent to an appropriate topological space having the $T_D$ separation property.

cs.DS

An Improved Physically-Based Surface Triangulation Method

This paper proposes improvements to the physically-based surface triangulation method, bubble meshing. The method simulates physical bubbles to automatically generate mesh vertices, resulting in high-quality Delaunay triangles. Despite its flexibility in local mesh size control and the advantage of local re-meshing, bubble meshing is constrained by high computational costs and slow convergence on complex surfaces. The proposed approach employs conformal mapping to simplify surface bubble packing by flattening the surface onto a plane. Surface triangulation is induced from the planar mesh, avoiding direct bubble movement on the surface. Optimizing bubble quantity control and separating it from the relaxation process accelerates convergence, cutting computation time by over 70%. The enhanced method enables efficient triangulation of disk topology surfaces, supports local size control, curvature adaptation, and re-meshing of discrete surfaces. Keywords: Adaptive triangulation, Surface remeshing, Bubble meshing, Conformal parameterization, Algorithm efficiency

cs.CG

Social Networks: Enumerating Maximal Community Patterns in $c$-Closed Graphs

Jacob Fox, C. Seshadhri, Tim Roughgarden, Fan Wei, and Nicole Wein introduced the model of $c$-closed graphs--a distribution-free model motivated by triadic closure, one of the most pervasive structural signatures of social networks. While enumerating maximal cliques in general graphs can take exponential time, it is known that in $c$-closed graphs, maximal cliques and maximal complete bipartite subgraphs can always be enumerated in polynomial time. These structures correspond to blow-ups of simple patterns: a single vertex or a single edge, with some vertices required to form cliques. In this work, we explore a natural extension: we study maximal blow-ups of arbitrary finite graphs $H$ in $c$-closed graphs. We prove that for any fixed graph $H$, the number of maximal blow-ups of $H$ in an $n$-vertex $c$-closed graph is always bounded by a polynomial in $n$. We further investigate the case of induced blow-ups and provide a precise characterization of the graphs $H$ for which the number of maximal induced blow-ups is also polynomially bounded in $n$. Finally, we study the analogue questions when $H$ ranges over an infinite family of graphs.

math.CO

On Domination Exponents for Pairs of Graphs

Understanding graph density profiles is notoriously challenging. Even for pairs of graphs, complete characterizations are known only in very limited cases, such as edges versus cliques. This paper explores a relaxation of the graph density profile problem by examining the homomorphism density domination exponent $C(H_1, H_2)$. This is the smallest real number $c \geq 0$ such that $t(H_1, T) \geq t(H_2, T)^c$ for all target graphs $T$ (if such a $c$ exists) where $t(H,T)$ is the homomorphism density from $H$ to $T$. We demonstrate that infinitely many families of graphs are required to realize $C(H_1, H_2)$ for all connected graphs $H_1$, $H_2$. We derive the homomorphism density domination exponent for a variety of graph pairs, including paths and cycles. As a couple of typical examples, we obtain exact values when $H_1$ is an even cycle and $H_2$ contains a Hamiltonian cycle, and provide asymptotically sharp bounds when both $H_1$ and $H_2$ are odd cycles.

math.CO

Density Measures for Language Generation

The recent successes of large language models (LLMs) have led to a surge of theoretical research into language generation. A recent line of work proposes an abstract view, called language generation in the limit, where generation is seen as a game between an adversary and an algorithm: the adversary generates strings from an unknown language $K$, chosen from a countable collection of candidate languages, and after seeing a finite set of these strings, the algorithm must generate new strings from $K$ that it has not seen before. This formalism highlights a key tension: the trade-off between validity (the algorithm should only produce strings from the language) and breadth (it should be able to produce many strings from the language). This trade-off is central in applied language generation as well, where it appears as a balance between hallucination (generating invalid utterances) and mode collapse (generating only a restricted set of outputs). Despite its importance, this trade-off has been challenging to study quantitatively. We develop ways to quantify this trade-off by formalizing breadth using measures of density. Existing algorithms for language generation in the limit produce output sets that can have zero density in the true language, and this important failure of breadth might seem unavoidable. We show, however, that such a failure is not necessary: we provide an algorithm for language generation in the limit whose outputs have strictly positive density in $K$. We also study the internal representations built by these algorithms, specifically the sequence of hypothesized candidate languages they consider, and show that achieving the strongest form of breadth may require oscillating indefinitely between high- and low-density representations. Our analysis introduces a novel topology on language families, with notions of convergence and limit points playing a key role.

math.CO

Constructing a Norm for Children's Scientific Drawing: Distribution Features Based on Semantic Similarity of Large Language Models

The use of children's drawings to examining their conceptual understanding has been proven to be an effective method, but there are two major problems with previous research: 1. The content of the drawings heavily relies on the task, and the ecological validity of the conclusions is low; 2. The interpretation of drawings relies too much on the subjective feelings of the researchers. To address this issue, this study uses the Large Language Model (LLM) to identify 1420 children's scientific drawings (covering 9 scientific themes/concepts), and uses the word2vec algorithm to calculate their semantic similarity. The study explores whether there are consistent drawing representations for children on the same theme, and attempts to establish a norm for children's scientific drawings, providing a baseline reference for follow-up children's drawing research. The results show that the representation of most drawings has consistency, manifested as most semantic similarity>0.8. At the same time, it was found that the consistency of the representation is independent of the accuracy (of LLM's recognition), indicating the existence of consistency bias. In the subsequent exploration of influencing factors, we used Kendall rank correlation coefficient to investigate the effects of "sample size", "abstract degree", and "focus points" on drawings, and used word frequency statistics to explore whether children represented abstract themes/concepts by reproducing what was taught in class. It was found that accuracy (of LLM's recognition) is the most sensitive indicator, and data such as sample size and semantic similarity are related to it; The consistency between classroom experiments and teaching purpose is also an important factor, many students focus more on the experiments themselves rather than what they explain.

cs.CL

Undecidability of polynomial inequalities in tournaments

Many fundamental problems in extremal combinatorics are equivalent to proving certain polynomial inequalities in graph homomorphism densities. In 2011, a breakthrough result by Hatami and Norine showed that it is undecidable to verify polynomial inequalities in graph homomorphism densities. Recently, Blekherman, Raymond and Wei extended this result by showing that it is also undecidable to determine the validity of polynomial inequalities in homomorphism densities for weighted graphs with edge weights taking real values. These two results resolved a question of Lov\'asz. In this paper, we consider the problem of determining the validity of polynomial inequalities in digraph homomorphism densities for tournaments. We prove that the answer to this problem is also undecidable.

math.CO