SearcharxivSearch

arXiv subjects

Zizhong Ding

Publications and source records attributed to Zizhong Ding.

4 recordsLinked to original sources

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs

Visual token pruning reduces the inference cost of multimodal large language models, but a fixed token ratio is poorly matched to text-rich inputs. In OCR-centric tasks, decisive evidence can be a small number, label, or field whose relevance is specified by the question; indiscriminate pruning can erase that evidence while retaining visually salient but irrelevant regions. We present ET-Prune, a training-free framework that casts pruning as evidence allocation. It derives question-conditioned evidence from a decoder-side partial query-key block, safeguards text-like spatial regions, and converts evidence uncertainty and density into a sample-specific token floor. Three progressive middle-layer events then move the sequence toward this budget, retaining more tokens for diffuse or text-dense evidence and pruning concentrated evidence more aggressively. At the observed point estimates from one deterministic pass per configuration, ET-Prune leads or ties among pruned methods in all six backbone-benchmark comparisons at roughly half tokens. On OCRBench-v2, it leads the strongest pruned baselines by 1.80 and 0.68 percentage points on Qwen3-VL-8B and InternVL3.5-8B, respectively, while retaining about half of the visual tokens; on MMBench v1.1, it reaches 0.8467 circular exact-matching accuracy versus 0.8437 for Vanilla at 54.45% average visual-token retention. These results show a favorable observed quality-cost trade-off for evidence-aware dynamic budgeting in text-rich multimodal inference.

cs.CV

AeroMesa: Efficient Data Management System for Multi-Dimensional Spatio-Temporal Trajectories

The proliferation of multi-dimensional trajectory data, fueled by large-scale IoT and the emerging low-altitude economy, particularly UAV operations, drives repositories to jointly support (x,y), (x,y,t), (x,y,z), and (x,y,z,t) queries within a single storage framework. Yet existing HBase-based systems fall short in three respects: severe row-key interval fragmentation when altitude is jointly encoded with horizontal coordinates, locality-unfriendly spatial encodings with workload-blind shape-code ordering, and coarse-grained temporal indexes that leave intra-slot boundary ambiguity unresolved. We present AeroMesa, an efficient data management system for multi-dimensional spatio-temporal trajectories built on Apache HBase and Redis, that natively supports (x,y), (x,y,t), (x,y,z), and (x,y,z,t) queries within a unified storage framework. AeroMesa addresses the above limitations through three designs: a decoupled horizontal-altitude architecture with a multi-granularity Height Spatio-Temporal Index (HTSI) that eliminates joint encoding fragmentation; Hilbert-BFS with Workload-Aware Jaccard (WAJ) reordering that improves spatial locality; and TI+, a dual-offset temporal index that resolves intra-slot false positives. Evaluations on T-Drive and an 87,537-trajectory high-fidelity UAV simulation demonstrate that AeroMesa reduces 3D/4D query latency by up to 30x over XZ3/TXZ3, lowers 2D latency by up to 17.9% over TMan, and cuts temporal candidates by up to 51.3% over MCTM, with sub-linear scalability confirmed under 200x data expansion, confirming AeroMesa's efficiency for multi-dimensional spatio-temporal trajectory management.

cs.DB

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models

The development of separate-encoder Unified multimodal models (UMMs) comes with a rapidly growing inference cost due to dense visual token processing. In this paper, we focus on understanding-side visual token reduction for improving the efficiency of separate-encoder UMMs. While this topic has been widely studied for MLLMs, existing methods typically rely on attention scores, text-image similarity and so on, implicitly assuming that the final objective is discriminative reasoning. This assumption does not hold for UMMs, where understanding-side visual tokens must also preserve the model's capabilities for editing images. We propose G$^2$TR, a generation-guided visual token reduction framework for separate-encoder UMMs. Our key insight is that the generation branch provides a task-agnostic signal for identifying understanding-side visual tokens that are not only semantically relevant but also important for latent-space image reconstruction and generation. G$^2$TR estimates token importance from consistency with VAE latent, performs balanced token selection, and merges redundant tokens into retained representatives to reduce information loss. The method is training-free, plug-and-play, and applied only after the understanding encoding stage, making it compatible with existing UMM inference pipelines. Experiments on image understanding and editing benchmarks show that G$^2$TR substantially reduces visual tokens and prefill computation by 1.94x while maintaining both reasoning accuracy and editing quality, outperforming baselines on almost all benchmarks. Code is at: https://github.com/lijunxian111/G2TR.

cs.CV

AdaTSQ: Pushing the Pareto Frontier of Diffusion Transformers via Temporal-Sensitivity Quantization

Diffusion Transformers (DiTs) have emerged as the state-of-the-art backbone for high-fidelity image and video generation. However, their massive computational cost and memory footprint hinder deployment on edge devices. While post-training quantization (PTQ) has proven effective for large language models (LLMs), directly applying existing methods to DiTs yields suboptimal results due to the neglect of the unique temporal dynamics inherent in diffusion processes. In this paper, we propose AdaTSQ, a novel PTQ framework that pushes the Pareto frontier of efficiency and quality by exploiting the temporal sensitivity of DiTs. First, we propose a Pareto-aware timestep-dynamic bit-width allocation strategy. We model the quantization policy search as a constrained pathfinding problem. We utilize a beam search algorithm guided by end-to-end reconstruction error to dynamically assign layer-wise bit-widths across different timesteps. Second, we propose a Fisher-guided temporal calibration mechanism. It leverages temporal Fisher information to prioritize calibration data from highly sensitive timesteps, seamlessly integrating with Hessian-based weight optimization. Extensive experiments on four advanced DiTs (e.g., Flux-Dev, Flux-Schnell, Z-Image, and Wan2.1) demonstrate that AdaTSQ significantly outperforms state-of-the-art methods like SVDQuant and ViDiT-Q. Our code will be released at https://github.com/Qiushao-E/AdaTSQ.

cs.CV