SearcharxivSearch

arXiv subjects

Zhihang Yi

Publications and source records attributed to Zhihang Yi.

5 recordsLinked to original sources

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning

Fine-grained visual reasoning requires multimodal large language models (MLLMs) to identify task-relevant visual evidence and ground their reasoning in local image regions. Existing agentic methods typically rely on reinforcement learning with verifiable rewards or supervised fine-tuning on large-scale annotated reasoning traces, leading to costly exploration, hand-designed verification rules, or heavy dependence on textual supervision. A natural way to avoid such external answer labels is to learn from trajectories sampled by the student itself, which points to On-Policy Distillation (OPD). To understand what OPD can and cannot provide for visual reasoning, we revisit it as negative-free stop-gradient alignment. This perspective shows that, although OPD provides effective token-level correction, its ceiling is constrained by the absence of trajectory-level discrimination. Motivated by these observations, we propose V-Zero, an answer-label-free framework for visual reasoning with contrastive evidence gating. V-Zero uses no annotated textual answer labels; instead, during training it pairs a question-relevant regional crop with a negative visual view to evaluate student-sampled trajectories and gate dense token-level distillation. Experiments on multiple visual reasoning benchmarks show that V-Zero consistently improves fine-grained visual reasoning while preserving strong generalization. Notably, V-Zero is more than 5$\times$ faster than previous supervised fine-tuning methods and more than 10$\times$ faster than reinforcement learning baselines. Code and dataset will be released at https://github.com/eVI-group-SCU/V-Zero

cs.CV

Multimodal Information Fusion for Chart Understanding: A Survey of MLLMs -- Evolution, Limitations, and Cognitive Enhancement

Chart understanding is a quintessential information fusion task, requiring the seamless integration of graphical and textual data to extract meaning. The advent of Multimodal Large Language Models (MLLMs) has revolutionized this domain, yet the landscape of MLLM-based chart analysis remains fragmented and lacks systematic organization. This survey provides a comprehensive roadmap of this nascent frontier by structuring the domain's core components. We begin by analyzing the fundamental challenges of fusing visual and linguistic information in charts. We then categorize downstream tasks and datasets, introducing a novel taxonomy of canonical and non-canonical benchmarks to highlight the field's expanding scope. Subsequently, we present a comprehensive evolution of methodologies, tracing the progression from classic deep learning techniques to state-of-the-art MLLM paradigms that leverage sophisticated fusion strategies. By critically examining the limitations of current models, particularly their perceptual and reasoning deficits, we identify promising future directions, including advanced alignment techniques and reinforcement learning for cognitive enhancement. This survey aims to equip researchers and practitioners with a structured understanding of how MLLMs are transforming chart information fusion and to catalyze progress toward more robust and reliable systems.

cs.CV

High Data-Rate Single-Symbol ML Decodable Distributed STBCs for Cooperative Networks

High data-rate Distributed Orthogonal Space-Time Block Codes (DOSTBCs) which achieve the single-symbol decodability and full diversity order are proposed in this paper. An upper bound of the data-rate of the DOSTBC is derived and it is approximately twice larger than that of the conventional repetition-based cooperative strategy. In order to facilitate the systematic constructions of the DOSTBCs achieving the upper bound of the data-rate, some special DOSTBCs, which have diagonal noise covariance matrices at the destination terminal, are investigated. These codes are referred to as the row-monomial DOSTBCs. An upper bound of the data-rate of the row-monomial DOSTBC is derived and it is equal to or slightly smaller than that of the DOSTBC. Lastly, the systematic construction methods of the row-monomial DOSTBCs achieving the upper bound of the data-rate are presented.

cs.IT

Finite-SNR Diversity-Multiplexing Tradeoff and Optimum Power Allocation in Bidirectional Cooperative Networks

This paper focuses on analog network coding (ANC) and time division broadcasting (TDBC) which are two major protocols used in bidirectional cooperative networks. Lower bounds of the outage probabilities of those two protocols are derived first. Those lower bounds are extremely tight in the whole signal-to-noise ratio (SNR) range irrespective of the values of channel variances. Based on those lower bounds, finite-SNR diversity-multiplexing tradeoffs of the ANC and TDBC protocols are obtained. Secondly, we investigate how to efficiently use channel state information (CSI) in those two protocols. Specifically, an optimum power allocation scheme is proposed for the ANC protocol. It simultaneously minimizes the outage probability and maximizes the total mutual information of this protocol. For the TDBC protocol, an optimum method to combine the received signals at the relay terminal is developed under an equal power allocation assumption. This method minimizes the outage probability and maximizes the total mutual information of the TDBC protocol at the same time.

cs.IT

The Impact of Noise Correlation and Channel Phase Information on the Data-Rate of the Single-Symbol ML Decodable Distributed STBCs

Very recently, we proposed the row-monomial distributed orthogonal space-time block codes (DOSTBCs) and showed that the row-monomial DOSTBCs achieved approximately twice higher bandwidth efficiency than the repetitionbased cooperative strategy [1]. However, we imposed two limitations on the row-monomial DOSTBCs. The first one was that the associated matrices of the codes must be row-monomial. The other was the assumption that the relays did not have any channel state information (CSI) of the channels from the source to the relays, although this CSI could be readily obtained at the relays without any additional pilot signals or any feedback overhead. In this paper, we first remove the row-monomial limitation; but keep the CSI limitation. In this case, we derive an upper bound of the data-rate of the DOSTBC and it is larger than that of the row-monomial DOSTBCs in [1]. Secondly, we abandon the CSI limitation; but keep the row-monomial limitation. Specifically, we propose the row-monomial DOSTBCs with channel phase information (DOSTBCs-CPI) and derive an upper bound of the data-rate of those codes. The rowmonomial DOSTBCs-CPI have higher data-rate than the DOSTBCs and the row-monomial DOSTBCs. Furthermore, we find the actual row-monomial DOSTBCs-CPI which achieve the upper bound of the data-rate.

cs.IT