SearcharxivSearch

arXiv subjects

Wen Zhong

Publications and source records attributed to Wen Zhong.

13 recordsLinked to original sources

Mirroring the Past: Exploring How Ancestral Digital Self Influences History Learning

Learners often perceive history as distant from themselves, which limits immersion and empathy in history learning. To bridge this gap, we introduce the "Ancestral Digital Self," an AI-generated pedagogical agent presented in prerecorded videos that mirrors the learner's facial features and vocal timbre, representing a historically situated version of the self. We developed a reproducible workflow for creating AI-generated historical learning videos and conducted a within-subjects study (N=36) comparing a Digital Self agent with a non-self pedagogical agent. The Digital Self agent enhanced experiential measures, including narrative transportation, perceived relatedness, self-other inclusion, and agent perception. However, it did not improve immediate learning outcomes: quiz scores were lower in the Digital Self condition, and Remember/Know judgments showed no reliable differences. Interviews further suggested that self-similarity increased familiarity and motivation, while novelty and uncanniness could draw attention away from historical content. These findings offer design implications for future educational environments supported by pedagogical agents.

cs.HC

VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing

The rapid rise of vlogs as a personalized storytelling medium has created a demand for automated systems to evaluate and refine vlog editing plans. However, vlog assessment is highly subjective and remains challenging due to a lack of standardized criteria, dataset and benchmark, and effective reward models. To address these challenges, we define a comprehensive vlog evaluation framework guided by professional vlog creators and product managers, establishing a taxonomy of six key dimensions, i.e., Creativity, Consistency, Concept Design, Cinematography, Narration, and Pacing. Subsequently, we curate a large-scale dataset of 100k vlog edits and a dedicated benchmark, VRMBench, to evaluate the vlog rewarding capabilities of Multimodal Large Language Models (MLLMs). Finally, we present VlogReward, a robust vlog reward model that can provide both fine-grained multi-dimensional scores and actionable feedback for iterative refinement. Technically, we enhance the Group Relative Policy Optimization (GRPO) framework by introducing an adjustable inter-group comparison reward, which mitigates the "direction blindness" issue of standard GRPO and enables the model to better distinguish varied-quality edits. VlogReward achieves state-of-the-art results that significantly outperform existing MLLMs, including GPT-5 and Gemini-3-Pro. We hope that our study can help vlog creators and foster automated vlog evaluation and refinement systems.

cs.AI

CoilDrop-MRI: Self-supervised physics-guided MRI reconstruction with coil dropout

Self-supervised deep learning-based methods have shown great promise for accelerated magnetic resonance imaging (MRI) reconstruction, achieving high image quality without requiring fully sampled data for training. These methods typically partition the acquired data into two disjoint subsets to construct input-target pairs for optimizing the reconstruction network. However, existing approaches perform this partition exclusively within the spatial frequency (k-space) domain, leaving the coil dimension unexplored. To enforce full exploitation of signal correlation across receiver coils, we propose CoilDrop-MRI, which applies coil-wise dropout to the input and uses the dropped data as training targets in a self-supervised framework. This method is integrated into unrolled architectures in both image-domain (SENSE) and k-space (SPIRiT) formulations. We further demonstrate its versatility by extending CoilDrop-MRI to multi-shot, phase-corrected diffusion MRI (dMRI) reconstruction. CoilDrop-MRI is extensively validated on multi-site, multi-field-strength (0.3T, 0.55T, and 3T), and multi-modality (T1-weighted, T2-weighted, T2-FLAIR, and dMRI) datasets and consistently outperforms state-of-the-art self-supervised methods, achieving quality comparable to supervised reconstruction methods without requiring fully sampled reference training data. Moreover, CoilDrop-MRI exhibits strong data efficiency and robust generalization across imaging conditions, establishing it as a practical and versatile framework for self-supervised parallel MRI reconstruction.

cs.CV

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing

Real-world video editing demands not only expert knowledge of cinematic techniques but also multimodal reasoning to select, align, and combine footage into coherent narratives. While recent Large Multimodal Models (LMMs) have shown remarkable progress in general video understanding, their abilities in multi-video reasoning and operational editing workflows remain largely unexplored. We introduce VEBENCH, the first comprehensive benchmark designed to evaluate both editing knowledge understanding and operational reasoning in realistic video editing scenarios. VEBENCH contains 3.9K high-quality edited videos (over 257 hours) and 3,080 human-verified QA pairs, built through a three-round human-AI collaborative annotation pipeline that ensures precise temporal labeling and semantic consistency. It features two complementary QA tasks: 1) Video Editing Technique Recognition, assessing models' ability to identify 7 editing techniques using multimodal cues; and 2) Video Editing Operation Simulation, modeling real-world editing workflows by requiring the selection and temporal localization of relevant clips from multiple candidates. Extensive experiments across proprietary (e.g., Gemini-2.5-Pro) and open-source LMMs reveal a large gap between current model performance and human-level editing cognition. These results highlight the urgent need for bridging video understanding with creative operational reasoning. We envision VEBENCH as a foundation for advancing intelligent video editing systems and driving future research on complex reasoning.

cs.CV

Vidi2.5: Large Multimodal Models for Video Understanding and Creation

Video has emerged as the primary medium for communication and creativity on the Internet, driving strong demand for scalable, high-quality video production. Vidi models continue to evolve toward next-generation video creation and have achieved state-of-the-art performance in multimodal temporal retrieval (TR). In its second release, Vidi2 advances video understanding with fine-grained spatio-temporal grounding (STG) and extends its capability to video question answering (Video QA), enabling comprehensive multimodal reasoning. Given a text query, Vidi2 can identify not only the corresponding timestamps but also the bounding boxes of target objects within the output time ranges. To enable comprehensive evaluation of STG, we introduce a new benchmark, VUE-STG, which offers critical improvements over existing STG datasets. In addition, we upgrade the previous VUE-TR benchmark to VUE-TR-V2, achieving a more balanced duration and query distribution. Remarkably, the Vidi2 model substantially outperforms leading proprietary systems, such as Gemini 3 Pro Preview and GPT-5, on both VUE-TR-V2 and VUE-STG, while achieving competitive results with popular open-source models with similar scale on video QA benchmarks. The latest Vidi2.5 offers significantly stronger STG capability and slightly better TR and Video QA performance over Vidi2. This update also introduces a Vidi2.5-Think model to handle plot understanding with complex plot reasoning. To comprehensively evaluate the performance of plot understanding, we propose VUE-PLOT benchmark with two tracks, Character and Reasoning. Notably, Vidi2.5-Think outperforms Gemini 3 Pro Preview on fine-grained character understanding with comparable performance on complex plot reasoning. Furthermore, we demonstrate the effectiveness of Vidi2.5 on a challenging real-world application, video editing planning.

cs.CV

Multishot Dual Polarity GRAPPA: Robust Nyquist Ghost Correction for multishot EPI

Purpose: This work aims to develop a robust Nyquist ghost correction method for multishot echo-planar imaging (EPI). The method helps correct challenging Nyquist ghosts, particularly on scanners with high-performance gradients or ultra-high fields. Methods: A method for multishot EPI ghost correction, called multishot dual-polarity GRAPPA (msDPG), is developed by extending the DPG concept to multishot readouts. msDPG employs tailored DPG kernels to address high-order phase differences between two EPI readout polarities, which cannot be fully addressed using linear phase correction (LPC). Advanced regularizers can be readily employed with the proposed msDPG for physiologic inter-shot phase variation correction during reconstruction. Additionally, a calibration refinement method is proposed to improve the quality of the DPG calibration data and enhance reconstruction performance. Results: Phantom and in vivo experiments on scanners with high-performance gradients and ultra-high fields demonstrated that msDPG achieved superior ghost correction performance than LPC, reducing the ghost-to-signal ratio (GSR) by over 50%. Compared to conventional DPG, msDPG provided images with lower noise amplification, particularly for acquisitions with large in-plane acceleration. Consequently, high-fidelity, submillimeter diffusion images were obtained using msDPG with regularized reconstruction. Conclusion: The proposed msDPG provides a robust Nyquist ghost correction method for multishot EPI, enabling submillimeter imaging with improved fidelity.

physics.med-ph

Vidi: Large Multimodal Models for Video Understanding and Editing

Humans naturally share information with those they are connected to, and video has become one of the dominant mediums for communication and expression on the Internet. To support the creation of high-quality large-scale video content, a modern pipeline requires a comprehensive understanding of both the raw input materials (e.g., the unedited footage captured by cameras) and the editing components (e.g., visual effects). In video editing scenarios, models must process multiple modalities (e.g., vision, audio, text) with strong background knowledge and handle flexible input lengths (e.g., hour-long raw videos), which poses significant challenges for traditional models. In this report, we introduce Vidi, a family of Large Multimodal Models (LMMs) for a wide range of video understand editing scenarios. The first release focuses on temporal retrieval, i.e., identifying the time ranges within the input videos corresponding to a given text query, which plays a critical role in intelligent editing. The model is capable of processing hour-long videos with strong temporal understanding capability, e.g., retrieve time ranges for certain queries. To support a comprehensive evaluation in real-world scenarios, we also present the VUE-TR benchmark, which introduces five key advancements. 1) Video duration: significantly longer than videos of existing temporal retrival datasets, 2) Audio support: includes audio-based queries, 3) Query format: diverse query lengths/formats, 4) Annotation quality: ground-truth time ranges are manually annotated. 5) Evaluation metric: a refined IoU metric to support evaluation over multiple time ranges. Remarkably, Vidi significantly outperforms leading proprietary models, e.g., GPT-4o and Gemini, on the temporal retrieval task, indicating its superiority in video editing scenarios.

cs.CV

Efficient Information Geometry Approach for Massive MIMO-OFDM Channel Estimation

We investigate the channel estimation for massive multiple-input multiple-output orthogonal frequency division multiplexing (MIMO-OFDM) systems. We revisit the information geometry approach (IGA) for massive MIMO-OFDM channel estimation. By using the constant magnitude property of the entries of the measurement matrix, we find that the second-order natural parameters of the distributions on all the auxiliary manifolds are equivalent to each other, and the first-order natural parameters are asymptotically equivalent to each other at the fixed point. Motivated by these results, we simplify the process of IGA and propose an efficient IGA (EIGA) for massive MIMO-OFDM channel estimation, which allows efficient implementation with fast Fourier transformation (FFT). We then establish a sufficient condition of its convergence and accordingly find a range of the damping factor for the convergence. We show that this range of damping factor is sufficiently wide by using the specific properties of the measurement matrices. Further, we prove that at the fixed point, the a posteriori mean obtained by EIGA is asymptotically optimal. Simulations confirm that EIGA can achieve the optimal performance with low complexity in a limited number of iterations.

cs.IT

Robust Transmission for Massive MIMO Downlink with Imperfect CSI

In this paper, the design of robust linear precoders for the massive multi-input multi-output (MIMO) downlink with imperfect channel state information (CSI) is investigated. The imperfect CSI for each UE obtained at the BS is modeled as statistical CSI under a jointly correlated channel model with both channel mean and channel variance information, which includes the effects of channel estimation error, channel aging and spatial correlation. The design objective is to maximize the expected weighted sum-rate. By combining the minorize-maximize (MM) algorithm with the deterministic equivalent method, an algorithm for robust linear precoder design is derived. The proposed algorithm achieves a stationary point of the expected weighted sum-rate maximization problem. To reduce the computational complexity, two low-complexity algorithms are then derived. One for the general case, and the other for the case when all the channel means are zeros. For the later case, it is proved that the beam domain transmission is optimal, and thus the precoder design reduces to the power allocation optimization in the beam domain. Simulation results show that the proposed robust linear precoder designs apply to various mobile scenarios and achieve high spectral efficiency.

cs.IT

MV-RNN: A Multi-View Recurrent Neural Network for Sequential Recommendation

Sequential recommendation is a fundamental task for network applications, and it usually suffers from the item cold start problem due to the insufficiency of user feedbacks. There are currently three kinds of popular approaches which are respectively based on matrix factorization (MF) of collaborative filtering, Markov chain (MC), and recurrent neural network (RNN). Although widely used, they have some limitations. MF based methods could not capture dynamic user's interest. The strong Markov assumption greatly limits the performance of MC based methods. RNN based methods are still in the early stage of incorporating additional information. Based on these basic models, many methods with additional information only validate incorporating one modality in a separate way. In this work, to make the sequential recommendation and deal with the item cold start problem, we propose a Multi-View Recurrent Neural Network (MV-RNN}) model. Given the latent feature, MV-RNN can alleviate the item cold start problem by incorporating visual and textual information. First, At the input of MV-RNN, three different combinations of multi-view features are studied, like concatenation, fusion by addition and fusion by reconstructing the original multi-modal data. MV-RNN applies the recurrent structure to dynamically capture the user's interest. Second, we design a separate structure and a united structure on the hidden state of MV-RNN to explore a more effective way to handle multi-view features. Experiments on two real-world datasets show that MV-RNN can effectively generate the personalized ranking list, tackle the missing modalities problem and significantly alleviate the item cold start problem.

cs.IR

Sequential Ensemble Learning for Outlier Detection: A Bias-Variance Perspective

Ensemble methods for classification and clustering have been effectively used for decades, while ensemble learning for outlier detection has only been studied recently. In this work, we design a new ensemble approach for outlier detection in multi-dimensional point data, which provides improved accuracy by reducing error through both bias and variance. Although classification and outlier detection appear as different problems, their theoretical underpinnings are quite similar in terms of the bias-variance trade-off [1], where outlier detection is considered as a binary classification task with unobserved labels but a similar bias-variance decomposition of error. In this paper, we propose a sequential ensemble approach called CARE that employs a two-phase aggregation of the intermediate results in each iteration to reach the final outcome. Unlike existing outlier ensembles which solely incorporate a parallel framework by aggregating the outcomes of independent base detectors to reduce variance, our ensemble incorporates both the parallel and sequential building blocks to reduce bias as well as variance by ($i$) successively eliminating outliers from the original dataset to build a better data model on which outlierness is estimated (sequentially), and ($ii$) combining the results from individual base detectors and across iterations (parallelly). Through extensive experiments on sixteen real-world datasets mainly from the UCI machine learning repository [2], we show that CARE performs significantly better than or at least similar to the individual baselines. We also compare CARE with the state-of-the-art outlier ensembles where it also provides significant improvement when it is the winner and remains close otherwise.

cs.LG

A Data-Driven Approach for Mapping Multivariate Data to Color

A wide variety of color schemes have been devised for mapping scalar data to color. Some use the data value to index a color scale. Others assign colors to different, usually blended disjoint materials, to handle areas where materials overlap. A number of methods can map low-dimensional data to color, however, these methods do not scale to higher dimensional data. Likewise, schemes that take a more artistic approach through color mixing and the like also face limits when it comes to the number of variables they can encode. We address the challenge of mapping multivariate data to color and avoid these limitations at the same time. It is a data driven method, which first gauges the similarity of the attributes and then arranges them according to the periphery of a convex 2D color space, such as HSL. The color of a multivariate data sample is then obtained via generalized barycentric coordinate (GBC) interpolation.

cs.GR

Channel Acquisition for Massive MIMO-OFDM with Adjustable Phase Shift Pilots

We propose adjustable phase shift pilots (APSPs) for channel acquisition in wideband massive multiple-input multiple-output (MIMO) systems employing orthogonal frequency division multiplexing (OFDM) to reduce the pilot overhead. Based on a physically motivated channel model, we first establish a relationship between channel space-frequency correlations and the channel power angle-delay spectrum in the massive antenna array regime, which reveals the channel sparsity in massive MIMO-OFDM. With this channel model, we then investigate channel acquisition, including channel estimation and channel prediction, for massive MIMO-OFDM with APSPs. We show that channel acquisition performance in terms of sum mean square error can be minimized if the user terminals' channel power distributions in the angle-delay domain can be made non-overlapping with proper phase shift scheduling. A simplified pilot phase shift scheduling algorithm is developed based on this optimal channel acquisition condition. The performance of APSPs is investigated for both one symbol and multiple symbol data models. Simulations demonstrate that the proposed APSP approach can provide substantial performance gains in terms of achievable spectral efficiency over the conventional phase shift orthogonal pilot approach in typical mobility scenarios.

cs.IT