SearcharxivSearch

arXiv subjects

Peishan Dai

Publications and source records attributed to Peishan Dai.

2 recordsLinked to original sources

Thinking in Video: Can Video Generators Really Reason About the Real World?

Recent advances in world models and video generation have given rise to an emerging reasoning paradigm that leverages video generative models to simulate, predict, and reason about real-world dynamics. We redefine this paradigm as Thinking in Video, where video is not merely an output artifact but a medium for constructing, extending, and verifying causal thought. However, this promise remains unverified: convincing rollouts may reflect memorized appearances rather than causal understanding, while existing metrics separate perceptual fidelity from semantic logic. To evaluate whether video generators support such reasoning, we introduce the Causal-Generative Dual-Judge (CGDJ), auditing World Model Consistency from two perspectives. Explicit Causal Perception tests whether a generator reads a video scenario as a reasoning problem through spatio-temporal flattened visual question answering, while Implicit Generative Perception-Prediction Gap evaluates whether it renders the causal consequence as a consistent future video. Applying CGDJ to representative open- and closed-source generators reveals a clear Perception-Prediction Gap: open-source models produce plausible dynamics despite near-zero explicit causal perception, whereas advanced closed-source systems show stronger but still limited alignment between reasoning and generation. Further analysis exposes audio-visual misalignment, where models verbalize correct causal logic more reliably than they render it, challenging the "world simulator" narrative.

cs.CV

Effective connectivity signatures in major depressive disorder: fMRI study using a multi-site dataset

Diagnosis of major depressive disorder (MDD) primarily relies on the patient's self-reported symptoms and a clinical evaluation. Effective connectivity (EC) from resting-state functional magnetic resonance imaging (rs-fMRI) analysis can reflect the directionality of connections between brain regions, making it a candidate method to classify MDD. This study used Granger causality analysis to extract EC features from a large multi-site MDD dataset. The ComBat algorithm and multivariate linear regression were used to harmonize site difference and to remove age and sex covariates, respectively. Two-sample t-tests and model-based feature selection methods were used to screen for highly discriminative EC features for MDD, and LightGBM was used to classify MDD. In this large-scale multi-site rs-fMRI dataset, 97 EC features deemed highly discriminative for MDD were screened. In the nested five-fold cross-validation, the best classification model with the 97 EC features achieved accuracy, sensitivity, and specificity of 94.35%, 93.52%, and 95.25%, respectively. In another independent large dataset, which tested the generalization performance of the 97 EC features, the best classification models achieved 94.74%, 90.59%, and 96.75% for accuracy, sensitivity, and specificity, respectively. This work demonstrated that EC had a reasonable discriminative ability and supported the notion for using EC to potentially assist clinical diagnosis of MDD.

q-bio.NC