Searcharxiv⌕ Search

arXiv subjects

Zhizheng Zhao

Publications and source records attributed to Zhizheng Zhao.

3 recordsLinked to original sources

Nuisance-Aware Muon Tomography

Cosmic-ray muon scattering tomography can image dense, shielded, or inaccessible objects without an artificial radiation source. In a compact magnet-free tracker, however, each accepted muon provides only a few hit positions and no event-by-event momentum measurement. The downstream hit residual is therefore a compound observable: target scattering, muon momentum, detector resolution, support material, air scattering, and track extrapolation all enter the same measured displacement. We introduce Nuisance-Aware Muon Tomography (NAMT), a residual-likelihood reconstruction method for magnet-free trackers. The upstream hits define the incident track, downstream hit residuals carry the scattering signal, and a radiation-length density field $λ=1/X_0$ predicts their material-induced variance through a path integral. NAMT marginalizes the unmeasured momentum with a shared event-level scattering scale and uses open-field blank scans to fix detector and environmental residuals before object reconstruction. On eight Geant4 benchmark scenes spanning strong, weak, and negative scattering contrast, NAMT-4P reaches a mean area under the ROC curve (AUC) of $0.916$ at $120$k effective muons and $1$ mm hit error, compared with $0.784$ for ASR, $0.749$ for MLS-EM, and $0.643$ for PoCA. NAMT-3P uses one downstream hit plane in reconstruction and still reaches $0.909$ mean AUC at the reference setting, while giving the highest reference mean contrast-to-noise ratio (CNR) and the best mean AUC at $30$k muons.

physics.ins-det↗

CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation

Recent video generation models have revealed the emergence of Chain-of-Frame (CoF) reasoning, enabling frame-by-frame visual inference. With this capability, video models have been successfully applied to various visual tasks (e.g., maze solving, visual puzzles). However, their potential to enhance text-to-image (T2I) generation remains largely unexplored due to the absence of a clearly defined visual reasoning starting point and interpretable intermediate states in the T2I generation process. To bridge this gap, we propose CoF-T2I, a model that integrates CoF reasoning into T2I generation via progressive visual refinement, where intermediate frames act as explicit reasoning steps and the final frame is taken as output. To establish such an explicit generation process, we curate CoF-Evol-Instruct, a dataset of CoF trajectories that model the generation process from semantics to aesthetics. To further improve quality and avoid motion artifacts, we enable independent encoding operation for each frame. Experiments show that CoF-T2I significantly outperforms the base video model and achieves competitive performance on challenging benchmarks, reaching 0.86 on GenEval and 7.468 on Imagine-Bench. These results indicate the substantial promise of video models for advancing high-quality text-to-image generation.

cs.CV↗

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Chain-of-Thought (CoT) reasoning has been extensively explored in large models to tackle complex understanding tasks. However, it still remains an open question whether such strategies can be applied to verifying and reinforcing image generation scenarios. In this paper, we provide the first comprehensive investigation of the potential of CoT reasoning to enhance autoregressive image generation. We focus on three techniques: scaling test-time computation for verification, aligning model preferences with Direct Preference Optimization (DPO), and integrating these techniques for complementary effects. Our results demonstrate that these approaches can be effectively adapted and combined to significantly improve image generation performance. Furthermore, given the pivotal role of reward models in our findings, we propose the Potential Assessment Reward Model (PARM) and PARM++, specialized for autoregressive image generation. PARM adaptively assesses each generation step through a potential assessment approach, merging the strengths of existing reward models, and PARM++ further introduces a reflection mechanism to self-correct the generated unsatisfactory image, which is the first to incorporate reflection in autoregressive image generation. Using our investigated reasoning strategies, we enhance a baseline model, Show-o, to achieve superior results, with a significant +24% improvement on the GenEval benchmark, surpassing Stable Diffusion 3 by +15%. We hope our study provides unique insights and paves a new path for integrating CoT reasoning with autoregressive image generation. Code and models are released at https://github.com/ZiyuGuo99/Image-Generation-CoT

cs.CV↗