SearcharxivSearch

arXiv subjects

Yucong Chen

Publications and source records attributed to Yucong Chen.

7 recordsLinked to original sources

InfVSR: Toward Consistency-Driven Streaming Generative Video Super-Resolution

Real-world videos often extend over thousands of frames. Existing generative video super-resolution (VSR) approaches, however, face two persistent challenges when processing long sequences: (1) inefficiency due to the heavy cost of multi-step denoising for full-length sequences; and (2) poor consistency is hindered by temporal decomposition that causes artifacts and discontinuities. To break these limits, we propose InfVSR, which reformulates VSR as an autoregressive-one-step-diffusion paradigm, and enables streaming inference with video diffusion priors. First, we adapt the pretrained DiT into a causal structure, maintaining both local and global coherence via rolling KV-cache and joint visual guidance. Second, we distill the diffusion process into a single step efficiently, with patch-wise pixel supervision and cross-chunk distribution matching. To fill the gap in long-form video evaluation, we build a new benchmark tailored for extended sequences and further introduce semantic-level metrics to comprehensively assess temporal consistency. Our method pushes the frontier of long-form VSR, achieves state-of-the-art quality with enhanced semantic consistency, and delivers up to 58x speed-up over existing methods such as MGLD-VSR. Our code and models are available at https://github.com/Kai-Liu001/InfVSR.

cs.CV

An Absorption Correction for Reliable Pair-Distribution Functions from Low Energy X-ray Sources

This paper explores the development and testing of a simple absorption correction model for processing x-ray powder diffraction data from Debye-Scherrer geometry laboratory x-ray experiments. This may be used as a pre-processing step before using PDFgetX3 to obtain reliable pair distribution functions (PDFs). The correction was found to depend only on muD, the product of the x-ray attenuation coefficient and capillary diameter. Various experimental and theoretical methods for estimating muD were explored, and the most appropriate muD values for correction were identified for different capillary diameters and x-ray beam sizes. We identify operational ranges of muD where reasonable signal to noise is possible after correction. A user-friendly software package, diffpy.labpdfproc, is presented that can help estimate muD and perform absorption corrections, with a rapid calculation for efficient processing.

cond-mat.mtrl-sci

Testing Protocols for Obtaining Reliable PDFs from Laboratory x-ray Sources Using PDFgetX3

In this work, we explored data acquisition protocols and improved data reduction protocols using PDFgetX3 to obtain reliable data for atomic pair distribution function (PDF) analysis from a laboratory-based Mo x-ray source. A variable counting scheme is described that preferentially counts in the high-angle region of the diffraction pattern. The effects on the resulting PDF are studied by varying the overall count time, the use of Soller slits, and limiting the out-of-plane divergence of the incident beam. The protocols are tested using an amorphous silica and a quartz sample. We also present a modification to the current PDFgetX3 data corrections to take care of sample absorption, which was previously neglected in the use of that program for high-energy synchrotron x-ray data. We show that, despite limitations in the Q-range and flux of laboratory instruments, reasonable data for PDF model fits may be obtained using the best protocols in a few hours of counting.

cond-mat.mtrl-sci

BitQ: Tailoring Block Floating Point Precision for Improved DNN Efficiency on Resource-Constrained Devices

Deep neural networks (DNNs) are powerful for cognitive tasks such as image classification, object detection, and scene segmentation. One drawback however is the significant high computational complexity and memory consumption, which makes them unfeasible to run real-time on embedded platforms because of the limited hardware resources. Block floating point (BFP) quantization is one of the representative compression approaches for reducing the memory and computational burden owing to their capability to effectively capture the broad data distribution of DNN models. Unfortunately, prior works on BFP-based quantization empirically choose the block size and the precision that preserve accuracy. In this paper, we develop a BFP-based bitwidth-aware analytical modeling framework (called ``BitQ'') for the best BFP implementation of DNN inference on embedded platforms. We formulate and resolve an optimization problem to identify the optimal BFP block size and bitwidth distribution by the trade-off of both accuracy and performance loss. Experimental results show that compared with an equal bitwidth setting, the BFP DNNs with optimized bitwidth allocation provide efficient computation, preserving accuracy on famous benchmarks. The source code and data are available at https://github.com/Cheliosoops/BitQ.

cs.CV

MV-ROPE: Multi-view Constraints for Robust Category-level Object Pose and Size Estimation

Recently there has been a growing interest in category-level object pose and size estimation, and prevailing methods commonly rely on single view RGB-D images. However, one disadvantage of such methods is that they require accurate depth maps which cannot be produced by consumer-grade sensors. Furthermore, many practical real-world situations involve a moving camera that continuously observes its surroundings, and the temporal information of the input video streams is simply overlooked by single-view methods. We propose a novel solution that makes use of RGB video streams. Our framework consists of three modules: a scale-aware monocular dense SLAM solution, a lightweight object pose predictor, and an object-level pose graph optimizer. The SLAM module utilizes a video stream and additional scale-sensitive readings to estimate camera poses and metric depth. The object pose predictor then generates canonical object representations from RGB images. The object pose is estimated through geometric registration of these canonical object representations with estimated object depth points. All per-view estimates finally undergo optimization within a pose graph, culminating in the output of robust and accurate canonical object poses. Our experimental results demonstrate that when utilizing public dataset sequences with high-quality depth information, the proposed method exhibits comparable performance to state-of-the-art RGB-D methods. We also collect and evaluate on new datasets containing depth maps of varying quality to further quantitatively benchmark the proposed method alongside previous RGB-D based methods. We demonstrate a significant advantage in scenarios where depth input is absent or the quality of depth sensing is limited.

cs.CV

Modeling Multimodal Aleatoric Uncertainty in Segmentation with Mixture of Stochastic Experts

Equipping predicted segmentation with calibrated uncertainty is essential for safety-critical applications. In this work, we focus on capturing the data-inherent uncertainty (aka aleatoric uncertainty) in segmentation, typically when ambiguities exist in input images. Due to the high-dimensional output space and potential multiple modes in segmenting ambiguous images, it remains challenging to predict well-calibrated uncertainty for segmentation. To tackle this problem, we propose a novel mixture of stochastic experts (MoSE) model, where each expert network estimates a distinct mode of the aleatoric uncertainty and a gating network predicts the probabilities of an input image being segmented in those modes. This yields an efficient two-level uncertainty representation. To learn the model, we develop a Wasserstein-like loss that directly minimizes the distribution distance between the MoSE and ground truth annotations. The loss can easily integrate traditional segmentation quality measures and be efficiently optimized via constraint relaxation. We validate our method on the LIDC-IDRI dataset and a modified multimodal Cityscapes dataset. Results demonstrate that our method achieves the state-of-the-art or competitive performance on all metrics.

cs.CV

Beam distribution reconstruction simulation for electron beam probe

Electron beam probe (EBP) is a new principle detector, which makes use of a low-intensity and low-energy electron beam to measure the transverse profile, bunch shape, beam neutralization and beam wake field of an intense beam with small dimensions. While can be applied to many aspects, we limit our analysis to beam distribution reconstruction. This kind of detector is almost non-interceptive for all of the beam and does not disturb the machine environment. In this paper, we present the theoretical aspects behind this technique for beam distribution measurement and some simulation results of the detector involved. First, a method to obtain parallel electron beam is introduced and a simulation code is developed. And then, EBP as a profile monitor for dense beam is simulated using fast scan method under various target beam profile, such as KV distribution, waterbag distribution, parabolic distribution, Gaussian distribution and halo distribution. Profile reconstruction from the deflected electron beam trajectory is implemented and compared with the actual one, and an expected agreement is achieved. Furthermore, Instead of fast scan, a slow scan, i.e. step-by-step scan, is considered, which lows the requirement for hardware, i.e. Radio Frequency deflector. we calculate the three dimensional electric field of Gaussian distribution and simulate the electron motion under this field. In addition, fast scan along the target beam direction and slow scan across the beam is also presented, and can provide a measurement of longitudinal distribution as well as transverse profile simultaneously. Final, simulation results for China Accelerator Driven Sub-critical System (CADS) and High Intensity Heavy Ion Accelerator Facility (HIAF) are given to investigate the quantitative behavior of EBP.

physics.acc-ph