SearcharxivSearch

arXiv subjects

Zhenyu Jiang

Publications and source records attributed to Zhenyu Jiang.

At least 19 recordsLinked to original sources

Rise Time and Charge Collection Efficiency of Graphene-Optimized 4H-SiC p-i-n Detector

Silicon carbide detectors exhibit good detection performance and are being considered for detection applications. However, the presence of surface electrode of detector limits the application of low-penetration particle detectors, photodetectors and heavy-ion detection. A graphene-optimized 4H-SiC detector has been fabricated to expand the application of SiC detectors.Its electrical properties and the charge collection performance of α particles are reported. The effective doping concentration of lightly doped 4H-SiC epitaxial layer is about 4.5\times10^{13}cm^{-3}, approaching the limit of the lowest doping level by the SiC epitaxial growth technique. The rise time of the graphene-optimized ring electrode detector is reduced by 24% at 200 V, compared to ring electrode detector. The charge collection efficiency (CCE) of graphene-optimized 4H-SiC PIN is 99.22%. When the irradiation dose is 2\times10^{11} n_{eq}/cm^2, the irradiation has no significant impact on the rise time and uniformity of the rise time for the graphene-optimized 4H-SiC detectors. This study proves that graphene has a certain radiation resistance. Graphene-optimized 4H-SiC detectors can not only reduce the signal rise time, but also improve uniformity of signal rise time and stability of charge collection. This research will expand the application of graphene-based 4H-SiC detectors in fields such as low energy ions, X-ray, UV light detection, particle physics, medical dosimetry and heavy-ion detection.

physics.ins-det

Towards Tight Robust Coresets for $k$-Medians Clustering

This paper considers coresets for the robust $k$-medians problem with $m$ outliers, and new constructions in various metric spaces are obtained. Specifically, for metric spaces with a bounded VC or doubling dimension $d$, the coreset size is $O(m) + \tilde{O}(kd\varepsilon^{-2})$, which is optimal up to logarithmic factors. For Euclidean spaces, the coreset size is $O(m\varepsilon^{-1}) + \tilde{O}(\min\{k^{4/3}\varepsilon^{-2}, k\varepsilon^{-3}\})$, improving upon a recent result by Jiang and Lou (ICALP 2025). These results also extend to robust $(k,z)$-clustering, yielding, for VC and doubling dimension, a coreset size of $O(m) + \tilde{O}(kd\varepsilon^{-2z})$ with the optimal linear dependence on $m$. This extended result improves upon the earlier work of Huang et al. (SODA 2025). The techniques introduce novel dataset decompositions, enabling chaining arguments to be applied jointly across multiple components.

cs.DS

Joint angle based learning to refine kinematic human pose estimation

Marker-free human pose estimation (HPE) has found increasing applications in various fields. Current HPE suffers from occasional errors in keypoint recognition and random fluctuation in keypoint trajectories when analyzing kinematic human poses. The performance of existing deep learning-based models for HPE refinement is considerably limited by inaccurate training datasets in which the keypoints are manually annotated. This paper proposed a novel method to overcome the difficulty, in which the key techniques include: (i) A robust joint angle-based description of kinematic human poses; (ii) Approximating temporal variation of joint angles using high order Fourier series to get reliable "ground truth"; (iii) A bidirectional recurrent network is designed as a post-processing module to refine the estimation of single image-based HPE models. Trained with the high-quality dataset constructed using our method, the network demonstrates outstanding performance to correct wrongly recognized joints and smooth their spatiotemporal trajectories. Tests show that joint angle-based refinement (JAR) outperforms the state-of-the-art HPE refinement network in challenging cases like figure skating and breaking. JAR also demonstrates great potential to rectify existing datasets.

cs.CV

Stable Charge Collection and Sub-45 ps Time Resolution in a 4H-SiC PIN Detector Irradiated With Low Fluence 16.5 MeV/u Ta Ions

A silicon carbide PIN detector was fabricated and its radiation tolerance under Ta heavy ion irradiation of 2370 MeV was evaluated. Its electrical properties, charge collection performance and time resolution of $β$-particles ($^{90}$Sr) are reported. The leakage currents for unirradiated and irradiated 4H-SiC PIN detectors are $1.47 \times 10^{-10}$~A @ 300 V and 1.49~$\times$ 10$^{-10}$A@ 300 V. The effective doping concentrations for unirradiated and irradiated 4H-SiC PIN detectors are $6.23\times 10^{13}$~cm$^{-3}$ and $6.13\times 10^{13}$~cm$^{-3}$. The irradiated detector exhibits good electrical performance and stable device architecture. The 4H-SiC PIN detector exhibits a charge collection efficiency (CCE) of 99.24\% under Ta Heavy Ion Irradiation. The time resolutions of the detector before and after irradiation are 40 ps and 45 ps, respectively. Experimental results indicate that the CCE and time resolution performance exhibit good stability before and after irradiation. These results demonstrate stable performance under Ta heavy ion irradiation, highlighting the detectors potential for radiation-hard applications in high-energy physics, space missions, and nuclear reactor monitoring.

physics.ins-det

Stability of Charge Collection Efficiency in a Novel Graphene-Optimized Silicon Carbide Detector Under 160 keV X-Ray Irradiation

A novel graphene-optimized silicon carbide PIN detector was fabricated. Its electrical properties, charge collection performance and signal rise time were evaluated under non-irradiated conditions and under X-ray irradiation with an energy of 160 keV at doses of 0.1 MGy and 1 MGy. The leakage currents of the detectors under non-irradiated, 0.1 MGy, and 1 MGy irradiation conditions are approximately 1.45e-10 A, 1.51e-10 A, and 1.57e-10 A, respectively. The effective doping concentration of the detector is approximately 8.08e13 cm^-3 before and after irradiation, with no significant change. The rise times of the signals from alpha particles signal detected by the detector under unirradiated, 0.1 MGy, and 1 MGy X-ray irradiation conditions are 336 ps, 368 ps, and 387 ps, respectively. The rise times of the beta particles signal detected by the detector under unirradiated, 0.1 MGy, and 1 MGy X-ray irradiation conditions are 342 ps, 375 ps, and 398 ps, respectively. After 0.1 MGy and 1 MGy X-ray irradiation, the charge collection efficiencies (CCEs) of the detector for alpha particles are 97.2% and 90.0%, respectively; for beta particles, they are 100.0% and 97.0%, respectively. Experiments confirm that 160 keV X-ray irradiation may not cause significant displacement damage in the 4H-SiC, and the minor performance degradation may be attributed to ionization induced changes in the graphene electrode. The detector exhibits excellent charge collection performance and fast time response. These results demonstrate stable performance under extreme X-ray exposure, highlighting the detector's potential for radiation-hard applications in high-energy physics, space missions, and nuclear reactor monitoring.

physics.ins-det

Stability of Charge Collection Efficiency and Time Resolution in a Novel Ultra-fast Graphene-Optimized Silicon Carbide Detector Under X-ray Irradiation

A graphene-optimized silicon carbide PIN detector was fabricated and its radiation tolerance under X-ray irradiation of 160 keV was evaluated. Its electrical properties, charge collection performance and time resolution of beta-particles (90Sr) are reported. After 1 MGy irradiation, the detector maintains an ultralow leakage current of approximately 2.2e-10 A @ 300 V and the C-V characteristics are basically consistent with full depletion at 120V. The time resolution of the graphene-optimized silicon carbide detector is 58.0 ps. The time resolution is comparable to that of state-of-the-art 4H-SiC low-gain avalanche detectors (LGADs). The G/RE 4H-SiC PIN detector exhibits outstanding time resolution performance. Compared with the time resolution of the RE 4H-SiC PIN detector, the time resolution of the G/RE 4H-SiC PIN detector has decreased by 39.6%. This demonstrates the significance of the graphene electrode design. The graphene detector exhibits a charge collection efficiency (CCE) of 99.24% after X-ray irradiation, along with excellent stability. The graphene-optimized silicon carbide detector maintains good timing resolution: 58.0ps before and 64.0ps after X-ray irradiation. Experimental results indicate that the CCE and time resolution performance exhibit good stability before and after irradiation. These results demonstrate stable performance under extreme X-ray exposure, highlighting the detectors potential for radiation-hard applications in high-energy physics, space missions, and nuclear reactor monitoring.

physics.ins-det

A Mechanistic Analysis of Sim-and-Real Co-Training in Generative Robot Policies

Co-training, which combines limited in-domain real-world data with abundant surrogate data such as simulation or cross-embodiment robot data, is widely used for training generative robot policies. Despite its empirical success, the mechanisms that determine when and why co-training is effective remain poorly understood. We investigate the mechanism of sim-and-real co-training through theoretical analysis and empirical study, and identify two intrinsic effects governing performance. The first, \textbf{``structured representation alignment"}, reflects a balance between cross-domain representation alignment and domain discernibility, and plays a primary role in downstream performance. The second, the \textbf{``importance reweighting effect"}, arises from domain-dependent modulation of action weighting and operates at a secondary level. We validate these effects with controlled experiments on a toy model and extensive sim-and-sim and sim-and-real robot manipulation experiments. Our analysis offers a unified interpretation of recent co-training techniques and motivates a simple method that consistently improves upon prior approaches. More broadly, our aim is to examine the inner workings of co-training and to facilitate research in this direction.

cs.RO

Automatic Road Subsurface Distress Recognition from Ground Penetrating Radar Images using Deep Learning-based Cross-verification

Ground penetrating radar (GPR) has become a rapid and non-destructive solution for road subsurface distress (RSD) detection. However, recognizing RSD from GPR images is labor-intensive and heavily relies on the expertise of inspectors. Deep learning-based automatic RSD recognition, though ameliorating the burden of data processing, suffers from insufficient capability to recognize defects. In this study, a novel cross-verification strategy was proposed to fully exploit the complementary abilities of region proposal networks in object recognition from different views of GPR images. Following this strategy, three YOLO-based models were used to detect the RSD (voids and loose structures) and manholes. Each model was trained with a specific view of 3D GPR dataset, which contains rigorously validated 2134 samples of diverse types obtained through field scanning. The cross-verification strategy achieves outstanding accuracy with a recall of over 98.6% in the tests using real field-scanning data. Field tests also show that deep learning-based automatic RSD recognition can reduce the human labor of inspection by around 90%.

cs.CV

Time Resolution of a Novel Ultra-fast Graphene-Optimized 4H-SiC PIN

Silicon carbide detectors exhibit good detection performance such as fast time resolution, high radiation tolerances, high breakdown voltage and low temperature sensitivity and have been studied for detection applications. Meanwhile, transient current technique (TCT) is a direct and effective method to evaluate the time resolution of semiconductor detectors. Conventional metal electrodes for TCT testing employ window structures, which lead to non-uniform electric field distribution and deteriorated time resolution. Graphene features high optical transmittance, ultrahigh carrier mobility, and excellent radiation resistance, making it an ideal transparent electrode material for semiconductor detectors. In this work, a graphene-optimized ring electrode (G/RE) 4H-SiC PIN detector and a reference ring electrode (RE) 4H-SiC PIN detector are fabricated. TCT measurements demonstrate that graphene integration improves the time resolution consistency, reducing the time resolution from 38 ps (reference RE detector) to 21 ps (G/RE detector) at the maximum scanning distance, while also achieving effective noise suppression. The graphene integration improves the stability of time resolution by 87% compared to the reference detector. Notably, the achieved time resolution of 21 ps is comparable to that of state-of-the-art 4H-SiC low-gain avalanche detectors (LGADs), which typically exhibit time resolutions better than 35 ps under single minimum ionizing particle (MIP) equivalent injection conditions, further validating the effectiveness of graphene-based electrode design.

physics.ins-det

Residual Off-Policy RL for Finetuning Behavior Cloning Policies

Recent advances in behavior cloning (BC) have enabled impressive visuomotor control policies. However, these approaches are limited by the quality of human demonstrations, the manual effort required for data collection, and the diminishing returns from offline data. In comparison, reinforcement learning (RL) trains an agent through autonomous interaction with the environment and has shown remarkable success in various domains. Still, training RL policies directly on real-world robots remains challenging due to sample inefficiency, safety concerns, and the difficulty of learning from sparse rewards for long-horizon tasks, especially for high-degree-of-freedom (DoF) systems. We present a recipe that combines the benefits of BC and RL through a residual learning framework. Our approach leverages BC policies as black-box bases and learns lightweight per-step residual corrections via sample-efficient off-policy RL. We demonstrate that our method requires only sparse binary reward signals and can effectively improve manipulation policies on high-degree-of-freedom (DoF) systems in both simulation and the real world. In particular, we demonstrate, to the best of our knowledge, the first successful real-world RL training on a humanoid robot with dexterous hands. Our results demonstrate state-of-the-art performance in various vision-based tasks, pointing towards a practical pathway for deploying RL in the real world.

cs.RO

MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos

We aim to enable humanoid robots to efficiently solve new manipulation tasks from a few video examples. In-context learning (ICL) is a promising framework for achieving this goal due to its test-time data efficiency and rapid adaptability. However, current ICL methods rely on labor-intensive teleoperated data for training, which restricts scalability. We propose using human play videos -- continuous, unlabeled videos of people interacting freely with their environment -- as a scalable and diverse training data source. We introduce MimicDroid, which enables humanoids to perform ICL using human play videos as the only training data. MimicDroid extracts trajectory pairs with similar manipulation behaviors and trains the policy to predict the actions of one trajectory conditioned on the other. Through this process, the model acquired ICL capabilities for adapting to novel objects and environments at test time. To bridge the embodiment gap, MimicDroid first retargets human wrist poses estimated from RGB videos to the humanoid, leveraging kinematic similarity. It also applies random patch masking during training to reduce overfitting to human-specific cues and improve robustness to visual differences. To evaluate few-shot learning for humanoids, we introduce an open-source simulation benchmark with increasing levels of generalization difficulty. MimicDroid outperformed state-of-the-art methods and achieved nearly twofold higher success rates in the real world. Additional materials can be found on: ut-austin-rpl.github.io/MimicDroid

cs.RO

VQualA 2025 Challenge on Face Image Quality Assessment: Methods and Results

Face images play a crucial role in numerous applications; however, real-world conditions frequently introduce degradations such as noise, blur, and compression artifacts, affecting overall image quality and hindering subsequent tasks. To address this challenge, we organized the VQualA 2025 Challenge on Face Image Quality Assessment (FIQA) as part of the ICCV 2025 Workshops. Participants created lightweight and efficient models (limited to 0.5 GFLOPs and 5 million parameters) for the prediction of Mean Opinion Scores (MOS) on face images with arbitrary resolutions and realistic degradations. Submissions underwent comprehensive evaluations through correlation metrics on a dataset of in-the-wild face images. This challenge attracted 127 participants, with 1519 final submissions. This report summarizes the methodologies and findings for advancing the development of practical FIQA approaches.

cs.CV

Quasi-Static IRS: 3D Shaped Beamforming for Area Coverage Enhancement

Intelligent reflecting surface (IRS) is a promising paradigm to reconfigure the wireless environment for enhanced communication coverage and quality. However, to compensate for the double pathloss effect, massive IRS elements are required, raising concerns on the scalability of cost and complexity. This paper introduces a new architecture of quasi-static IRS (QS-IRS), which tunes element phases via mechanical adjustment or manually re-arranging the array topology. QS-IRS relies on massive production/assembly of purely passive elements only, and thus is suitable for ultra low-cost and large-scale deployment to enhance long-term coverage. To achieve this end, an IRS-aided area coverage problem is formulated, which explicitly considers the element radiation pattern (ERP), with the newly introduced shape masks for the mainlobe, and the sidelobe constraints to reduce energy leakage. An alternating optimization (AO) algorithm based on the difference-of-convex (DC) and successive convex approximation (SCA) procedure is proposed, which achieves shaped beamforming with power gains close to that of the joint optimization algorithm, but with significantly reduced computational complexity.

cs.IT

Neural Channel Knowledge Map Assisted Scheduling Optimization of Active IRSs in Multi-User Systems

Intelligent Reflecting Surfaces (IRSs) have potential for significant performance gains in next-generation wireless networks but face key challenges, notably severe double-pathloss and complex multi-user scheduling due to hardware constraints. Active IRSs partially address pathloss but still require efficient scheduling in cell-level multi-IRS multi-user systems, whereby the overhead/delay of channel state acquisition and the scheduling complexity both rise dramatically as the user density and channel dimensions increase. Motivated by these challenges, this paper proposes a novel scheduling framework based on neural Channel Knowledge Map (CKM), designing Transformer-based deep neural networks (DNNs) to predict ergodic spectral efficiency (SE) from historical channel/throughput measurements tagged with user positions. Specifically, two cascaded networks, LPS-Net and SE-Net, are designed to predict link power statistics (LPS) and ergodic SE accurately. We further propose a low-complexity Stable Matching-Iterative Balancing (SM-IB) scheduling algorithm. Numerical evaluations verify that the proposed neural CKM significantly enhances prediction accuracy and computational efficiency, while the SM-IB algorithm effectively achieves near-optimal max-min throughput with greatly reduced complexity.

cs.IT

NTIRE 2025 Challenge on Short-form UGC Video Quality Assessment and Enhancement: Methods and Results

This paper presents a review for the NTIRE 2025 Challenge on Short-form UGC Video Quality Assessment and Enhancement. The challenge comprises two tracks: (i) Efficient Video Quality Assessment (KVQ), and (ii) Diffusion-based Image Super-Resolution (KwaiSR). Track 1 aims to advance the development of lightweight and efficient video quality assessment (VQA) models, with an emphasis on eliminating reliance on model ensembles, redundant weights, and other computationally expensive components in the previous IQA/VQA competitions. Track 2 introduces a new short-form UGC dataset tailored for single image super-resolution, i.e., the KwaiSR dataset. It consists of 1,800 synthetically generated S-UGC image pairs and 1,900 real-world S-UGC images, which are split into training, validation, and test sets using a ratio of 8:1:1. The primary objective of the challenge is to drive research that benefits the user experience of short-form UGC platforms such as Kwai and TikTok. This challenge attracted 266 participants and received 18 valid final submissions with corresponding fact sheets, significantly contributing to the progress of short-form UGC VQA and image superresolution. The project is publicly available at https://github.com/lixinustc/KVQE- ChallengeCVPR-NTIRE2025.

eess.IV

Sim-and-Real Co-Training: A Simple Recipe for Vision-Based Robotic Manipulation

Large real-world robot datasets hold great potential to train generalist robot models, but scaling real-world human data collection is time-consuming and resource-intensive. Simulation has great potential in supplementing large-scale data, especially with recent advances in generative AI and automated data generation tools that enable scalable creation of robot behavior datasets. However, training a policy solely in simulation and transferring it to the real world often demands substantial human effort to bridge the reality gap. A compelling alternative is to co-train the policy on a mixture of simulation and real-world datasets. Preliminary studies have recently shown this strategy to substantially improve the performance of a policy over one trained on a limited amount of real-world data. Nonetheless, the community lacks a systematic understanding of sim-and-real co-training and what it takes to reap the benefits of simulation data for real-robot learning. This work presents a simple yet effective recipe for utilizing simulation data to solve vision-based robotic manipulation tasks. We derive this recipe from comprehensive experiments that validate the co-training strategy on various simulation and real-world datasets. Using two domains--a robot arm and a humanoid--across diverse tasks, we demonstrate that simulation data can enhance real-world task performance by an average of 38%, even with notable differences between the simulation and real-world data. Videos and additional results can be found at https://co-training.github.io/

cs.RO

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

General-purpose robots need a versatile body and an intelligent mind. Recent advancements in humanoid robots have shown great promise as a hardware platform for building generalist autonomy in the human world. A robot foundation model, trained on massive and diverse data sources, is essential for enabling the robots to reason about novel situations, robustly handle real-world variability, and rapidly learn new tasks. To this end, we introduce GR00T N1, an open foundation model for humanoid robots. GR00T N1 is a Vision-Language-Action (VLA) model with a dual-system architecture. The vision-language module (System 2) interprets the environment through vision and language instructions. The subsequent diffusion transformer module (System 1) generates fluid motor actions in real time. Both modules are tightly coupled and jointly trained end-to-end. We train GR00T N1 with a heterogeneous mixture of real-robot trajectories, human videos, and synthetically generated datasets. We show that our generalist robot model GR00T N1 outperforms the state-of-the-art imitation learning baselines on standard simulation benchmarks across multiple robot embodiments. Furthermore, we deploy our model on the Fourier GR-1 humanoid robot for language-conditioned bimanual manipulation tasks, achieving strong performance with high data efficiency.

cs.RO

HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots

Humanoid whole-body control requires adapting to diverse tasks such as navigation, loco-manipulation, and tabletop manipulation, each demanding a different mode of control. For example, navigation relies on root velocity tracking, while tabletop manipulation prioritizes upper-body joint angle tracking. Existing approaches typically train individual policies tailored to a specific command space, limiting their transferability across modes. We present the key insight that full-body kinematic motion imitation can serve as a common abstraction for all these tasks and provide general-purpose motor skills for learning multiple modes of whole-body control. Building on this, we propose HOVER (Humanoid Versatile Controller), a multi-mode policy distillation framework that consolidates diverse control modes into a unified policy. HOVER enables seamless transitions between control modes while preserving the distinct advantages of each, offering a robust and scalable solution for humanoid control across a wide range of modes. By eliminating the need for policy retraining for each control mode, our approach improves efficiency and flexibility for future humanoid applications.

cs.RO