Searcharxiv⌕ Search

arXiv subjects

Guanzhong Wang

Publications and source records attributed to Guanzhong Wang.

At least 19 recordsLinked to original sources

Rendering-in-the-Loop: An Execution-Driven Agent for Interactive Web Development

Multimodal large language models have achieved remarkable progress in front-end web development, generating interactive webpages from multimodal references such as screenshots and interaction videos. However, existing work largely emphasizes visual metrics such as aesthetics and layout similarity, while overlooking the more critical validation of interactive functionality. We present RILA, an execution-driven agent that puts browser rendering in the loop, iteratively editing generated code from runtime interaction feedback. RILA introduces an Action Interaction Verification (AIV) module that replays the reference interaction trajectory on the generated webpage to collect grounded execution-aware observations, and an Execution-aware Rendering Score (ERS) that jointly measures interaction correctness and visual fidelity to guide iterative optimization. We further build an execution-verified data synthesis pipeline that produces diverse, high-quality training data, offering gains complementary to inference-time optimization. On IWR-Bench, RILA consistently improves both interaction and visual fidelity across foundation models. Notably, with our training pipeline, RILA lifts the compact Qwen3.5-9B backbone from 40.40% to 57.52%, surpassing far larger one-shot generators, including the 1T-parameter Kimi-K2.6 (55.61%) and the proprietary GPT-5.5 (55.74%).

cs.CV↗

Cross-Platform Autonomous Control of Minimal Kitaev Chains

Contemporary quantum devices are reaching new limits in size and complexity, allowing for the experimental exploration of emergent quantum modes. However, this increased complexity introduces significant challenges in device tuning and control. Here, we demonstrate autonomous tuning of emergent Poor Man's Majorana zero modes in a minimal realization of a Kitaev chain. We achieve this task using cross-platform transfer learning. First, we train a tuning model on a theory model. Next, we retrain it using a Kitaev chain realization in a two-dimensional electron gas. Finally, we apply this model to tune a Kitaev chain realized in quantum dots coupled through a semiconductor-superconductor section in a one-dimensional nanowire. Utilizing a convolutional neural network, we predict the tunneling and Cooper pair splitting rates from differential conductance measurements, employing these predictions to adjust the electrochemical potential to a Poor Man's Majorana sweet spot. The algorithm successfully converges to an immediate vicinity of a sweet spot (within 1.5 mV in 67.6% of attempts and within 4.5 mV in 80.9% of cases), typically finding a sweet spot in 45 minutes or less. This advancement is a stepping stone towards autonomous tuning of emergent modes in interacting systems, and towards foundational tuning machine learning models that can be deployed across a range of experimental platforms.

cond-mat.mes-hall↗

Sortblock: Similarity-Aware Feature Reuse for Diffusion Model

Diffusion Transformers (DiTs) have demonstrated remarkable generative capabilities, particularly benefiting from Transformer architectures that enhance visual and artistic fidelity. However, their inherently sequential denoising process results in high inference latency, limiting their deployment in real-time scenarios. Existing training-free acceleration approaches typically reuse intermediate features at fixed timesteps or layers, overlooking the evolving semantic focus across denoising stages and Transformer blocks.To address this, we propose Sortblock, a training-free inference acceleration framework that dynamically caches block-wise features based on their similarity across adjacent timesteps. By ranking the evolution of residuals, Sortblock adaptively determines a recomputation ratio, selectively skipping redundant computations while preserving generation quality. Furthermore, we incorporate a lightweight linear prediction mechanism to reduce accumulated errors in skipped blocks.Extensive experiments across various tasks and DiT architectures demonstrate that Sortblock achieves over 2$\times$ inference speedup with minimal degradation in output quality, offering an effective and generalizable solution for accelerating diffusion-based generative models.

cs.CV↗

Forecasting When to Forecast: Accelerating Diffusion Models with Confidence-Gated Taylor

Diffusion Transformers (DiTs) have demonstrated remarkable performance in visual generation tasks. However, their low inference speed limits their deployment in low-resource applications. Recent training-free approaches exploit the redundancy of features across timesteps by caching and reusing past representations to accelerate inference. Building on this idea, TaylorSeer instead uses cached features to predict future ones via Taylor expansion. However, its module-level prediction across all transformer blocks (e.g., attention or feedforward modules) requires storing fine-grained intermediate features, leading to notable memory and computation overhead. Moreover, it adopts a fixed caching schedule without considering the varying accuracy of predictions across timesteps, which can lead to degraded outputs when prediction fails. To address these limitations, we propose a novel approach to better leverage Taylor-based acceleration. First, we shift the Taylor prediction target from the module level to the last block level, significantly reducing the number of cached features. Furthermore, observing strong sequential dependencies among Transformer blocks, we propose to use the error between the Taylor-estimated and actual outputs of the first block as an indicator of prediction reliability. If the error is small, we trust the Taylor prediction for the last block; otherwise, we fall back to full computation, thereby enabling a dynamic caching mechanism. Empirical results show that our method achieves a better balance between speed and quality, achieving a 3.17x acceleration on FLUX, 2.36x on DiT, and 4.14x on Wan Video with negligible quality drop. The Project Page is \href{https://cg-taylor-acce.github.io/CG-Taylor/}{here.}

cs.CV↗

SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards

Building upon large language models (LLMs), recent large multimodal models (LMMs) unify cross-model understanding and generation into a single framework. However, LMMs still struggle to achieve accurate vision-language alignment, prone to generating text responses contradicting the visual input or failing to follow the text-to-image prompts. Current solutions require external supervision (e.g., human feedback or reward models) and only address unidirectional tasks-either understanding or generation. In this work, based on the observation that understanding and generation are naturally inverse dual tasks, we propose \textbf{SUDER} (\textbf{S}elf-improving \textbf{U}nified LMMs with \textbf{D}ual s\textbf{E}lf-\textbf{R}ewards), a framework reinforcing the understanding and generation capabilities of LMMs with a self-supervised dual reward mechanism. SUDER leverages the inherent duality between understanding and generation tasks to provide self-supervised optimization signals for each other. Specifically, we sample multiple outputs for a given input in one task domain, then reverse the input-output pairs to compute the dual likelihood within the model as self-rewards for optimization. Extensive experimental results on visual understanding and generation benchmarks demonstrate that our method can effectively enhance the performance of the model without any external supervision, especially achieving remarkable improvements in text-to-image tasks.

cs.AI↗

Single-shot parity readout of a minimal Kitaev chain

Protecting qubits from noise is essential for building reliable quantum computers. Topological qubits offer a route to this goal by encoding quantum information non-locally, using pairs of Majorana zero modes. These modes form a shared fermionic state whose occupation -- either even or odd -- defines the fermionic parity that encodes the qubit. Crucially, this parity cannot be accessed by any measurement that probes only one Majorana mode. This reflects the non-local nature of the encoding and its inherent protection against noise. A promising platform for realizing such qubits is the Kitaev chain, implemented in quantum dots coupled via superconductors. Even a minimal chain of two dots can host a pair of Majorana modes and store quantum information in their joint parity. Here we introduce a new technique for reading out this parity, based on quantum capacitance. This global probe senses the joint state of the chain and enables real-time, single-shot discrimination of the parity state. By comparing with simultaneous local charge sensing, we confirm that only the global signal resolves the parity. We observe random telegraph switching and extract parity lifetimes exceeding one millisecond. These results establish the essential readout step for time-domain control of Majorana qubits, resolving a long-standing experimental challenge.

cond-mat.mes-hall↗

PP-DocBee: Improving Multimodal Document Understanding Through a Bag of Tricks

With the rapid advancement of digitalization, various document images are being applied more extensively in production and daily life, and there is an increasingly urgent need for fast and accurate parsing of the content in document images. Therefore, this report presents PP-DocBee, a novel multimodal large language model designed for end-to-end document image understanding. First, we develop a data synthesis strategy tailored to document scenarios in which we build a diverse dataset to improve the model generalization. Then, we apply a few training techniques, including dynamic proportional sampling, data preprocessing, and OCR postprocessing strategies. Extensive evaluations demonstrate the superior performance of PP-DocBee, achieving state-of-the-art results on English document understanding benchmarks and even outperforming existing open source and commercial models in Chinese document understanding. The source code and pre-trained models are publicly available at \href{https://github.com/PaddlePaddle/PaddleMIX}{https://github.com/PaddlePaddle/PaddleMIX}.

cs.CV↗

PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding

This report introduces PP-DocBee2, an advanced version of the PP-DocBee, designed to enhance multimodal document understanding. Built on a large multimodal model architecture, PP-DocBee2 addresses the limitations of its predecessor through key technological improvements, including enhanced synthetic data quality, improved visual feature fusion strategy, and optimized inference methodologies. These enhancements yield an $11.4\%$ performance boost on internal benchmarks for Chinese business documents, and reduce inference latency by $73.0\%$ to the vanilla version. A key innovation of our work is a data quality optimization strategy for multimodal document tasks. By employing a large-scale multimodal pre-trained model to evaluate data, we apply a novel statistical criterion to filter outliers, ensuring high-quality training data. Inspired by insights into underutilized intermediate features in multimodal models, we enhance the ViT representational capacity by decomposing it into layers and applying a novel feature fusion strategy to improve complex reasoning. The source code and pre-trained model are available at \href{https://github.com/PaddlePaddle/PaddleMIX}{https://github.com/PaddlePaddle/PaddleMIX}.

cs.CV↗

Enabling Versatile Controls for Video Diffusion Models

Despite substantial progress in text-to-video generation, achieving precise and flexible control over fine-grained spatiotemporal attributes remains a significant unresolved challenge in video generation research. To address these limitations, we introduce VCtrl (also termed PP-VCtrl), a novel framework designed to enable fine-grained control over pre-trained video diffusion models in a unified manner. VCtrl integrates diverse user-specified control signals-such as Canny edges, segmentation masks, and human keypoints-into pretrained video diffusion models via a generalizable conditional module capable of uniformly encoding multiple types of auxiliary signals without modifying the underlying generator. Additionally, we design a unified control signal encoding pipeline and a sparse residual connection mechanism to efficiently incorporate control representations. Comprehensive experiments and human evaluations demonstrate that VCtrl effectively enhances controllability and generation quality. The source code and pre-trained models are publicly available and implemented using the PaddlePaddle framework at http://github.com/PaddlePaddle/PaddleMIX/tree/develop/ppdiffusers/examples/ppvctrl.

cs.CV↗

RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer

In this report, we present RT-DETRv2, an improved Real-Time DEtection TRansformer (RT-DETR). RT-DETRv2 builds upon the previous state-of-the-art real-time detector, RT-DETR, and opens up a set of bag-of-freebies for flexibility and practicality, as well as optimizing the training strategy to achieve enhanced performance. To improve the flexibility, we suggest setting a distinct number of sampling points for features at different scales in the deformable attention to achieve selective multi-scale feature extraction by the decoder. To enhance practicality, we propose an optional discrete sampling operator to replace the grid_sample operator that is specific to RT-DETR compared to YOLOs. This removes the deployment constraints typically associated with DETRs. For the training strategy, we propose dynamic data augmentation and scale-adaptive hyperparameters customization to improve performance without loss of speed. Source code and pre-trained models will be available at https://github.com/lyuwenyu/RT-DETR.

cs.CV↗

DETRs Beat YOLOs on Real-time Object Detection

The YOLO series has become the most popular framework for real-time object detection due to its reasonable trade-off between speed and accuracy. However, we observe that the speed and accuracy of YOLOs are negatively affected by the NMS. Recently, end-to-end Transformer-based detectors (DETRs) have provided an alternative to eliminating NMS. Nevertheless, the high computational cost limits their practicality and hinders them from fully exploiting the advantage of excluding NMS. In this paper, we propose the Real-Time DEtection TRansformer (RT-DETR), the first real-time end-to-end object detector to our best knowledge that addresses the above dilemma. We build RT-DETR in two steps, drawing on the advanced DETR: first we focus on maintaining accuracy while improving speed, followed by maintaining speed while improving accuracy. Specifically, we design an efficient hybrid encoder to expeditiously process multi-scale features by decoupling intra-scale interaction and cross-scale fusion to improve speed. Then, we propose the uncertainty-minimal query selection to provide high-quality initial queries to the decoder, thereby improving accuracy. In addition, RT-DETR supports flexible speed tuning by adjusting the number of decoder layers to adapt to various scenarios without retraining. Our RT-DETR-R50 / R101 achieves 53.1% / 54.3% AP on COCO and 108 / 74 FPS on T4 GPU, outperforming previously advanced YOLOs in both speed and accuracy. We also develop scaled RT-DETRs that outperform the lighter YOLO detectors (S and M models). Furthermore, RT-DETR-R50 outperforms DINO-R50 by 2.2% AP in accuracy and about 21 times in FPS. After pre-training with Objects365, RT-DETR-R50 / R101 achieves 55.3% / 56.2% AP. The project page: https://zhao-yian.github.io/RTDETR.

cs.CV↗

Signatures of Majorana protection in a three-site Kitaev chain

Majorana zero modes (MZMs) are non-Abelian excitations predicted to emerge at the edges of topological superconductors. One proposal for realizing a topological superconductor in one dimension involves a chain of spinless fermions, coupled through $p$-wave superconducting pairing and electron hopping. This concept is also known as the Kitaev chain. A minimal two-site Kitaev chain has recently been experimentally realized using quantum dots (QDs) coupled through a superconductor. In such a minimal chain, MZMs are quadratically protected against global perturbations of the QD electrochemical potentials. However, they are not protected from perturbations of the inter-QD couplings. In this work, we demonstrate that extending the chain to three sites offers greater protection than the two-site configuration. The enhanced protection is evidenced by the stability of the zero-energy modes, which is robust against variations in both the coupling amplitudes and the electrochemical potential variations in the constituent QDs. While our device offers all the desired control of the couplings it does not allow for superconducting phase control. Our experimental observations are in good agreement with numerical simulated conductances with phase averaging. Our work pioneers the development of longer Kitaev chains, a milestone towards topological protection in QD-based chains.

cond-mat.supr-con↗

Robust poor man's Majorana zero modes using Yu-Shiba-Rusinov states

The recent realization of a two-site Kitaev chain featuring "poor man's Majorana" states demonstrates a path forward in the field of topological superconductivity. Harnessing the potential of these states for quantum information processing, however, requires increasing their robustness to external perturbations. Here, we form a two-site Kitaev chain using proximitized quantum dots hosting Yu-Shiba-Rusinov states. The strong hybridization between such states and the superconductor enables the creation of poor man's Majorana states with a gap larger than $70 \mathrm{~μeV}$. It also greatly reduces the charge dispersion compared to Kitaev chains made with non-proximitized quantum dots. The large gap and reduced sensitivity to charge fluctuations will benefit qubit manipulation and demonstration of non-abelian physics using poor man's Majorana states.

cond-mat.mes-hall↗

Charge sensing the parity of an Andreev molecule

The proximity effect of superconductivity on confined states in semiconductors gives rise to various bound states such as Andreev bound states (ABSs), Andreev molecules and Majorana zero modes. While such bound states do not conserve charge, their Fermion parity is a good quantum number. One way to measure parity is to convert it to charge first, which is then sensed. In this work, we sense the charge of ABSs and Andreev molecules in an InSb-Al hybrid nanowire using an integrated quantum dot operated as a charge sensor. We show how charge sensing measurements can resolve the even and odd states of an Andreev molecule, without affecting the parity. Such an approach can be further utilized for parity measurements of Majorana zero modes in Kitaev chains based on quantum dots.

cond-mat.mes-hall↗

Realization of a minimal Kitaev chain in coupled quantum dots

Majorana bound states constitute one of the simplest examples of emergent non-Abelian excitations in condensed matter physics. A toy model proposed by Kitaev shows that such states can arise at the ends of a spinless $p$-wave superconducting chain. Practical proposals for its realization require coupling neighboring quantum dots in a chain via both electron tunneling and crossed Andreev reflection. While both processes have been observed in semiconducting nanowires and carbon nanotubes, crossed-Andreev interaction was neither easily tunable nor strong enough to induce coherent hybridization of dot states. Here we demonstrate the simultaneous presence of all necessary ingredients for an artificial Kitaev chain: two spin-polarized quantum dots in an InSb nanowire strongly coupled by both elastic co-tunneling and crossed Andreev reflection. We fine-tune this system to a sweet spot where a pair of Poor Man's Majorana states is predicted to appear. At this sweet spot, the transport characteristics satisfy the theoretical predictions for such a system, including pairwise correlation, zero charge and stability against local perturbations. While the simple system presented here can be scaled to simulate a full Kitaev chain with an emergent topological order, it can also be used imminently to explore relevant physics related to non-Abelian anyons.

cond-mat.mes-hall↗

Singlet and triplet Cooper pair splitting in hybrid superconducting nanowires

In most naturally occurring superconductors, electrons with opposite spins are paired up to form Cooper pairs. This includes both conventional $s$-wave superconductors such as aluminum as well as high-$T_\text{c}$, $d$-wave superconductors. Materials with intrinsic $p$-wave superconductivity, hosting Cooper pairs made of equal-spin electrons, have not been conclusively identified, nor synthesized, despite promising progress. Instead, engineered platforms where $s$-wave superconductors are brought into contact with magnetic materials have shown convincing signatures of equal-spin pairing. Here, we directly measure equal-spin pairing between spin-polarized quantum dots. This pairing is proximity-induced from an $s$-wave superconductor into a semiconducting nanowire with strong spin-orbit interaction. We demonstrate such pairing by showing that breaking a Cooper pair can result in two electrons with equal spin polarization. Our results demonstrate controllable detection of singlet and triplet pairing between the quantum dots. Achieving such triplet pairing in a sequence of quantum dots will be required for realizing an artificial Kitaev chain.

cond-mat.mes-hall↗

Crossed Andreev reflection and elastic co-tunneling in a three-site Kitaev chain nanowire device

The formation of a topological superconducting phase in a quantum-dot-based Kitaev chain requires nearest neighbor crossed Andreev reflection and elastic co-tunneling. Here we report on a hybrid InSb nanowire in a three-site Kitaev chain geometry - the smallest system with well-defined bulk and edge - where two superconductor-semiconductor hybrids separate three quantum dots. We demonstrate pairwise crossed Andreev reflection and elastic co-tunneling between both pairs of neighboring dots and show sequential tunneling processes involving all three quantum dots. These results are the next step towards the realization of topological superconductivity in long Kitaev chain devices with many coupled quantum dots.

cond-mat.mes-hall↗

Spin-filtered measurements of Andreev bound states

A semiconductor nanowire brought in proximity to a superconductor can form discrete, particle-hole symmetric states, known as Andreev bound states (ABSs). An ABS can be found in its ground or excited states of different spin and parity, such as a spin-zero singlet state with an even number of electrons or a spin-1/2 doublet state with an odd number of electrons. Considering the difference between spin of the even and odd states, spin-filtered measurements have the potential to reveal the underlying ground state. To directly measure the spin of single-electron excitations, we probe an ABS using a spin-polarized quantum dot that acts as a bipolar spin filter, in combination with a non-polarized tunnel junction in a three-terminal circuit. We observe a spin-polarized excitation spectrum of the ABS, which in some cases is fully spin-polarized, despite the presence of strong spin-orbit interaction in the InSb nanowires. In addition, decoupling the hybrid from the normal lead blocks the ABS relaxation resulting in a current blockade where the ABS is trapped in an excited state. Spin-polarized spectroscopy of hybrid nanowire devices, as demonstrated here, is proposed as an experimental tool to support the observation of topological superconductivity.

cond-mat.mes-hall↗