SearcharxivSearch

arXiv subjects

Binbin Yang

Publications and source records attributed to Binbin Yang.

18 recordsLinked to original sources

Development, Evaluation, and Multicenter Clinical-Trial Application of an Artificial Intelligence-Assisted MRI Method for Quantitative Knee Cartilage Morphometry

Objective: To develop and evaluate an AI-assisted MRI method for quantitative knee cartilage morphometry in a multicenter phase III knee osteoarthritis trial. Methods: AI pre-segmentation used 3D full-resolution nnU-Net. Version 1.0 used separate femorotibial- and patellar-cartilage models, whereas version 2.0 used a unified three-class model trained on gold-standard annotations. Trial images then underwent two-reader correction and third-reader adjudication. Adjudicated masks were partitioned into medial/lateral femoral and tibial cartilage plus patellar cartilage. Cartilage volume was measured in physical coordinates, mean thickness by 3D ray tracing (3D-RT), and surface area with local thickness <1.5 mm by a 3D ray-based area method (3D-RBA). Evaluation included 1,189 phase III MRI examinations, reader agreement, 20 synthetic thinning models, and a 69-participant longitudinal comparison with 3D-PMA and three comparator thickness methods. Results: Overall pre-segmentation Dice was 0.964 +/- 0.030 (median 0.970), with 78.7% achieving Dice >=0.95. Inter-reader ICCs for cartilage volume were 0.959-0.995. In the 69-participant subset, total cartilage volume increased from 14,184.366 mm^3 at V0 to 15,359.345 mm^3 at V8; 3D-RBA and 3D-PMA decreased by 4.70% and 6.88%, and all four thickness measures were highest at V8. In 20 geometric experiments, MAPE was 5.73%, CCC 0.822, and Dice 0.956. The workflow was applied to 1,188 MRI examinations from 416 participants. From V0 to V8, the treatment group showed +3.45% total cartilage volume, +2.46% mean thickness, and -4.54% 3D-RBA, versus -2.08%, -1.32%, and +0.16% in controls. Conclusion: This workflow provided a reproducible MRI cartilage assessment framework for a multicenter KOA trial. Cross-method agreement and geometric validation supported 3D-RT and 3D-RBA for therapeutic efficacy evaluation.

eess.IV

DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior

Storyboard synthesis plays a crucial role in visual storytelling, aiming to generate coherent shot sequences that visually narrate cinematic events with consistent characters, scenes, and transitions. However, existing approaches are mostly adapted from text-to-image diffusion models, which struggle to maintain long-range temporal coherence, consistent character identities, and narrative flow across multiple shots. In this paper, we introduce DreamShot, a video generative model based storyboard framework that fully exploits powerful video diffusion priors for controllable multi-shot synthesis. DreamShot supports both Text-to-Shot and Reference-to-Shot generation, as well as story continuation conditioned on previous frames, enabling flexible and context-aware storyboard generation. By leveraging the spatial-temporal consistency inherent in video generative models, DreamShot produces visually and semantically coherent sequences with improved narrative fidelity and character continuity. Furthermore, DreamShot incorporates a multi-reference role conditioning module that accepts multiple character reference images and enforces identity alignment via a Role-Attention Consistency Loss, explicitly constraining attention between reference and generated roles. Extensive experiments demonstrate that DreamShot achieves superior scene coherence, role consistency, and generation efficiency compared to state-of-the-art text-to-image storyboard models, establishing a new direction toward controllable video model-driven visual storytelling.

cs.CV

UniVid: Pyramid Diffusion Model for High Quality Video Generation

Diffusion-based text-to-video generation (T2V) or image-to-video (I2V) generation have emerged as a prominent research focus. However, there exists a challenge in integrating the two generative paradigms into a unified model. In this paper, we present a unified video generation model (UniVid) with hybrid conditions of the text prompt and reference image. Given these two available controls, our model can extract objects' appearance and their motion descriptions from textual prompts, while obtaining texture details and structural information from image clues to guide the video generation process. Specifically, we scale up the pre-trained text-to-image diffusion model for generating temporally coherent frames via introducing our temporal-pyramid cross-frame spatial-temporal attention modules and convolutions. To support bimodal control, we introduce a dual-stream cross-attention mechanism, whose attention scores can be freely re-weighted for interpolation of between single and two modalities controls during inference. Extensive experiments showcase that our UniVid achieves superior temporal coherence on T2V, I2V and (T+I)2V tasks.

cs.CV

Unveiling Perceptual Artifacts: A Fine-Grained Benchmark for Interpretable AI-Generated Image Detection

Current AI-Generated Image (AIGI) detection approaches predominantly rely on binary classification to distinguish real from synthetic images, often lacking interpretable or convincing evidence to substantiate their decisions. This limitation stems from existing AIGI detection benchmarks, which, despite featuring a broad collection of synthetic images, remain restricted in their coverage of artifact diversity and lack detailed, localized annotations. To bridge this gap, we introduce a fine-grained benchmark towards eXplainable AI-Generated image Detection, named X-AIGD, which provides pixel-level, categorized annotations of perceptual artifacts, spanning low-level distortions, high-level semantics, and cognitive-level counterfactuals. These comprehensive annotations facilitate fine-grained interpretability evaluation and deeper insight into model decision-making processes. Our extensive investigation using X-AIGD provides several key insights: (1) Existing AIGI detectors demonstrate negligible reliance on perceptual artifacts, even at the most basic distortion level. (2) While AIGI detectors can be trained to identify specific artifacts, they still substantially base their judgment on uninterpretable features. (3) Explicitly aligning model attention with artifact regions can increase the interpretability and generalization of detectors. The data and code are available at: https://github.com/Coxy7/X-AIGD.

cs.CV

Magneto-optical spectroscopy based on pump-probe strobe light

We demonstrate a pump-probe strobe light spectroscopy for sensitive detection of magneto-optical dynamics in the context of hybrid magnonics. The technique uses a combinatorial microwave-optical pump-probe scheme, leveraging both the high-energy resolution of microwaves and the high-efficiency detection using optical photons. In contrast to conventional stroboscopy using a continuous-wave light, we apply microwave and optical pulses with varying pulse widths, and demonstrate magnetooptical detection of magnetization dynamics in Y3Fe5O12 films. The detected magneto-optical signals strongly depend on the characteristics of both the microwave and the optical pulses as well as their relative time delays. We show that good magneto-optical sensitivity and coherent stroboscopic character are maintained even at a microwave pump pulse of 1.5 ns and an optical probe pulse of 80 ps, under a 7 megahertz clock rate, corresponding to a pump-probe footprint of ~1% in one detection cycle. Our results show that time-dependent strobe light measurement of magnetization dynamics can be achieved in the gigahertz frequency range under a pump-probe detection scheme.

cond-mat.mes-hall

Tuning Magneto-Optical Zero-Reflection via Dual-Channel Hybrid Magnonics

Multi-channel coupling in hybrid systems makes an attractive testbed not only because of the distinct advantages entailed in each constituent mode, but also the opportunity to leverage interference among the various excitation pathways. Here, via combined analytical calculation and experiment, we demonstrate that the phase of the magnetization precession at the interface of a coupled yttrium iron garnet(YIG)/permalloy(Py) bilayer is collectively controlled by the microwave photon field torque and the interlayer exchange torque, manifesting a coherent, dual-channel excitation scheme that effectively tunes the magneto-optic spectrum. The different torque contributions vary with frequency, external bias field, and types of interlayer coupling between YIG and Py, which further results in destructive or constructive interferences between the two excitation channels, and hence, selective suppression or amplification of the hybridized magnon modes.

cond-mat.mtrl-sci

Phase-resolving spin-wave microscopy using infrared strobe light

The needs for sensitively and reliably probing magnetization dynamics have been increasing in various contexts such as studying novel hybrid magnonic systems, in which the spin dynamics strongly and coherently couple to other excitations, including microwave photons, light photons, or phonons. Recent advances in quantum magnonics also highlight the need for employing magnon phase as quantum state variables, which is to be detected and mapped out with high precision in on-chip micro- and nano-scale magnonic devices. Here, we demonstrate a facile optical technique that can directly perform concurrent spectroscopic and imaging functionalities with spatial- and phase-resolutions, using infrared strobe light operating at 1550-nm wavelength. To showcase the methodology, we spectroscopically studied the phase-resolved spin dynamics in a bilayer of Permalloy and Y3Fe5O12 (YIG), and spatially imaged the backward volume spin wave modes of YIG in the dipolar spin wave regime. Using the strobe light probe, the detected precessional phase contrast can be directly used to construct the map of the spin wave wavefront, in the continuous-wave regime of spin-wave propagation and in the stationary state, without needing any optical reference path. By selecting the applied field, frequency, and detection phase, the spin wave images can be made sensitive to the precession amplitude and phase. Our results demonstrate that infrared optical strobe light can serve as a versatile platform for magneto-optical probing of magnetization dynamics, with potential implications in investigating hybrid magnonic systems.

cond-mat.mtrl-sci

Hybrid Magnonics with Localized Spoof Surface Plasmon Polaritons

Hybrid magnonic systems have emerged as a promising direction for information propagation with preserved coherence. Due to high tunability of magnons, their interactions with microwave photons can be engineered to probe novel phenomena based on strong photon-magnon coupling. Improving the photon-magnon coupling strength can be done by tuning the structure of microwave resonators to better interact with the magnon counterpart. Planar resonators have been explored due to their potential for on-chip integration, but only common modes from stripline-based resonators have been used. Here, we present a microwave spiral resonator supporting the spoof localized surface plasmons (LSPs) and implement it to the investigation of photon-magnon coupling for hybrid magnonic applications. We showcase strong magnon-LSP photon coupling using a ferrimagnetic yttrium iron garnet sphere. We discuss the dependence of the spiral resonator design to the engineering capacity of the photon mode frequency and spatial field distributions, via both experiment and simulation. By the localized photon mode profiles, the resulting magnetic field concentrates near the surface dielectrics, giving rise to an enhanced magnetic filling factor. The strong coupling and large engineering space render the spoof LSPs an interesting contender in developing novel hybrid magnonic systems and functionalities.

cond-mat.mtrl-sci

Shape Synthesis and 3D Ceramic Printing of Non-canonical MIMO Dielectric Resonator Antennas

In this paper, we report a shape synthesis method for multi-mode dielectric resonator antennas (DRA) using characteristic mode theory (CMT) and a binary genetic algorithm (BGA). By including the antenna's characteristic modal responses (resonance frequencies and quality factors) in the cost function, the shape synthesis process is conducted without including excitation feeds. Through the optimization procedure, a non-canonical dielectric body is formed from tetrahedral elements to support the required modal properties. As a demonstration of the proposed design approach, two three-mode MIMO DRAs are synthesized from both a rectangular and a cylindrical volume to operate at 2.45 GHz. The synthesized MIMO DRA's complex shape (based on rectangle) is then fabricated using Nanoparticle jetted zirconia. A combination of probe and slot feeds are employed to excite the desired modes. Due to the orthogonality of the characteristic modes and the careful design of the feeding network, isolation $>20$ dB is achieved between all ports.

physics.app-ph

Zippo: Zipping Color and Transparency Distributions into a Single Diffusion Model

Beyond the superiority of the text-to-image diffusion model in generating high-quality images, recent studies have attempted to uncover its potential for adapting the learned semantic knowledge to visual perception tasks. In this work, instead of translating a generative diffusion model into a visual perception model, we explore to retain the generative ability with the perceptive adaptation. To accomplish this, we present Zippo, a unified framework for zipping the color and transparency distributions into a single diffusion model by expanding the diffusion latent into a joint representation of RGB images and alpha mattes. By alternatively selecting one modality as the condition and then applying the diffusion process to the counterpart modality, Zippo is capable of generating RGB images from alpha mattes and predicting transparency from input images. In addition to single-modality prediction, we propose a modality-aware noise reassignment strategy to further empower Zippo with jointly generating RGB images and its corresponding alpha mattes under the text guidance. Our experiments showcase Zippo's ability of efficient text-conditioned transparent image generation and present plausible results of Matte-to-RGB and RGB-to-Matte translation.

cs.CV

Combinatorial split-ring and spiral meta-resonator for efficient magnon-photon coupling

Developing hybrid materials and structures for electromagnetic wave engineering has been a promising route towards novel functionalities and tunabilities in many modern applications and perspectives in new quantum technologies. Despite its established success in engineering optical light and terahertz waves, the implementation of meta-resonators operating at the microwave band is still emerging, especially those that allow for on-chip integration and size miniaturization, which has turned out crucial to developing hybrid quantum systems at the microwave band. In this work, we present a microwave meta-resonator consisting of split-ring and and spiral resonators, and implement it to the investigation of photon-magnon coupling for hybrid magnonic applications. We observe broadened bandwidth to the split ring modes augmented by the additional spiral resonator, and, by coupling the modes to a magnetic sample, the resultant photon-magnon coupling can be significantly enhanced to more than ten-fold. Our work suggests that combinatorial, hybrid microwave resonators may be a promising approach towards future development and implementation of photon-magnon coupling in hybrid magnonic systems.

physics.app-ph

DiffCloth: Diffusion Based Garment Synthesis and Manipulation via Structural Cross-modal Semantic Alignment

Cross-modal garment synthesis and manipulation will significantly benefit the way fashion designers generate garments and modify their designs via flexible linguistic interfaces.Current approaches follow the general text-to-image paradigm and mine cross-modal relations via simple cross-attention modules, neglecting the structural correspondence between visual and textual representations in the fashion design domain. In this work, we instead introduce DiffCloth, a diffusion-based pipeline for cross-modal garment synthesis and manipulation, which empowers diffusion models with flexible compositionality in the fashion domain by structurally aligning the cross-modal semantics. Specifically, we formulate the part-level cross-modal alignment as a bipartite matching problem between the linguistic Attribute-Phrases (AP) and the visual garment parts which are obtained via constituency parsing and semantic segmentation, respectively. To mitigate the issue of attribute confusion, we further propose a semantic-bundled cross-attention to preserve the spatial structure similarities between the attention maps of attribute adjectives and part nouns in each AP. Moreover, DiffCloth allows for manipulation of the generated results by simply replacing APs in the text prompts. The manipulation-irrelevant regions are recognized by blended masks obtained from the bundled attention maps of the APs and kept unchanged. Extensive experiments on the CM-Fashion benchmark demonstrate that DiffCloth both yields state-of-the-art garment synthesis results by leveraging the inherent structural information and supports flexible manipulation with region consistency.

cs.CV

LAW-Diffusion: Complex Scene Generation by Diffusion with Layouts

Thanks to the rapid development of diffusion models, unprecedented progress has been witnessed in image synthesis. Prior works mostly rely on pre-trained linguistic models, but a text is often too abstract to properly specify all the spatial properties of an image, e.g., the layout configuration of a scene, leading to the sub-optimal results of complex scene generation. In this paper, we achieve accurate complex scene generation by proposing a semantically controllable Layout-AWare diffusion model, termed LAW-Diffusion. Distinct from the previous Layout-to-Image generation (L2I) methods that only explore category-aware relationships, LAW-Diffusion introduces a spatial dependency parser to encode the location-aware semantic coherence across objects as a layout embedding and produces a scene with perceptually harmonious object styles and contextual relations. To be specific, we delicately instantiate each object's regional semantics as an object region map and leverage a location-aware cross-object attention module to capture the spatial dependencies among those disentangled representations. We further propose an adaptive guidance schedule for our layout guidance to mitigate the trade-off between the regional semantic alignment and the texture fidelity of generated objects. Moreover, LAW-Diffusion allows for instance reconfiguration while maintaining the other regions in a synthesized image by introducing a layout-aware latent grafting mechanism to recompose its local regional semantics. To better verify the plausibility of generated scenes, we propose a new evaluation metric for the L2I task, dubbed Scene Relation Score (SRS) to measure how the images preserve the rational and harmonious relations among contextual objects. Comprehensive experiments demonstrate that our LAW-Diffusion yields the state-of-the-art generative performance, especially with coherent object relations.

cs.CV

Synthesis of General Decoupling Networks Using Transmission Lines

In this paper, we introduce a synthesis technique for transmission line based decoupling networks, which find application in coupled systems such as multiple-antenna systems and compact antenna arrays. Employing the generalized $\pi$-network and the transmission line analysis technique, we reduce the decoupling network design into simple matrix calculations. The synthesized decoupling network is essentially a generalized $\pi$-network with transmission lines at all branches. A standard electrical length of $3\lambda/8$ and $5\lambda/8$ are chosen to simplify the physical implementation, leaving the characteristic impedances of the transmission line branches the main design parameters. The advantage of this proposed decoupling network is that it can be implemented using transmission lines, ensuring better control on loss, performance consistency and higher power handling capability when compared with lumped components, and can be easily scaled for operation at different frequencies. A two-port microstrip antenna system at 1.2 GHz and a three-port monopole antenna system at 1 GHz are investigated respectively to demonstrate the validity of the proposed synthesis method, and perfect decoupling ($S_{21}<-50$dB) are achieved at both design frequencies.

eess.SY

Continual Object Detection via Prototypical Task Correlation Guided Gating Mechanism

Continual learning is a challenging real-world problem for constructing a mature AI system when data are provided in a streaming fashion. Despite recent progress in continual classification, the researches of continual object detection are impeded by the diverse sizes and numbers of objects in each image. Different from previous works that tune the whole network for all tasks, in this work, we present a simple and flexible framework for continual object detection via pRotOtypical taSk corrElaTion guided gaTing mechAnism (ROSETTA). Concretely, a unified framework is shared by all tasks while task-aware gates are introduced to automatically select sub-models for specific tasks. In this way, various knowledge can be successively memorized by storing their corresponding sub-model weights in this system. To make ROSETTA automatically determine which experience is available and useful, a prototypical task correlation guided Gating Diversity Controller(GDC) is introduced to adaptively adjust the diversity of gates for the new task based on class-specific prototypes. GDC module computes class-to-class correlation matrix to depict the cross-task correlation, and hereby activates more exclusive gates for the new task if a significant domain gap is observed. Comprehensive experiments on COCO-VOC, KITTI-Kitchen, class-incremental detection on VOC and sequential learning of four tasks show that ROSETTA yields state-of-the-art performance on both task-based and class-based continual object detection.

cs.CV

Fundamental Limits on Substructure Dielectric Resonator Antennas

We show theoretically that the characteristic modes of dielectric resonator antennas (DRAs) must be capacitive in the low frequency limit, and show that as a consequence of this constraint and the Poincaré Separation Theorem, the modes of any DRA consisting of partial elements of an encompassing super-structure cannot resonate at a frequency that is lower than that of the encompassing structure. Thus, design techniques relying on complex sub-structures to miniaturize the antenna, including topology optimization and meandered windings, cannot apply to DRAs. Due to the capacitive nature of the DRA modes, it is also shown that the Q factor of any DRA sub-structure will be bounded from below by that of the super-structure at frequencies below the first self-resonance of the super-structure. We demonstrate these bounding relations with numerical examples.

physics.app-ph

FD-FCN: 3D Fully Dense and Fully Convolutional Network for Semantic Segmentation of Brain Anatomy

In this paper, a 3D patch-based fully dense and fully convolutional network (FD-FCN) is proposed for fast and accurate segmentation of subcortical structures in T1-weighted magnetic resonance images. Developed from the seminal FCN with an end-to-end learning-based approach and constructed by newly designed dense blocks including a dense fully-connected layer, the proposed FD-FCN is different from other FCN-based methods and leads to an outperformance in the perspective of both efficiency and accuracy. Compared with the U-shaped architecture, FD-FCN discards the upsampling path for model fitness. To alleviate the problem of parameter explosion, the inputs of dense blocks are no longer directly passed to subsequent layers. This architecture of FD-FCN brings a great reduction on both memory and time consumption in training process. Although FD-FCN is slimmed down, in model competence it gains better capability of dense inference than other conventional networks. This benefits from the construction of network architecture and the incorporation of redesigned dense blocks. The multi-scale FD-FCN models both local and global context by embedding intermediate-layer outputs in the final prediction, which encourages consistency between features extracted at different scales and embeds fine-grained information directly in the segmentation process. In addition, dense blocks are rebuilt to enlarge the receptive fields without significantly increasing parameters, and spectral coordinates are exploited for spatial context of the original input patch. The experiments were performed over the IBSR dataset, and FD-FCN produced an accurate segmentation result of overall Dice overlap value of 89.81% for 11 brain structures in 53 seconds, with at least 3.66% absolute improvement of dice accuracy than state-of-the-art 3D FCN-based methods.

eess.IV

Depthwise Non-local Module for Fast Salient Object Detection Using a Single Thread

Recently deep convolutional neural networks have achieved significant success in salient object detection. However, existing state-of-the-art methods require high-end GPUs to achieve real-time performance, which makes them hard to adapt to low-cost or portable devices. Although generic network architectures have been proposed to speed up inference on mobile devices, they are tailored to the task of image classification or semantic segmentation, and struggle to capture intra-channel and inter-channel correlations that are essential for contrast modeling in salient object detection. Motivated by the above observations, we design a new deep learning algorithm for fast salient object detection. The proposed algorithm for the first time achieves competitive accuracy and high inference efficiency simultaneously with a single CPU thread. Specifically, we propose a novel depthwise non-local moudule (DNL), which implicitly models contrast via harvesting intra-channel and inter-channel correlations in a self-attention manner. In addition, we introduce a depthwise non-local network architecture that incorporates both depthwise non-local modules and inverted residual blocks. Experimental results show that our proposed network attains very competitive accuracy on a wide range of salient object detection datasets while achieving state-of-the-art efficiency among all existing deep learning based algorithms.

cs.CV