SearcharxivSearch

arXiv subjects

Yinan Chen

Publications and source records attributed to Yinan Chen.

At least 19 recordsLinked to original sources

Development, Evaluation, and Multicenter Clinical-Trial Application of an Artificial Intelligence-Assisted MRI Method for Quantitative Knee Cartilage Morphometry

Objective: To develop and evaluate an AI-assisted MRI method for quantitative knee cartilage morphometry in a multicenter phase III knee osteoarthritis trial. Methods: AI pre-segmentation used 3D full-resolution nnU-Net. Version 1.0 used separate femorotibial- and patellar-cartilage models, whereas version 2.0 used a unified three-class model trained on gold-standard annotations. Trial images then underwent two-reader correction and third-reader adjudication. Adjudicated masks were partitioned into medial/lateral femoral and tibial cartilage plus patellar cartilage. Cartilage volume was measured in physical coordinates, mean thickness by 3D ray tracing (3D-RT), and surface area with local thickness <1.5 mm by a 3D ray-based area method (3D-RBA). Evaluation included 1,189 phase III MRI examinations, reader agreement, 20 synthetic thinning models, and a 69-participant longitudinal comparison with 3D-PMA and three comparator thickness methods. Results: Overall pre-segmentation Dice was 0.964 +/- 0.030 (median 0.970), with 78.7% achieving Dice >=0.95. Inter-reader ICCs for cartilage volume were 0.959-0.995. In the 69-participant subset, total cartilage volume increased from 14,184.366 mm^3 at V0 to 15,359.345 mm^3 at V8; 3D-RBA and 3D-PMA decreased by 4.70% and 6.88%, and all four thickness measures were highest at V8. In 20 geometric experiments, MAPE was 5.73%, CCC 0.822, and Dice 0.956. The workflow was applied to 1,188 MRI examinations from 416 participants. From V0 to V8, the treatment group showed +3.45% total cartilage volume, +2.46% mean thickness, and -4.54% 3D-RBA, versus -2.08%, -1.32%, and +0.16% in controls. Conclusion: This workflow provided a reproducible MRI cartilage assessment framework for a multicenter KOA trial. Cross-method agreement and geometric validation supported 3D-RT and 3D-RBA for therapeutic efficacy evaluation.

eess.IV

Bang-bang protocol for nondispersive qubit readout

Fast, precise, and quantum-non-demolition (QND) readout of superconducting qubits is a fundamental component of high-fidelity quantum sensing and computation. Conventional approaches typically operate in the dispersive regime, where the qubit-resonator coupling $g$ is weak compared to the detuning $\Delta$. While exhibiting good QND properties, the readout rate is limited to $\sim g^2\sqrt{N}/\Delta\ll g$, where $N$ is the number of photons in the resonator. QND readout in the nondispersive regime, where the readout rate reaches its full potential $\sim g$, relies on parameter sweeps that may encounter resonances, leading to measurement-induced state transitions (MIST). In this work, we study a nondispersive readout protocol that replaces these sweeps by sudden quenches of the coupling constant, using a resonator that is preloaded with photons. We call this protocol bang-bang readout, and show that it realizes single-shot projective measurements. The fidelity and QNDness of the qubit post-measurement are remarkably high, with an error that decreases like $1/N$. To arrive at these findings, we develop an analytical theory for the dynamics and measurements of the Jaynes-Cummings (JC) model, including a systematic expansion of correction terms in powers of $1/\sqrt{N}$. We show that the protocol can also be implemented without preloading the resonator by instead strongly driving the qubit, e.g., with a classical flux drive.

quant-ph

RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States

Learning-based memory systems for self-evolving LLM agents face two tightly coupled challenges. First, trajectory-indexed utilities grow with the interaction history, thereby dispersing limited feedback over an ever-expanding state space. Second, because trajectory-level rewards are jointly assigned to co-retrieved memories, irrelevant experiences may receive misleading utility updates and consequently enter the memory-reward trap. To address these challenges, we introduce Reduced-Order Memory Reinforcement Learning (RoMeRL), which represents the growing trajectory-indexed utility space using a fixed-dimensional per-task memory state factorized by outcome polarity and memory dynamics. RoMeRL incorporates new experiences through a fixed set of semantic coordinates whose contents are updated or replaced over time, thereby concentrating feedback over a bounded utility support. Theoretically, we show that this reduced-order parameterization increases the average feedback received by each utility coordinate and characterize the steady-state occupancy of erroneous coordinates under a generic coordinate-transition model. Empirically, across ALFWorld and LifelongAgentBench, RoMeRL improves task performance, reduces the Cold-Q ratio by 80.0%, increases feedback density by approximately 6.0 times, reduces the maintained memory size by 84.4%, and cuts LLM calls by 21.1%. These results show that reduced-order utility states support efficient self-evolving agent memory while limiting persistent reward contamination. Code is available at: https://github.com/YOUNG-fnxm/RoMeRL

cs.LG

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation

While instruction-based video editing has seen significant progress, joint audio-visual editing remains constrained by the absence of dedicated datasets and benchmarks. To bridge this gap, we present JAVEdit-100k, the first large-scale, high-quality dataset tailored for instruction-guided joint audio-visual editing. Focusing on human-centric videos, JAVEdit-100k comprises approximately 100K editing triplets spanning five distinct categories, including subject editing and speech editing. This dataset is rigorously constructed via four meticulously designed generation pipelines, seamlessly paired with an agent-in-the-loop quality control mechanism. Furthermore, to address the lack of standardized evaluation within the field, we introduce JAVEditBench, a comprehensive benchmark featuring curated source videos and human-aligned instructions across all editing categories. Finally, we propose JAVEdit, a pioneering baseline model for instruction-guided joint audio-visual editing. Experiments show that \model\ outperforms all baselines on five of six evaluation metrics.

cs.CV

Quantum metrology via partial quantum error correction

We introduce a method for error-corrected quantum metrology where only partial quantum error correction (QEC) is needed to suppress local noise and maintain the probe states' super-standard-quantum-limit (super-SQL) sensing performance. This stands in contrast to the existing QEC-assisted sensing schemes in Phys. Rev. Lett. 112, 080801 (2014) and Phys. Rev. Lett. 112, 150802 (2014), where a probe state is encoded into the logical subspace of a quantum code and error correction involves measurements on all checks of the code. Here, we encode the probe states into superpositions of energetically different states of the underlying quantum code. For our probe states, error correction using a subset of checks is enough to suppress noise both before and after phase imprinting. We analyze the tradeoff in noise suppression. For noise parallel to our phase imprinter of weight $l$, we achieve a suppression of $p^\delta$ where $p$ is the noise strength and $\delta = \lfloor (l+1)/2 \rfloor$. We propose an adaptive imprinter weight increasing strategy to maintain super-SQL performance as we scale up the system. In all our examples, checks and phase imprinters are chosen to be local operators avoiding non-local connectivity.

quant-ph

CREWS: Collaborative Robust Edge WiFi Sensing with Asynchronous and Incomplete Observations

Existing collaborative WiFi sensing systems rely on perfect node synchronization and complete data availability. However, real-world edge deployments suffer from heterogeneous computing and network dropouts, leading to asynchronous and incomplete features. We propose CREWS, a robust collaborative sensing framework that inherently resists these network volatility. First, CREWS employs a topology-agnostic aggregator invariant to the arrival order and subset size of incoming features. Second, rather than discarding delayed observations, it utilizes a staleness-aware adaptive replay mechanism. By treating stale features from lagging nodes as system-induced hard samples, CREWS transforms synchronization delays into beneficial training regularization. We theoretically prove the joint convergence of this architecture and demonstrate how replay bounds the bias-variance trade-off. Extensive evaluations and an 8-node heterogeneous hardware testbed demonstrate its superior resilience. Under severe conditions i.e., 50\% transient dropout rate or out-of-distribution jitter, CREWS restricts accuracy degradation to merely 2.2 percentage points, substantially outperforming state-of-the-art baselines.

cs.NI

Quantum sensing with critical systems: impact of symmetry, imperfections, and decoherence

Entangled many-body states enable high-precision quantum sensing beyond the standard quantum limit. We develop interferometric sensing protocols based on quantum critical wavefunctions and compare their performance with Greenberger-Horne-Zeilinger (GHZ) and spin-squeezed states. Building on the idea of symmetries as a metrological resource, we introduce a symmetry-based algorithm to identify optimal measurement strategies. We illustrate this algorithm both for magnetic systems with internal symmetries and Rydberg-atom arrays with spatial symmetries. We study the robustness of criticality for quantum sensing under non-unitary deformations, symmetry-preserving and symmetry-breaking decoherence, and qubit loss -- identifying regimes where critical systems outperform GHZ states and showing that non-unitary deformation can even enhance sensing precision. Combined with recent results on log-depth preparation of critical wavefunctions, interferometric sensing in this setting appears increasingly promising.

quant-ph

A New Approach from Lattice of Subgroup Sets to Generalized Solvable Extension Formations

In this paper, we establish the decomposition of morphisms from lattice of subgroup sets to generalized solvable extension formations. To achieve this, we develop a unified framework involving maximal subgroup functors, generating formation morphism and contraction-extension functors. In particular, solvability-induced sets of maximal subgroups are determined and generating formation morphism gives rise to generalized solvable extension formations.

math.GR

IVEBench: Modern Benchmark Suite for Instruction-Guided Video Editing Assessment

Instruction-guided video editing has emerged as a rapidly advancing research direction, offering new opportunities for intuitive content transformation while also posing significant challenges for systematic evaluation. Existing video editing benchmarks fail to support the evaluation of instruction-guided video editing adequately and further suffer from limited source diversity, narrow task coverage and incomplete evaluation metrics. To address the above limitations, we introduce IVEBench, a modern benchmark suite specifically designed for instruction-guided video editing assessment. IVEBench comprises a diverse database of 600 high-quality source videos, spanning seven semantic dimensions, and covering video lengths ranging from 32 to 1,024 frames. It further includes 8 categories of editing tasks with 35 subcategories, whose prompts are generated and refined through large language models and expert review. Crucially, IVEBench establishes a three-dimensional evaluation protocol encompassing video quality, instruction compliance and video fidelity, integrating both traditional metrics and multimodal large language model-based assessments. Extensive experiments demonstrate the effectiveness of IVEBench in benchmarking state-of-the-art instruction-guided video editing methods, showing its ability to provide comprehensive and human-aligned evaluation outcomes.

cs.CV

Bosonic Entanglement and Quantum Sensing from Energy Transfer in two-tone Floquet Systems

Quantum-enhanced sensors, which surpass the standard quantum limit (SQL) and approach the fundamental precision limits dictated by quantum mechanics, are finding applications across a wide range of scientific fields. This quantum advantage becomes particularly significant when a large number of particles are included in the sensing circuit. Achieving such enhancement requires introducing and preserving entanglement among many particles, posing significant experimental challenges. In this work, we integrate concepts from Floquet theory and quantum information to design an entangler capable of generating the desired entanglement between two paths of a quantum interferometer. We demonstrate that our path-entangled states enable sensing beyond the SQL, reaching the fundamental Heisenberg limit (HL) of quantum mechanics. Moreover, we show that a decoding parity measurement maintains the HL when specific conditions from Floquet theory are satisfied$\unicode{x2013}$particularly those related to the periodic driving parameters that preserve entanglement during evolution. We address the effects of a priori phase uncertainty and imperfect transmission, showing that our method remains robust under realistic conditions. Finally, we propose a superconducting-circuit implementation of our sensor in the microwave regime, highlighting its potential for practical applications in high-precision measurements.

quant-ph

UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions

The quality of the video dataset (image quality, resolution, and fine-grained caption) greatly influences the performance of the video generation model. The growing demand for video applications sets higher requirements for high-quality video generation models. For example, the generation of movie-level Ultra-High Definition (UHD) videos and the creation of 4K short video content. However, the existing public datasets cannot support related research and applications. In this paper, we first propose a high-quality open-sourced UHD-4K (22.4\% of which are 8K) text-to-video dataset named UltraVideo, which contains a wide range of topics (more than 100 kinds), and each video has 9 structured captions with one summarized caption (average of 824 words). Specifically, we carefully design a highly automated curation process with four stages to obtain the final high-quality dataset: \textit{i)} collection of diverse and high-quality video clips. \textit{ii)} statistical data filtering. \textit{iii)} model-based data purification. \textit{iv)} generation of comprehensive, structured captions. In addition, we expand Wan to UltraWan-1K/-4K, which can natively generate high-quality 1K/4K videos with more consistent text controllability, demonstrating the effectiveness of our data curation.We believe that this work can make a significant contribution to future research on UHD video generation. UltraVideo dataset and UltraWan models are available at https://xzc-zju.github.io/projects/UltraVideo.

cs.CV

Multi-Prototype Embedding Refinement for Semi-Supervised Medical Image Segmentation

Medical image segmentation aims to identify anatomical structures at the voxel-level. Segmentation accuracy relies on distinguishing voxel differences. Compared to advancements achieved in studies of the inter-class variance, the intra-class variance receives less attention. Moreover, traditional linear classifiers, limited by a single learnable weight per class, struggle to capture this finer distinction. To address the above challenges, we propose a Multi-Prototype-based Embedding Refinement method for semi-supervised medical image segmentation. Specifically, we design a multi-prototype-based classification strategy, rethinking the segmentation from the perspective of structural relationships between voxel embeddings. The intra-class variations are explored by clustering voxels along the distribution of multiple prototypes in each class. Next, we introduce a consistency constraint to alleviate the limitation of linear classifiers. This constraint integrates different classification granularities from a linear classifier and the proposed prototype-based classifier. In the thorough evaluation on two popular benchmarks, our method achieves superior performance compared with state-of-the-art methods. Code is available at https://github.com/Briley-byl123/MPER.

eess.IV

Ruddlesden-Popper defects act as a free surface: role in formation and photophysical properties of CsPbI3

The perovskite semiconductor, CsPbI3, holds excellent promise for solar cell applications due to its suitable bandgap. However, achieving phase-stable CsPbI3 solar cells with high power conversion efficiency remains a major challenge. Ruddlesden-Popper (RP) defects have been identified in a range of perovskite semiconductors, including CsPbI3. However, there is limited understanding as to why they form or their impact on stability and photophysical properties. Here we increase the prevalence of RP defects with increased Cs-excess in vapour-deposited CsPbI3 thin films and observe superior structural stability but inferior photophysical properties. Significantly, using electron microscopy, we find that the atomic positions at the planar defect are comparable to those of a free surface, revealing their role in phase stabilisation. Density functional theory (DFT) calculations reveal the RP planes are electronically benign, however, antisites observed at RP turning points are likely to be malign. We therefore propose that increasing RP planes while reducing RP turning points could offer a breakthrough for improving both phase stability and photophysical performance. The formation mechanism revealed here may well apply more generally to RP structures in other perovskite systems.

cond-mat.mtrl-sci

Image Inversion: A Survey from GANs to Diffusion and Beyond

Image inversion is a fundamental task in generative models, aiming to map images back to their latent representations to enable downstream applications such as editing, restoration, and style transfer. This paper provides a comprehensive review of the latest advancements in image inversion techniques, focusing on two main paradigms: Generative Adversarial Network (GAN) inversion and diffusion model inversion. We categorize these techniques based on their optimization methods. For GAN inversion, we systematically classify existing methods into encoder-based approaches, latent optimization approaches, and hybrid approaches, analyzing their theoretical foundations, technical innovations, and practical trade-offs. For diffusion model inversion, we explore training-free strategies, fine-tuning methods, and the design of additional trainable modules, highlighting their unique advantages and limitations. Additionally, we discuss several popular downstream applications and emerging applications beyond image tasks, identifying current challenges and future research directions. By synthesizing the latest developments, this paper aims to provide researchers and practitioners with a valuable reference resource, promoting further advancements in the field of image inversion. We keep track of the latest works at https://github.com/RyanChenYN/ImageInversion

cs.CV

Strongly Enhanced Electronic Bandstructure Renormalization by Light in Nanoscale Strained Regions of Monolayer MoS$_2$/Au(111) Heterostructures

Understanding and controlling the photoexcited quasiparticle (QP) dynamics in monolayer transition metal dichalcogenides lays the foundation for exploring the strongly interacting, non-equilibrium 2D quasiparticle and polaritonic states in these quantum materials and for harnessing the properties emerging from these states for optoelectronic applications. In this study, scanning tunneling microscopy/spectroscopy with light illumination at the tunneling junction is performed to investigate the QP dynamics in monolayer MoS$_2$ on an Au(111) substrate with nanoscale corrugations. The corrugations on the surface of the substrate induce nanoscale local strain in the overlaying monolayer MoS$_2$ single crystal, which result in energetically favorable spatial regions where photoexcited QPs, including excitons, trions, and electron-hole plasmas, accumulate. These strained regions exhibit pronounced electronic bandstructure renormalization as a function of the photoexcitation wavelength and intensity as well as the strain gradient, implying strong interplay among nanoscale structures, strain, and photoexcited QPs. In conjunction with the experimental work, we construct a theoretical framework that integrates non-uniform nanoscale strain into the electronic bandstructure of a monolayer MoS$_2$ lattice using a tight-binding approach combined with first-principle calculations. This methodology enables better understanding of the experimental observation of photoexcited QP localization in the nanoscale strain-modulated electronic bandstructure landscape. Our findings illustrate the feasibility of utilizing nanoscale architectures and optical excitations to manipulate the local electronic bandstructure of monolayer TMDs and to enhance the many-body interactions of excitons, which is promising for the development of nanoscale energy-adjustable optoelectronic and photonic technologies.

cond-mat.mtrl-sci

DiffSeg: A Segmentation Model for Skin Lesions Based on Diffusion Difference

Weakly supervised medical image segmentation (MIS) using generative models is crucial for clinical diagnosis. However, the accuracy of the segmentation results is often limited by insufficient supervision and the complex nature of medical imaging. Existing models also only provide a single outcome, which does not allow for the measurement of uncertainty. In this paper, we introduce DiffSeg, a segmentation model for skin lesions based on diffusion difference which exploits diffusion model principles to ex-tract noise-based features from images with diverse semantic information. By discerning difference between these noise features, the model identifies diseased areas. Moreover, its multi-output capability mimics doctors' annotation behavior, facilitating the visualization of segmentation result consistency and ambiguity. Additionally, it quantifies output uncertainty using Generalized Energy Distance (GED), aiding interpretability and decision-making for physicians. Finally, the model integrates outputs through the Dense Conditional Random Field (DenseCRF) algorithm to refine the segmentation boundaries by considering inter-pixel correlations, which improves the accuracy and optimizes the segmentation results. We demonstrate the effectiveness of DiffSeg on the ISIC 2018 Challenge dataset, outperforming state-of-the-art U-Net-based methods.

cs.CV

A Sentiment Analysis of Medical Text Based on Deep Learning

The field of natural language processing (NLP) has made significant progress with the rapid development of deep learning technologies. One of the research directions in text sentiment analysis is sentiment analysis of medical texts, which holds great potential for application in clinical diagnosis. However, the medical field currently lacks sufficient text datasets, and the effectiveness of sentiment analysis is greatly impacted by different model design approaches, which presents challenges. Therefore, this paper focuses on the medical domain, using bidirectional encoder representations from transformers (BERT) as the basic pre-trained model and experimenting with modules such as convolutional neural network (CNN), fully connected network (FCN), and graph convolutional networks (GCN) at the output layer. Experiments and analyses were conducted on the METS-CoV dataset to explore the training performance after integrating different deep learning networks. The results indicate that CNN models outperform other networks when trained on smaller medical text datasets in combination with pre-trained models like BERT. This study highlights the significance of model selection in achieving effective sentiment analysis in the medical domain and provides a reference for future research to develop more efficient model architectures.

cs.CL

Generative Software Engineering

The rapid development of deep learning techniques, improved computational power, and the availability of vast training data have led to significant advancements in pre-trained models and large language models (LLMs). Pre-trained models based on architectures such as BERT and Transformer, as well as LLMs like ChatGPT, have demonstrated remarkable language capabilities and found applications in Software engineering. Software engineering tasks can be divided into many categories, among which generative tasks are the most concern by researchers, where pre-trained models and LLMs possess powerful language representation and contextual awareness capabilities, enabling them to leverage diverse training data and adapt to generative tasks through fine-tuning, transfer learning, and prompt engineering. These advantages make them effective tools in generative tasks and have demonstrated excellent performance. In this paper, we present a comprehensive literature review of generative tasks in SE using pre-trained models and LLMs. We accurately categorize SE generative tasks based on software engineering methodologies and summarize the advanced pre-trained models and LLMs involved, as well as the datasets and evaluation metrics used. Additionally, we identify key strengths, weaknesses, and gaps in existing approaches, and propose potential research directions. This review aims to provide researchers and practitioners with an in-depth analysis and guidance on the application of pre-trained models and LLMs in generative tasks within SE.

cs.SE