SearcharxivSearch

arXiv subjects

Yongming Li

Publications and source records attributed to Yongming Li.

At least 19 recordsLinked to original sources

How to improve the discrimination power of classically simulable measurements?

Classically simulable measurements (CSMs) constitute an important class of restricted measurements in the odd-prime-dimensional magic resource theory, referring to those measurements with positive discrete Wigner functions. Since their discrimination power is weaker than that of global measurements, it is necessary to study how to improve the discrimination power of CSMs. In this paper, we consider three methods to improve the discrimination power of CSMs, including adding magic resources, using quantum catalysts, and using quantum memories. Specifically, we relate measurements with positive discrete Wigner functions to completely positive Wigner-preserving measurement channels, thereby transforming the problem of improving the discrimination power of CSMs into the problem of determining how many magic resources are required to simulate quantum channels using free operations. Based on this, we derive the lower and upper bounds of the simulation cost. Moreover, we provide a concrete example for which these bounds coincide and prove that consumable magic resources can enhance the discrimination power of CSMs. Finally, we establish a no-go theorem, which shows that for discriminating a pair of states with positive Wigner functions, neither finite-dimensional quantum catalysts nor finite-dimensional quantum memories can improve the optimal success probability of discrimination using CSMs.

quant-ph

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation

Video generation is rapidly evolving from single-shot synthesis to complex multi-shot audio-video (MSAV) narratives to meet real-world demands. However, evaluating such frontier models remains a fundamental challenge. Existing benchmarks are limited in scope and data diversity, and rely on rigid evaluation pipelines, preventing systematic and reliable assessment of modern MSAV models. To bridge these gaps, we introduce MSAVBench, the first comprehensive benchmark and adaptive hybrid evaluation framework for multi-shot audio-video generation. Our benchmark spans four key dimensions, video, audio, shot, and reference, covering diverse task settings, varying shot counts of up to 15, and challenging non-realistic scenarios. Our evaluation framework improves robustness through an adaptive self-correction mechanism for shot segmentation, instance-wise rubrics for subjective metrics, and tool-grounded evidence extraction for complex judgments. Furthermore, MSAVBench achieves high alignment with human judgments, reaching a Spearman rank correlation of 91.5%. Our systematic evaluation of 19 state-of-the-art closed- and open-source models shows that current systems still struggle with director-level control and fine-grained audio-visual synchronization, while modular or agentic generation pipelines offer a promising path toward narrowing the gap between open- and closed-source models. The benchmark data and evaluation code are publicly available at https://github.com/ali-vilab/MSAVBench.

cs.CV

Skill-Evolving Grounded Reasoning for Free-Text Promptable 3D Medical Image Segmentation

Free-text promptable 3D medical image segmentation offers an intuitive and clinically flexible interaction paradigm. However, current methods are highly sensitive to linguistic variability: minor changes in phrasing can cause substantial performance degradation despite identical clinical intent. Existing approaches attempt to improve robustness through stronger vision-language fusion or larger vocabularies, yet they lack mechanisms to consistently align ambiguous free-form expressions with anatomically grounded representations. We propose Skill-Evolving grounded Reasoning (SEER), a novel framework for free-text promptable 3D medical image segmentation that explicitly bridges linguistic variability and anatomical precision through a reasoning-driven design. First, we curate the SEER-Trace dataset, which pairs raw clinical requests with image-grounded, skill-tagged reasoning traces, establishing a reproducible benchmark. Second, SEER constructs an evidence-aligned target representation via a vision-language reasoning chain that verifies clinical intent against image-derived anatomical evidence, thereby enforcing semantic consistency before voxel-level decoding. Third, we introduce SEER-Loop, a dynamic skill-evolving strategy that distills high-reward reasoning trajectories into reusable skill artifacts and progressively integrates them into subsequent inference, enabling structured self-refinement and improved robustness to diverse linguistic expressions. Extensive experiments demonstrate superior performance of SEER over state-of-the-art baselines. Under linguistic perturbations, SEER reduces performance variance by 81.94% and improves worst-case Dice by 18.60%. Project page: https://seer-medseg.github.io.

eess.IV

Possibilistic Computation Tree Logic: Decidability and Complete Axiomatization

Possibilistic computation tree Logic (PoCTL) is one kind of branching temporal logic combined with uncertain information in possibility theory, which was introduced in order to cope with the systematic verification on systems with uncertain information in possibility theory. There are two decision problems related to PoCTL: the model checking problem and the satisfiability problem. The model checking problem of PoCTL has been studied, while the satisfiability problem of PoCTL was not discussed. One of the purpose of this work is to study the satisfiability problem of PoCTL. By introducing some techniques to extract possibility information from PoCTL formulae and constructing their possibilistic Hintikka structures, we show that the satisfiability problem of PoCTL is decidable in exponential time. Furthermore, we give a complete axiomatization of PoCTL, which is another important inference problem of PoCTL.

cs.LO

Asymptotic stability of solitary waves for the 1D focusing cubic Schr\"odinger equation

We establish the full asymptotic stability of solitary wave solutions for the 1D focusing cubic Schr\"odinger equation on the line under small perturbations in weighted Sobolev spaces, building upon our results in [58]. The proof integrates the space-time resonances approach, based on the distorted Fourier transform, with modulation techniques to show modified scattering for the radiation term and convergence for the modulation parameters. A key challenge throughout the nonlinear analysis is the slow local decay of the radiation term, caused by threshold resonances in the linearized operator. The presence of favorable null structures in the quadratic nonlinearities mitigates this problem through the use of normal form transformations. Another essential step in the proof involves developing a variant of the local smoothing estimate that incorporates a moving center.

math.AP

Optimal Compilation Strategies for QFT Circuits in Neutral-Atom Quantum Computing

Neutral-atom quantum computing (NAQC) offers distinct advantages such as dynamic qubit reconfigurability, long coherence times, and high gate fidelities, making it a promising platform for scalable quantum computing. Despite these strengths, efficiently implementing quantum circuits like the Quantum Fourier Transform (QFT) remains a significant challenge due to atom movement overheads and connectivity constraints. This paper introduces optimal compilation strategies tailored to QFT circuits and NAQC systems, addressing these challenges for both linear and grid-like architectures. By minimizing atom movements, the proposed methods achieve theoretical lower bounds in atom movements while preserving high circuit fidelity. Comparative evaluations against state-of-the-art compilers demonstrate the superior performance of the proposed methods. These methods could serve as benchmarks for evaluating the performance of NAQC compilers.

quant-ph

All-Angle Scanning Leaky-Wave Antennas and Surface-Wave Routing by Reconfigurable Metasurfaces

In this work, we show that propagating waves can be fully converted into surface waves and back using geometrically periodic arrays of simple electrically small metal elements loaded by adjustable reactive loads. The proposed approach allows the creation of all-angle scanning leaky-wave antennas with perfect or even superdirective aperture efficiency at all scan angles. Moreover, it is possible to co-design such leaky-wave antenna arrays with surface-wave waveguides that can guide the received power to the load or to another leaky-wave antenna section. That second section can either reradiate the received power into any direction or perform some other transformation of the reradiated wave front, for example, focusing the power at a point. These and other functionalities are realized by global optimization of the reactive loads of array elements. This global optimization, together with the use of arrays with a subwavelength geometrical period, allows proper control over both propagating and evanescent-field distributions, ensuring theoretically perfect performance at arbitrary scan angles. The proposed technique can be used in antenna engineering and in advanced designs of reconfigurable intelligent surfaces.

physics.optics

Drug classification based on X-ray spectroscopy combined with machine learning

The proliferation of new types of drugs necessitates the urgent development of faster and more accurate detection methods. Traditional detection methods have high requirements for instruments and environments, making the operation complex. X-ray absorption spectroscopy, a non-destructive detection technique, offers advantages such as ease of operation, penetrative observation, and strong substance differentiation capabilities, making it well-suited for application in the field of drug detection and identification. In this study, we constructed a classification model using Convolutional Neural Networks (CNN), Support Vector Machines (SVM), and Particle Swarm Optimization (PSO) to classify and identify drugs based on their X-ray spectral profiles. In the experiments, we selected 14 chemical reagents with chemical formulas similar to drugs as samples. We utilized CNN to extract features from the spectral data of these 14 chemical reagents and used the extracted features to train an SVM model. We also utilized PSO to optimize two critical initial parameters of the SVM. The experimental results demonstrate that this model achieved higher classification accuracy compared to two other common methods, with a prediction accuracy of 99.14%. Additionally, the model exhibited fast execution speed, mitigating the drawback of a drastic increase in running time and efficiency reduction that may result from the direct fusion of PSO and SVM. Therefore, the combined approach of X-ray absorption spectroscopy with CNN, PSO, and SVM provides a rapid, highly accurate, and reliable classification and identification method for the field of drug detection, holding promising prospects for widespread application.

cs.CV

AMNet: An Acoustic Model Network for Enhanced Mandarin Speech Synthesis

This paper presents AMNet, an Acoustic Model Network designed to improve the performance of Mandarin speech synthesis by incorporating phrase structure annotation and local convolution modules. AMNet builds upon the FastSpeech 2 architecture while addressing the challenge of local context modeling, which is crucial for capturing intricate speech features such as pauses, stress, and intonation. By embedding a phrase structure parser into the model and introducing a local convolution module, AMNet enhances the model's sensitivity to local information. Additionally, AMNet decouples tonal characteristics from phonemes, providing explicit guidance for tone modeling, which improves tone accuracy and pronunciation. Experimental results demonstrate that AMNet outperforms baseline models in subjective and objective evaluations. The proposed model achieves superior Mean Opinion Scores (MOS), lower Mel Cepstral Distortion (MCD), and improved fundamental frequency fitting $F0 (R^2)$, confirming its ability to generate high-quality, natural, and expressive Mandarin speech.

cs.SD

BadRefSR: Backdoor Attacks Against Reference-based Image Super Resolution

Reference-based image super-resolution (RefSR) represents a promising advancement in super-resolution (SR). In contrast to single-image super-resolution (SISR), RefSR leverages an additional reference image to help recover high-frequency details, yet its vulnerability to backdoor attacks has not been explored. To fill this research gap, we propose a novel attack framework called BadRefSR, which embeds backdoors in the RefSR model by adding triggers to the reference images and training with a mixed loss function. Extensive experiments across various backdoor attack settings demonstrate the effectiveness of BadRefSR. The compromised RefSR network performs normally on clean input images, while outputting attacker-specified target images on triggered input images. Our study aims to alert researchers to the potential backdoor risks in RefSR. Codes are available at https://github.com/xuefusiji/BadRefSR.

cs.CV

In-Situ Mode: Generative AI-Driven Characters Transforming Art Engagement Through Anthropomorphic Narratives

Art appreciation serves as a crucial medium for emotional communication and sociocultural dialogue. In the digital era, fostering deep user engagement on online art appreciation platforms remains a challenge. Leveraging generative AI technologies, we present EyeSee, a system designed to engage users through anthropomorphic characters. We implemented and evaluated three modes (Narrator, Artist, and In-Situ) acting as a third-person narrator, a first-person creator, and first-person created objects, respectively, across two sessions: Narrative and Recommendation. We conducted a within-subject study with 24 participants. In the Narrative session, we found that the In-Situ and Artist modes had higher aesthetic appeal than the Narrator mode, although the Artist mode showed lower perceived usability. Additionally, from the Narrative to Recommendation session, we found that user-perceived relatability and believability within each interaction mode were sustained, but the user-perceived consistency and stereotypicality changed. Our findings suggest novel implications for applying anthropomorphic in-situ narratives to other educational settings.

cs.HC

Asymptotic stability of solitary waves for the 1D focusing cubic Schr\"odinger equation under even perturbations

We establish the full asymptotic stability of solitary waves for the focusing cubic Schr\"odinger equation on the line under small even perturbations in weighted Sobolev norms. The strategy of our proof combines a space-time resonances approach based on the distorted Fourier transform to capture modified scattering effects with modulation techniques to take into account the symmetries of the problem, namely the invariance under scaling and phase shifts. A major challenge is the slow local decay of the radiation term caused by the threshold resonances of the non-selfadjoint linearized matrix Schr\"odinger operator around the solitary waves. Our analysis hinges on two remarkable null structures that we uncover in the quadratic nonlinearities of the evolution equation for the radiation term as well as of the modulation equations.

math.AP

VNet: A GAN-based Multi-Tier Discriminator Network for Speech Synthesis Vocoders

Since the introduction of Generative Adversarial Networks (GANs) in speech synthesis, remarkable achievements have been attained. In a thorough exploration of vocoders, it has been discovered that audio waveforms can be generated at speeds exceeding real-time while maintaining high fidelity, achieved through the utilization of GAN-based models. Typically, the inputs to the vocoder consist of band-limited spectral information, which inevitably sacrifices high-frequency details. To address this, we adopt the full-band Mel spectrogram information as input, aiming to provide the vocoder with the most comprehensive information possible. However, previous studies have revealed that the use of full-band spectral information as input can result in the issue of over-smoothing, compromising the naturalness of the synthesized speech. To tackle this challenge, we propose VNet, a GAN-based neural vocoder network that incorporates full-band spectral information and introduces a Multi-Tier Discriminator (MTD) comprising multiple sub-discriminators to generate high-resolution signals. Additionally, we introduce an asymptotically constrained method that modifies the adversarial loss of the generator and discriminator, enhancing the stability of the training process. Through rigorous experiments, we demonstrate that the VNet model is capable of generating high-fidelity speech and significantly improving the performance of the vocoder.

eess.AS

Reconfigurable Superdirective and Superabsorptive Aperiodic Metasurfaces

In this paper, we present a general theory of aperiodic subwavelength arrays for controlling electromagnetic waves. The considered platform is formed by an array of electrically small loaded scatterers above a ground plane. While the array is geometrically periodic, all the loads can be in general different, so that the distributions of currents induced by plane waves are not periodic. To allow analytical solutions, we study arrays of thin wires or strips loaded by bulk loads. We demonstrate a practical way of creating tunable and reconfigurable multifunctional devices, on examples of superdirective beam splitters, focusing lenses establishing subdiffraction focusing, and absorbers going beyond perfect absorption. Contrary to the constraints imposed by the Floquet theorem in periodic counterparts like periodic metasurfaces or metagratings, where a fixed angle of incidence and period dictate the propagating directions of reflected waves, the proposed aperiodic designs allow controlling all propagating modes in any direction, which provides more freedom in manipulating electromagnetic waves. We hope that these results can be useful in multiple applications, such as telecommunications, radar techniques, signal processing, and energy harnessing.

physics.app-ph

Going Beyond Perfect Absorption: Reconfigurable Super-directive Absorbers

In the context of electromagnetic absorption, it is obvious that for an infinite planar periodic structure illuminated by a plane wave, the maximum attainable absorptance, i.e., perfect absorption, is theoretically limited to 100% of the incident power. Here we show that an intriguing possibility of overcoming this limit arises in finite-size resonant absorbing arrays. We present a comprehensive analysis of a simple two-dimensional strip array over an infinite perfectly conducting plane, where the strips are loaded by reconfigurable impedance loads. The absorptance is defined as the ratio of the dissipated power per unit length of the strips to the incident power on the unit length of the array width. The results show that even regular arrays of impedance strips can slightly overcome the limit of 100% absorptance, while using aperiodic arrays with optimized loads, absorptance can be significantly increased as compared with the scenario where the strips are identical. In principle, by tuning the reconfigurable loads, high super-unity absorptance can be realized for all angles of illumination.

physics.app-ph

Simultaneous High-Efficiency Anomalous Reflection and Angle of Arrival Sensing in Reconfigurable Intelligent Surfaces

In this work, we introduce reconfigurable intelligent surfaces designed to simultaneously perform reflection of single or multiple incident waves toward the receiver or receivers and sensing the angles of arrival. We achieve anomalous reflection with strongly suppressed parasitic scattering through an in-situ optimization of either the currents flowing on array elements or the far field in the receiver direction. The suppression of parasitic scattering allows us to accurately and without additional measurements or computations detect the angles of arrival of the illuminations through the spatial Fourier transform of the optimized current distribution through the controllable reactive loads. Therefore, unlike other recently proposed methods, our scheme of integrated sensing and communication does not require any pre-computed data sets and works for an arbitrary number of simultaneous illuminations. As a proof of principle, we design and analyze with full-wave simulations several reconfigurable intelligent surfaces consisting of an array of loaded wires above a ground plane.

physics.app-ph

Synergizing Human-AI Agency: A Guide of 23 Heuristics for Service Co-Creation with LLM-Based Agents

This empirical study serves as a primer for interested service providers to determine if and how Large Language Models (LLMs) technology will be integrated for their practitioners and the broader community. We investigate the mutual learning journey of non-AI experts and AI through CoAGent, a service co-creation tool with LLM-based agents. Engaging in a three-stage participatory design processes, we work with with 23 domain experts from public libraries across the U.S., uncovering their fundamental challenges of integrating AI into human workflows. Our findings provide 23 actionable "heuristics for service co-creation with AI", highlighting the nuanced shared responsibilities between humans and AI. We further exemplar 9 foundational agency aspects for AI, emphasizing essentials like ownership, fair treatment, and freedom of expression. Our innovative approach enriches the participatory design model by incorporating AI as crucial stakeholders and utilizing AI-AI interaction to identify blind spots. Collectively, these insights pave the way for synergistic and ethical human-AI co-creation in service contexts, preparing for workforce ecosystems where AI coexists.

cs.HC

Dispersive estimates for 1D matrix Schrödinger operators with threshold resonance

We establish dispersive estimates and local decay estimates for the time evolution of non-self-adjoint matrix Schrödinger operators with threshold resonances in one space dimension. In particular, we show that the decay rates in the weighted setting are the same as in the regular case after subtracting a finite rank operator corresponding to the threshold resonances. Such matrix Schrödinger operators naturally arise from linearizing a focusing nonlinear Schrödinger equation around a solitary wave. It is known that the linearized operator for the 1D focusing cubic NLS equation exhibits a threshold resonance. We also include an observation of a favorable structure in the quadratic nonlinearity of the evolution equation for perturbations of solitary waves of the 1D focusing cubic NLS equation.

math.AP