SearcharxivSearch

arXiv subjects

Yingying Wang

Publications and source records attributed to Yingying Wang.

At least 19 recordsLinked to original sources

Epidemiological Causal Graph Identification: Challenges, Identifiability and Algorithms

Causal discovery from observational data is fundamental to statistics and machine learning, yet determining causal direction without interventions necessitates structural assumptions. Existing identifiability research primarily focuses on continuous variables under additive noise models, often neglecting mixed datasets containing ordinal scales, counts, and continuous measurements. This paper investigates causal discovery in Directed Acyclic Graphs (DAGs) where nodes follow either an ordinal distribution (via an ordered logit model) or a regular one-parameter exponential family distribution. We prove that the edge direction between an ordinal and an exponential family node is distributionally identifiable for generic parameter values. Our findings generalize previous Ordinal-Poisson results to the broader exponential family. Computationally, we introduce a score-based exhaustive search and a masked continuous optimization framework using DAGMA for larger graphs. Numerical results validate the theory, recovering edge orientations within a Markov equivalence class that are unidentifiable under classical structural equation models.

cs.LG

Selecting among Missingness Models for Sequential Outcomes with Nonignorable Nonresponse

Sequential outcomes in longitudinal studies and multi-wave surveys may be missing not at random at both earlier and later occasions. We study graphical models in which at least one outcome is self-censoring and the response indicator for a later outcome may depend on either the earlier response indicator or the realized earlier outcome. These restrictions define two candidate families under which the relevant full-data distributions are identifiable and whose observed-data models overlap; graphs containing both dependencies form a broader class outside the prespecified comparison. For each candidate family, we establish identification of the full-data distribution under rank or completeness conditions and develop likelihood-based estimation. We then propose a two-stage Vuong-type procedure. The first stage determines whether the candidate models are observationally distinguishable; only after distinguishability is established does the second stage compare their Kullback--Leibler divergences from the true observed-data distribution. We also show that ordinary Wald inference remains asymptotically valid for the selected model-specific functional when the selected model has a fixed positive expected log-likelihood advantage. Simulations evaluate the two-stage procedure across graph classes. We finally apply the procedure to compare candidate missingness models in the Job Corps data and perform downstream functional estimation under the selected model.

stat.ME

Phase singularity enabled polarization switchable analog spatial differentiation in an atomic MoS$_2$ planar Fabry-Pérot cavity

Reconfigurable analog optical computing requires rapid and efficient switching between core mathematical operations, such as first- and second-order spatial differentiation. Here, we demonstrate a monolayer MoS$_2$ integrated a planar Fabry-Pérot (F-P) cavity that performs polarization switchable analog spatial differentiation under oblique incidence. By exploiting polarization dependent phase singularities, the device satisfies distinct optical transfer functions for different-order differentiation at the same operating condition. As a result, first-order and second-order derivatives of input images are experimentally realized by simply switching the incident polarization. These results establish the planar cavity as a compact reconfigurable spatial differentiator, whose computational order is controlled solely by light polarization. This approach provides a fast, convenient, and integration friendly strategy for tunable optical computing, and enables polarization switchable edge detection for image processing, with potential applications in real-time object recognition, feature extraction, and optical data compression.

physics.optics

Decentralized LLM-Driven Coordination of Acoustic Robots for Contactless Object Manipulation

Natural language interfaces can simplify interaction with multi-robot systems, especially when non-expert users need to issue high-level commands. Acoustic manipulation using ultrasonic phased arrays also enables contactless object handling for applications such as healthcare, laboratory automation, and precision transport. However, combining large language models (LLMs) with distributed acoustic mobile robots remains underexplored. This paper presents a decentralized framework for natural language-driven coordination of acoustic robots for contactless object manipulation. The system converts spoken instructions into executable multi-robot task plans using Whisper-based speech recognition, LLM-based semantic parsing, structured JSON task representation, and distributed scheduling. The JSON schema encodes robot assignments, temporal dependencies, spatial constraints, and synchronization requirements for sequential, parallel, and synchronized execution. The system is implemented on two TurtleBot3-based acoustic robots, each equipped with an ultrasonic phased array for contactless object transport. Experiments were conducted in three scenarios: sequential execution, parallel multi-robot transport, and synchronized cooperative manipulation. The system achieved task success rates of 96 percent for sequential tasks, 86 percent for parallel execution, and 70 percent for synchronized collaborative transport. These results show that natural language commands can be transformed into distributed robot actions for contactless manipulation, highlighting the potential of LLM-driven automation for human-robot interaction in distributed robotic systems.

cs.RO

TripleWin: Fixed-Point Equilibrium Pricing for Data-Model Coupled Markets

The rise of the machine learning (ML) model economy has intertwined markets for training datasets and pre-trained models. However, most pricing approaches still separate data and model transactions or rely on broker-centric pipelines that favor one side. Recent studies of data markets with externalities capture buyer interactions but do not yield a simultaneous and symmetric mechanism across data sellers, model producers, and model buyers. We propose a unified data-model coupled market that treats dataset and model trading as a single system. A supply-side mapping transforms dataset payments into buyer-visible model quotations, while a demand-side mapping propagates buyer prices back to datasets through Shapley-based allocation. Together, they form a closed loop that links four interactions: supply-demand propagation in both directions and mutual coupling among buyers and among sellers. We prove that the joint operator is a standard interference function (SIF), guaranteeing existence, uniqueness, and global convergence of equilibrium prices. Experiments demonstrate efficient convergence and improved fairness compared with broker-centric and one-sided baselines. The code is available on https://github.com/HongrunRen1109/Triple-Win-Pricing.

cs.LG

Bound preserving and mass conservative methods for the nonlocal Cahn-Hilliard equation with the logarithmic Flory-Huggins potential

It is well known that the exponential time differencing (ETD) method has been successfully applied to the classic Cahn-Hilliard equation with double well potential. However, this numerical method can not be extended to the Cahn-Hilliard equation with Flory-Huggins potential directly due to the fact that the the numerical solution may go beyond the physical interval which leads the non-physical solution. In this paper, we develop and analyze first- and second-order numerical schemes for the nonlocal Cahn-Hilliard equation with the classic Flory-Huggins energy potential. In more detail, the ETD method is firstly used to obtain the prediction solution, and then this prediction solution is corrected by the projection method to avoid non-physical solution. The proposed method is shown to preserve bound and mass conservation in discrete settings. In addition, error estimates for the numerical solution are rigorously obtained for both schemes. Extensive numerical tests and comparisons are conducted to demonstrate the performance of the proposed schemes.

math.NA

Design of Angular-offset Interstitial-Tube- Assisted Hollow-Core Fibers with Ultrahigh Mode Purity and Ultralow Loss

Antiresonant hollow-core fibres (AR-HCFs) have recently reached attenuation far below the Rayleigh-scattering limit of silica, but their inherently multimode nature remains a major challenge for practical systems requiring high modal purity. In particular, suppressing higher-order modes (HOMs) at the 1 dB/m level while maintaining sub-0.1 dB/km fundamental-mode (FM) loss is difficult because conventional filtering strategies rely on tuning nested-tube dimensions, a design freedom that becomes increasingly restricted in the ultralow-loss regime. Here, we propose a new HOM-control mechanism in an interstitial-tube-assisted double nested anti-resonant nodeless fiber (IT-DNANF) by introducing angular offset of the interstitial tubes. Instead of using nested cavities as the primary tuning element, the proposed approach exploits the gap region between adjacent cladding tubes as a leakage-adjacent modal-control interface. Numerical simulations show that the offset increases both FM and HOM losses, but with a substantially stronger sensitivity for HOMs, leading to rapid enhancement of differential modal loss. Furthermore, when the gap-region FM is tuned into phase matching with the core HOM, strong coupling to a high-leakage state is induced, resulting in a pronounced HOM-loss peak. Using the practical criterion of HOM losslarger than 1 dB/m, we identify optimized IT-DNANF designs that achieve rapid HOM stripping while maintaining FM loss below 0.05 dB/km at 1550 nm. This work establishes angular offset as a physically distinct and manufacturability-friendly degree of freedom for mode purification in ultralow-loss hollow-core fibres.

physics.optics

Cross-Scale Pansharpening via ScaleFormer and the PanScale Benchmark

Pansharpening aims to generate high-resolution multi-spectral images by fusing the spatial detail of panchromatic images with the spectral richness of low-resolution MS data. However, most existing methods are evaluated under limited, low-resolution settings, limiting their generalization to real-world, high-resolution scenarios. To bridge this gap, we systematically investigate the data, algorithmic, and computational challenges of cross-scale pansharpening. We first introduce PanScale, the first large-scale, cross-scale pansharpening dataset, accompanied by PanScale-Bench, a comprehensive benchmark for evaluating generalization across varying resolutions and scales. To realize scale generalization, we propose ScaleFormer, a novel architecture designed for multi-scale pansharpening. ScaleFormer reframes generalization across image resolutions as generalization across sequence lengths: it tokenizes images into patch sequences of the same resolution but variable length proportional to image scale. A Scale-Aware Patchify module enables training for such variations from fixed-size crops. ScaleFormer then decouples intra-patch spatial feature learning from inter-patch sequential dependency modeling, incorporating Rotary Positional Encoding to enhance extrapolation to unseen scales. Extensive experiments show that our approach outperforms SOTA methods in fusion quality and cross-scale generalization. The datasets and source code are available at https://github.com/caoke-963/ScaleFormer.

cs.CV

Geometry and cohomology of compactified Deligne--Lusztig varieties

For connected reductive groups together with a Frobenius root $F$, we show that the cohomology of the structure sheaf and respectively the canonical sheaf for compactified Deligne--Lusztig varieties associated to an element in the free monoid generated by the simple reflections is isomorphic to that of a minimal length element in an $F$-conjugacy class in the Weyl group.

math.AG

Self-assembly of flexible patchy nanoparticles in solution

The self-assembly of polymer grafted nanoparticles is more and more used in the field of functional materials. However, there is still a lack of analysis on the dynamic transformation paths of different self-assembly morphologies, which makes it impossible to achieve further precise regulation and targeted design in experiments and industrial production. In this work the effects of patchy property, grafted chain length, ratio and grafting density on the self-assembly behavior and structure of polymer grafted flexible patchy nanoparticles are investigated by dissipative particle dynamics simulation method through the construction of coarse-grained model of polymer grafted ternary nanoparticles. The influence and regulation mechanisms of these factors on the self-assembly structure transformation of flexible patchy nanoparticles are systematically studied, and a variety of structures such as dendritic structure, columnar structure, and bilayer membrane are obtained. The self-assembly structure of flexible patchy nanoparticles obtained in this work (such as bilayer membrane structure) provides a potential application basis for designing drug carriers. By precisely regulating the specific structural characteristics of the system, it is possible to achieve efficient loading of drugs and targeted delivery functions, thus significantly improving the bioavailability and effect of drugs.

cond-mat.soft

MedAD-R1: Eliciting Consistent Reasoning in Interpretible Medical Anomaly Detection via Consistency-Reinforced Policy Optimization

Medical Anomaly Detection (MedAD) presents a significant opportunity to enhance diagnostic accuracy using Large Multimodal Models (LMMs) to interpret and answer questions based on medical images. However, the reliance on Supervised Fine-Tuning (SFT) on simplistic and fragmented datasets has hindered the development of models capable of plausible reasoning and robust multimodal generalization. To overcome this, we introduce MedAD-38K, the first large-scale, multi-modal, and multi-center benchmark for MedAD featuring diagnostic Chain-of-Thought (CoT) annotations alongside structured Visual Question-Answering (VQA) pairs. On this foundation, we propose a two-stage training framework. The first stage, Cognitive Injection, uses SFT to instill foundational medical knowledge and align the model with a structured think-then-answer paradigm. Given that standard policy optimization can produce reasoning that is disconnected from the final answer, the second stage incorporates Consistency Group Relative Policy Optimization (Con-GRPO). This novel algorithm incorporates a crucial consistency reward to ensure the generated reasoning process is relevant and logically coherent with the final diagnosis. Our proposed model, MedAD-R1, achieves state-of-the-art (SOTA) performance on the MedAD-38K benchmark, outperforming strong baselines by more than 10\%. This superior performance stems from its ability to generate transparent and logically consistent reasoning pathways, offering a promising approach to enhancing the trustworthiness and interpretability of AI for clinical decision support.

cs.CV

High-Polarization-Extinction Raman Conversion in Gas-Filled Polarization-Maintaining Hollow-Core Fibers

Gas-filled hollow-core fibers (HCFs) have emerged as a versatile platform for high-power nonlinear optics, enabling phenomena from ultrafast pulse compression to broadband frequency generation. However, the lack of robust polarization control has remained a critical obstacle to the deployment of gas-based fiber sources. Here, we overcome this bottleneck by demonstrating the generation of highly-polarized Stokes light via stimulated Raman scattering (SRS) in a nitrogen-filled polarization-maintaining anti-resonant hollow-core fiber (PM-HCF). By exploiting the strong structural birefringence of the fiber, the Raman interaction becomes polarization-decoupled along the principal birefringence axes, leading to threshold-selective Raman amplification and an intrinsic polarization purification mechanism. As a result, the vibrational Raman Stokes emission exhibits a polarization extinction ratio (PER) of 35 dB, even when the incident pump PER is as low as ~2 dB. Through analytical theory and numerical modeling, we validate the underlying polarization-selective Raman dynamics and identify the fiber platform as the dominant factor governing the observed PER saturation. We further show that this high polarization purity and high conversion efficiency is maintained under tight bending conditions with radii down to 5 cm, in stark contrast to conventional non-PM-HCF. These results establish PM-HCFs as a robust and scalable architecture for generating polarization-stable, frequency-shifted light, and indicate that polarization may be treated as an actively engineerable degree of freedom in gas photonics, paving the way toward deployment-ready gas-based fiber sources for precision metrology, quantum communication, and coherent sensing.

physics.optics

Polarization-Differential Loss Enabled High Polarization Extinction in Hollow-Core Fibers

Delivering a well defined state of polarization over hollow core fibres (HCFs) is pivotal for next generation ultra stable photonic systems. Yet in all existing HCFs, whether birefringent or not, their polarization extinction ratio (PER) rapidly deteriorates during propagation or under mechanical disturbance, leaving no practical high and stable PER solution. Here, we break this impasse by embedding a polarization differential loss (PDL) mechanism directly into the cladding architecture.

physics.optics

Self-supervised Multiplex Consensus Mamba for General Image Fusion

Image fusion integrates complementary information from different modalities to generate high-quality fused images, thereby enhancing downstream tasks such as object detection and semantic segmentation. Unlike task-specific techniques that primarily focus on consolidating inter-modal information, general image fusion needs to address a wide range of tasks while improving performance without increasing complexity. To achieve this, we propose SMC-Mamba, a Self-supervised Multiplex Consensus Mamba framework for general image fusion. Specifically, the Modality-Agnostic Feature Enhancement (MAFE) module preserves fine details through adaptive gating and enhances global representations via spatial-channel and frequency-rotational scanning. The Multiplex Consensus Cross-modal Mamba (MCCM) module enables dynamic collaboration among experts, reaching a consensus to efficiently integrate complementary information from multiple modalities. The cross-modal scanning within MCCM further strengthens feature interactions across modalities, facilitating seamless integration of critical information from both sources. Additionally, we introduce a Bi-level Self-supervised Contrastive Learning Loss (BSCL), which preserves high-frequency information without increasing computational overhead while simultaneously boosting performance in downstream tasks. Extensive experiments demonstrate that our approach outperforms state-of-the-art (SOTA) image fusion algorithms in tasks such as infrared-visible, medical, multi-focus, and multi-exposure fusion, as well as downstream visual tasks.

cs.CV

MMMamba: A Versatile Cross-Modal In Context Fusion Framework for Pan-Sharpening and Zero-Shot Image Enhancement

Pan-sharpening aims to generate high-resolution multispectral (HRMS) images by integrating a high-resolution panchromatic (PAN) image with its corresponding low-resolution multispectral (MS) image. To achieve effective fusion, it is crucial to fully exploit the complementary information between the two modalities. Traditional CNN-based methods typically rely on channel-wise concatenation with fixed convolutional operators, which limits their adaptability to diverse spatial and spectral variations. While cross-attention mechanisms enable global interactions, they are computationally inefficient and may dilute fine-grained correspondences, making it difficult to capture complex semantic relationships. Recent advances in the Multimodal Diffusion Transformer (MMDiT) architecture have demonstrated impressive success in image generation and editing tasks. Unlike cross-attention, MMDiT employs in-context conditioning to facilitate more direct and efficient cross-modal information exchange. In this paper, we propose MMMamba, a cross-modal in-context fusion framework for pan-sharpening, with the flexibility to support image super-resolution in a zero-shot manner. Built upon the Mamba architecture, our design ensures linear computational complexity while maintaining strong cross-modal interaction capacity. Furthermore, we introduce a novel multimodal interleaved (MI) scanning mechanism that facilitates effective information exchange between the PAN and MS modalities. Extensive experiments demonstrate the superior performance of our method compared to existing state-of-the-art (SOTA) techniques across multiple tasks and benchmarks.

cs.CV

Pan-LUT: Efficient Pan-sharpening via Learnable Look-Up Tables

Recently, deep learning-based pan-sharpening algorithms have achieved notable advancements over traditional methods. However, deep learning-based methods incur substantial computational overhead during inference, especially with large images. This excessive computational demand limits the applicability of these methods in real-world scenarios, particularly in the absence of dedicated computing devices such as GPUs and TPUs. To address these challenges, we propose Pan-LUT, a novel learnable look-up table (LUT) framework for pan-sharpening that strikes a balance between performance and computational efficiency for large remote sensing images. Our method makes it possible to process 15K*15K remote sensing images on a 24GB GPU. To finely control the spectral transformation, we devise the PAN-guided look-up table (PGLUT) for channel-wise spectral mapping. To effectively capture fine-grained spatial details, we introduce the spatial details look-up table (SDLUT). Furthermore, to adaptively aggregate channel information for generating high-resolution multispectral images, we design an adaptive output look-up table (AOLUT). Our model contains fewer than 700K parameters and processes a 9K*9K image in under 1 ms using one RTX 2080 Ti GPU, demonstrating significantly faster performance compared to other methods. Experiments reveal that Pan-LUT efficiently processes large remote sensing images in a lightweight manner, bridging the gap to real-world applications. Furthermore, our model surpasses SOTA methods in full-resolution scenes under real-world conditions, highlighting its effectiveness and efficiency.

cs.CV

Understanding the Representation of Older Adults in Motion Capture Locomotion Datasets

The Internet of Things (IoT) sensors have been widely employed to capture human locomotions to enable applications such as activity recognition, human pose estimation, and fall detection. Motion capture (MoCap) systems are frequently used to generate ground truth annotations for human poses when training models with data from wearable or ambient sensors, and have been shown to be effective to synthesize data in these modalities. However, the representation of older adults, an increasingly important demographic in healthcare, in existing MoCap locomotion datasets has not been thoroughly examined. This work surveyed 41 publicly available datasets, identifying eight that include older adult motions and four that contain motions performed by younger actors annotated as old style. Older adults represent a small portion of participants overall, and few datasets provide full-body motion data for this group. To assess the fidelity of old-style walking motions, quantitative metrics are introduced, defining high fidelity as the ability to capture age-related differences relative to normative walking. Using gait parameters that are age-sensitive, robust to noise, and resilient to data scarcity, we found that old-style walking motions often exhibit overly controlled patterns and fail to faithfully characterize aging. These findings highlight the need for improved representation of older adults in motion datasets and establish a method to quantitatively evaluate the quality of old-style walking motions.

cs.CY

CT Radiomics-Based Explainable Machine Learning Model for Accurate Differentiation of Malignant and Benign Endometrial Tumors: A Two-Center Study

Aimed to develop and validate a CT radiomics-based explainable machine learning model for precise diagnosing malignancy and benignity specifically in endometrial cancer (EC) patients. A total of 83 EC patients from two centers, including 46 with malignant and 37 with benign conditions, were included, with data split into a training set (n=59) and a testing set (n=24). The regions of interest (ROIs) were manually segmented from pre-surgical CT scans, and 1132 radiomic features were extracted from the pre-surgical CT scans using Pyradiomics. Six explainable machine learning (ML) modeling algorithms were implemented respectively, for determining the optimal radiomics pipeline. The diagnostic performance of the radiomic model was evaluated by using sensitivity, specificity, accuracy, precision, F1 score, AUROC, and AUPRC. To enhance clinical understanding and usability, we separately implemented SHAP analysis and feature mapping visualization, and evaluated the calibration curve and decision curve. By comparing six modeling strategies, the Random Forest model emerged as the optimal choice for diagnosing EC, with a training AUROC of 1.00 and a testing AUROC of 0.96. SHAP identified the most important radiomic features, revealing that all selected features were significantly associated with EC (P < 0.05). Radiomics feature maps also provide a feasible assessment tool for clinical applications. Decision Curve Analysis (DCA) indicated a higher net benefit for our model compared to the "All" and "None" strategies, suggesting its clinical utility in identifying high-risk cases and reducing unnecessary interventions. In conclusion, the CT radiomics-based explainable ML model achieved high diagnostic performance, which could be used as an intelligent auxiliary tool for the diagnosis of endometrial cancer.

eess.IV