SearcharxivSearch

arXiv subjects

Dongdong Wang

Publications and source records attributed to Dongdong Wang.

At least 19 recordsLinked to original sources

Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection

Online log anomaly detection is critical for maintaining the reliability of large-scale computing systems. Although recent language model-based log anomaly detectors achieve strong detection performance, their confidence estimates remain poorly calibrated. We show that these detectors frequently assign excessive confidence to incorrect predictions, particularly for anomalous logs under severe class imbalance. Moreover, confidence on erroneous predictions remains persistently high even when conventional calibration metrics indicate good calibration, creating a critical reliability gap for operational monitoring systems. To address this issue, we propose Log Reconstruction and Distance (LoRD), a lightweight post-hoc calibration framework for reliable log anomaly detection. LoRD learns prediction-route-specific reliability models from latent representations of correctly classified validation samples and estimates prediction reliability through route-wise reconstruction distances. Based on the estimated reliability, LoRD selectively recalibrates high-risk predictions to suppress overconfident errors while preserving reliable predictions. Extensive experiments on four large-scale log benchmark datasets and multiple language model-based detectors demonstrate that LoRD consistently improves confidence reliability and substantially reduces overconfident anomaly-related errors without sacrificing anomaly detection performance.

cs.LG

GoldenRetriever: Non-Interactive Homomorphic Encrypted Retrieval for Privacy-Preserving RAG

Retrieval-Augmented Generation (RAG) enhances large language models by incorporating external knowledge, but existing pipelines typically operate on plaintext data, raising significant privacy concerns. Prior work on privacy-preserving retrieval leverages cryptographic techniques such as homomorphic encryption (HE) and private information retrieval (PIR), but often relies on interactive protocols or ranking-based selection mechanisms that incur high latency and potential information leakage. In this paper, we propose a practical non-interactive encrypted retrieval framework for RAG based on threshold selection. Instead of performing expensive top-$k$ ranking under encryption, our approach selects documents whose similarity scores exceed a predefined threshold, reducing computational complexity from quadratic to linear in the corpus size. We implement this design using CKKS-based homomorphic computation, enabling fully encrypted similarity evaluation and document selection without revealing query content, intermediate scores, or selected indices. To bridge the gap between approximate encrypted computation and discrete token reconstruction, we introduce a precision-stable mask polarization method that ensures accurate recovery of selected documents. Experiments on standard retrieval benchmarks demonstrate that our approach achieves competitive retrieval effectiveness while significantly reducing latency compared to ranking-based encrypted methods. These results highlight threshold-based selection as a practical foundation for scalable and secure RAG systems.

cs.CR

Quiescent and traveling solitons in the fractional parametrically driven damped nonlinear Schrödinger equation

We systematically investigate the existence, stability, and dynamics of optical solitons in the framework of the one-dimensional nonlinear Schrödinger equation with the Riesz-fractional diffraction operator, cubic self-focusing, and linear loss, balanced by a linear parametric drive. The model, which can be realized in a laser cavity, produces standing and moving solitons, the latter ones existing below a critical velocity. One of the soliton species is stable in a wide range of parameters, while others are unstable. The fractional diffraction significantly alters the existence conditions and stability thresholds of the solitons. Collision between moving solitons are considered too. The results essentially expand the variety of nonlinear modes in media with fractional diffraction.

nlin.PS

Do VLMs See What Sensors Feel? A Scalable Expert-Guided Design for Wheelchair Accessibility Assessment from Street View

Assessing built-environment interaction, such as wheelchair accessibility, is difficult because real-world mobility is shaped by distributed, context-dependent, and temporary barriers that are hard to capture at scale. To support scalable assessment, this paper examines whether vision-language models (VLMs) can identify accessibility barriers from Google Street View (GSV) imagery. We propose an expert-guided retrieval-augmented framework that combines GSV images, ADA-informed guidance, and expert-derived rubrics to evaluate accessibility dimensions. We collect a campus-scale dataset at the University of Florida, linking 407 unique GSV locations with GPS-derived wheelchair dwell behavior as a mobility-friction signal. Results show that VLM ratings are both negatively correlated and distributionally similar with dwell time, indicating partial but consistent alignment with a behavioral proxy for mobility friction. Visual cue analysis shows that certain environmental objects, such as curb ramps and crosswalks, are associated with higher VLM accessibility scores, while alignment remains limited for subtle surface conditions, transient obstructions, and viewpoint-dependent barriers. Overall, our findings show the potential of expert-guided VLMs for scalable accessibility assessment aligning with sensor-derived indicators of real-world wheelchair navigation.

cs.CV

Parametrically driven pure-quartic solitons

Parametrically driven solitons are self-trapped modes in various physical settings, including optics, magnetics, etc. So far, the analysis was focused on the existence, stability, and dynamics of such solitons in systems including the second-order group-velocity dispersion (GVD), linear loss, parametric gain, and cubic nonlinearity. Here, we report the existence of quiescent parametrically driven pure-quartic solitons (PDPQSs) in the full system, and moving PDPQSs in the absence of losses. A systematic analysis reveals stability domains for the solitons in the system's parameter space. Evolution of unstable states is explored too, and it is demonstrated that collisions between traveling stable PDPQSs are elastic.

physics.optics

Physics-Informed Teacher-Student Ensemble Learning for Traffic State Estimation with a Varying Speed Limit Scenario

Physics-informed deep learning (PIDL) neural networks have shown their capability as a useful instrument for transportation practitioners in utilizing the underlying relationship between the state variables for traffic state estimation (TSE). Another efficient traffic management approach is implementing varying speed limits (VSLs) on transportation corridors to control traffic and mitigate congestion. However, the existing training architecture of PIDL in the literature cannot accommodate the changing traffic characteristics on a freeway with VSL. To tackle this challenge, we propose a novel framework integrating teacher-student ensemble training with PIDL neural networks for TSE under VSL scenarios. The physics of flow conservation law is encoded locally in the teacher models by PIDL, and the student model uses a multi-layer perceptron classifier (MLP) to identify traffic characteristics and selects the ensemble member of PIDL neural networks for TSE. This integrated framework provides a natural solution for capturing the heterogeneity of VSL and accurately addressing the TSE problem. The case study results validate the proposed ensemble approach, demonstrating its superior performance in TSE compared to other popular baseline methods, as indicated by relative L2 error.

cs.LG

Cloud-top infrared observations reveal the four-dimensional precipitation structure

Accurate four-dimensional (4D) precipitation information is essential for understanding the Earth's energy and water cycles, yet remains observationally unresolved at global scales. Conventional theory holds that geostationary infrared observations primarily sense cloud-top properties, with limited sensitivity to sub-cloud precipitation. Here we show that cloud-top infrared measurements nevertheless encode sufficient information to recover the four-dimensional structure of precipitation, revealing a previously unexploited observability of sub-cloud processes. We introduce a physically constrained deep learning framework, 4DPrecipNet, in which a moisture-first constraint requires the latent representation to recover precipitable water vapour, anchoring the model in thermodynamic consistency. By integrating multi-channel infrared radiances with these constraints and radar-derived precipitation profiles, we reconstruct the vertical and temporal evolution of precipitation systems from geostationary orbit. The framework captures deep convective structures and their evolution, with robust performance across large samples and independent radar comparisons. These results demonstrate that sub-cloud precipitation is physically encoded in cloud-top infrared observations, establishing a new pathway for continuous global monitoring of precipitation structure.

cs.CV

Built Environment Reasoning from Remote Sensing Imagery Using Large Vision--Language Models

This work investigates the use of large language models (LLMs) for tasks in smart cities. The core idea is to leverage remote sensing imagery to characterize the built environment, including design suggestions, constructability assessment, landuse patterns, and risk identification. We examine remote sensing imagery at multiple spatial scales as inputs for multimodal language modeling and evaluate their effects on built-environment-related reasoning. In addition, we compare state-of-the-art LLMs, including InternVL and Qwen, in terms of accuracy and reliability when generating built environment recommendations. The results demonstrate the potential of integrating remote sensing imagery with large language models to assist smart cities and decision-making.

cs.CL

Impacts of Climate Change on Photovoltaic Potential in Africa

Africa holds the world's highest solar irradiance yet has <2% of global photovoltaic (PV) capacity, leaving 600 million people without electricity access. However, climate change impacts on its 10 TW potential remain understudied. Using four decades of ERA5 reanalysis data (1980-2020) at 0.25 degree resolution, we quantify the contributions of key climate factors to historical changes in African PV potential through multivariate decomposition. Continental PV potential increased by 3.2%, driven primarily by enhanced solar radiation (+1.2 degree Celsius, contributing -23%). East Africa gained >6% from radiation enhancement, while North Africa declined by 0.5% as extreme heat (+2 degree Celsius) overwhelmed radiation benefits. Critically, stability analysis using the coefficient of variation (CV) reveals that high-irradiance subtropical zones are highly variable (CV=0.4), in contrast to stable equatorial regions (CV=0.1), challenging the assumption that resource abundance ensures reliability. These findings reframe Africa's solar strategy: North Africa requires prioritizing heat-resilient technology over capacity maximization; subtropical zones demand grid-storage co-investment; and East Africa presents globally competitive opportunities for rapid, stable deployment. By resolving spatiotemporal heterogeneities and quantifying climate-driver contributions, our analysis provides an actionable framework for climate-resilient solar deployment, critical for Africa's energy transition and climate mitigation.

physics.soc-ph

Transformation of topological optical states via spiral modulation in fractional-diffraction systems

We propose a scheme for manipulations of a variety of topological states in fractional optical systems through spiral modulation of the local refraction index. An analytical approximation, based on a truncated finite-mode system, and direct simulations reveal that the spiral modulation supports direct mutual conversion between eigenmodes with topological-charge difference Δm=1, driven by the resonant coupling between the modes and the underlying spiral modulation. The conversion between eigenmodes with Δm=2 requires involvement of an intermediary mode and precise tuning of the resonance condition. We further explore the conversion of degenerate modes under the action of azimuthal modulation. Modulated degenerate modes exhibit incomplete conversion, evolving into intermediate states with odd parity symmetry. Finally, we examine nonlinear effects on the spiral-modulation-induced mode conversion, identifying an essential nonlinearity-induced resonance-frequency shift that critically affects the conversion efficiency.

physics.optics

A novel fast sweeping method for computing the attenuation operator $t^*$ in absorbing media

$t^*$ represents the total path attenuation and characterizes the amplitude decay of a propagating seismic wave. Calculating the attenuation operator $t^*$ is typically required in seismic attenuation tomography. Traditional methods for calculating $t^*$ require determining the ray path explicitly. However, ray tracing can be computationally intensive when processing large datasets, and conventional ray tracing techniques may fail even in mildly heterogeneous media. In this study, we propose a modified fast sweeping method (MFSM) to solve the governing equation for $t^*$ without explicitly calculating the ray path. The approach consists of two main steps. First, the traveltime field is calculated by numerically solving the eikonal equation using the fast sweeping method. Second, $t^*$ is computed by solving its governing equation with the MFSM, based on the discretization of the gradient of $t^*$ using an upwinding scheme derived from the traveltime gradient. The MFSM is rigorously validated through comparisons with analytical solutions and by examining $t^*$ errors under grid refinement in both simple and complex models. Key performance metrics, including convergence, number of iterations, and computation time, are evaluated. Two versions of the MFSM are developed for both Cartesian and spherical coordinate systems. We demonstrate the practical applicability of the developed MFSM in calculating $t^*$ in North Island, and discuss the method's efficiency in estimating earthquake response spectra.

physics.geo-ph

Multimodal Human-AI Synergy for Medical Imaging Quality Control: A Hybrid Intelligence Framework with Adaptive Dataset Curation and Closed-Loop Evaluation

Medical imaging quality control (QC) is essential for accurate diagnosis, yet traditional QC methods remain labor-intensive and subjective. To address this challenge, in this study, we establish a standardized dataset and evaluation framework for medical imaging QC, systematically assessing large language models (LLMs) in image quality assessment and report standardization. Specifically, we first constructed and anonymized a dataset of 161 chest X-ray (CXR) radiographs and 219 CT reports for evaluation. Then, multiple LLMs, including Gemini 2.0-Flash, GPT-4o, and DeepSeek-R1, were evaluated based on recall, precision, and F1 score to detect technical errors and inconsistencies. Experimental results show that Gemini 2.0-Flash achieved a Macro F1 score of 90 in CXR tasks, demonstrating strong generalization but limited fine-grained performance. DeepSeek-R1 excelled in CT report auditing with a 62.23\% recall rate, outperforming other models. However, its distilled variants performed poorly, while InternLM2.5-7B-chat exhibited the highest additional discovery rate, indicating broader but less precise error detection. These findings highlight the potential of LLMs in medical imaging QC, with DeepSeek-R1 and Gemini 2.0-Flash demonstrating superior performance.

cs.CL

Exploratory analysis of injury severity under different levels of driving automation (SAE Level 2-5) using multi-source data

Vehicles equipped with automated driving capabilities have shown potential to improve safety and operations. Advanced driver assistance systems (ADAS) and automated driving systems (ADS) have been widely developed to support vehicular automation. Although the studies on the injury severity outcomes that involve automated vehicles are ongoing, there is limited research investigating the difference between injury severity outcomes for the ADAS and ADS equipped vehicles. To ensure a comprehensive analysis, a multi-source dataset that includes 1,001 ADAS crashes (SAE Level 2 vehicles) and 548 ADS crashes (SAE Level 4 vehicles) is used. Two random parameters multinomial logit models with heterogeneity in the means of random parameters are considered to gain a better understanding of the variables impacting the crash injury severity outcomes for the ADAS (SAE Level 2) and ADS (SAE Level 4) vehicles. It was found that while 67 percent of crashes involving the ADAS equipped vehicles in the dataset took place on a highway, 94 percent of crashes involving ADS took place in more urban settings. The model estimation results also reveal that the weather indicator, driver type indicator, differences in the system sophistication that are captured by both manufacture year and high/low mileage as well as rear and front contact indicators all play a role in the crash injury severity outcomes. The results offer an exploratory assessment of safety performance of the ADAS and ADS equipped vehicles using the real-world data and can be used by the manufacturers and other stakeholders to dictate the direction of their deployment and usage.

stat.AP

RiD-kit: Software package designed to do enhanced sampling using reinforced dynamics

Developing an efficient method to accelerate the speed of molecular dynamics is a central theme in the field of molecular simulation. One category among the methods are collective-variable-based methods, which rely on predefined collective variables (CVs). The difficulty of selecting a few important CVs hinders the methods to be applied to large systems easily. Here we present a CV-based enhanced sampling method RiD-kit, which could handle a large number of CVs and perform efficient sampling. The method could be applied to various kinds of systems, including biomolecules, chemical reactions and materials. In this protocol, we guide the users through all phases of the RiD-kit workflow, from preparing the input files, setting the simulation parameters and analyzing the results. The RiD-kit workflow provides an efficient and user-friendly command line tool which could submit jobs to various kinds of platforms including the high-performance computers (HPC), cloud server and local machines.

physics.chem-ph

Video-to-Text Pedestrian Monitoring (VTPM): Leveraging Computer Vision and Large Language Models for Privacy-Preserve Pedestrian Activity Monitoring at Intersections

Computer vision has advanced research methodologies, enhancing system services across various fields. It is a core component in traffic monitoring systems for improving road safety; however, these monitoring systems don't preserve the privacy of pedestrians who appear in the videos, potentially revealing their identities. Addressing this issue, our paper introduces Video-to-Text Pedestrian Monitoring (VTPM), which monitors pedestrian movements at intersections and generates real-time textual reports, including traffic signal and weather information. VTPM uses computer vision models for pedestrian detection and tracking, achieving a latency of 0.05 seconds per video frame. Additionally, it detects crossing violations with 90.2% accuracy by incorporating traffic signal data. The proposed framework is equipped with Phi-3 mini-4k to generate real-time textual reports of pedestrian activity while stating safety concerns like crossing violations, conflicts, and the impact of weather on their behavior with latency of 0.33 seconds. To enhance comprehensive analysis of the generated textual reports, Phi-3 medium is fine-tuned for historical analysis of these generated textual reports. This fine-tuning enables more reliable analysis about the pedestrian safety at intersections, effectively detecting patterns and safety critical events. The proposed VTPM offers a more efficient alternative to video footage by using textual reports reducing memory usage, saving up to 253 million percent, eliminating privacy issues, and enabling comprehensive interactive historical analysis.

cs.CV

Perturbing Attention Gives You More Bang for the Buck: Subtle Imaging Perturbations That Efficiently Fool Customized Diffusion Models

Diffusion models (DMs) embark a new era of generative modeling and offer more opportunities for efficient generating high-quality and realistic data samples. However, their widespread use has also brought forth new challenges in model security, which motivates the creation of more effective adversarial attackers on DMs to understand its vulnerability. We propose CAAT, a simple but generic and efficient approach that does not require costly training to effectively fool latent diffusion models (LDMs). The approach is based on the observation that cross-attention layers exhibits higher sensitivity to gradient change, allowing for leveraging subtle perturbations on published images to significantly corrupt the generated images. We show that a subtle perturbation on an image can significantly impact the cross-attention layers, thus changing the mapping between text and image during the fine-tuning of customized diffusion models. Extensive experiments demonstrate that CAAT is compatible with diverse diffusion models and outperforms baseline attack methods in a more effective (more noise) and efficient (twice as fast as Anti-DreamBooth and Mist) manner.

cs.CV

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

We present Hunyuan-DiT, a text-to-image diffusion transformer with fine-grained understanding of both English and Chinese. To construct Hunyuan-DiT, we carefully design the transformer structure, text encoder, and positional encoding. We also build from scratch a whole data pipeline to update and evaluate data for iterative model optimization. For fine-grained language understanding, we train a Multimodal Large Language Model to refine the captions of the images. Finally, Hunyuan-DiT can perform multi-turn multimodal dialogue with users, generating and refining images according to the context. Through our holistic human evaluation protocol with more than 50 professional human evaluators, Hunyuan-DiT sets a new state-of-the-art in Chinese-to-image generation compared with other open-source models. Code and pretrained models are publicly available at github.com/Tencent/HunyuanDiT

cs.CV

Enhancing Traffic Safety with Parallel Dense Video Captioning for End-to-End Event Analysis

This paper introduces our solution for Track 2 in AI City Challenge 2024. The task aims to solve traffic safety description and analysis with the dataset of Woven Traffic Safety (WTS), a real-world Pedestrian-Centric Traffic Video Dataset for Fine-grained Spatial-Temporal Understanding. Our solution mainly focuses on the following points: 1) To solve dense video captioning, we leverage the framework of dense video captioning with parallel decoding (PDVC) to model visual-language sequences and generate dense caption by chapters for video. 2) Our work leverages CLIP to extract visual features to more efficiently perform cross-modality training between visual and textual representations. 3) We conduct domain-specific model adaptation to mitigate domain shift problem that poses recognition challenge in video understanding. 4) Moreover, we leverage BDD-5K captioned videos to conduct knowledge transfer for better understanding WTS videos and more accurate captioning. Our solution has yielded on the test set, achieving 6th place in the competition. The open source code will be available at https://github.com/UCF-SST-Lab/AICity2024CVPRW

cs.CV