SearcharxivSearch

arXiv subjects

Yanhong Wu

Publications and source records attributed to Yanhong Wu.

At least 19 recordsLinked to original sources

FreoStream:Enhancing Stream Guardrails via Future-Aware Reasoning and Safety-Aligned Optimization

Stream guardrails enable token-level safety detection before full responses are generated. However, they often make overly conservative judgements and block those sensitive but safe tokens, which is known as over-refusal. Due to lack of full context, they also fail to detect implicitly harmful content from jailbreaking. To address these challenges, we propose FreoStream, a novel streaming guardrail framework. Specifically, FreoStream fine-tunes a LoRA module to perform Future-Aware Reasoning when the base guardrail detects unsafe tokens. The reasoning process follows a Future-Reason-Judge paradigm: predict the future, reason about the full context and give the final judgement. This design can effectively reduce over-refusal by incorporating the future information. Moreover, we introduce the Safety-Aligned Optimization module that extracts the safety-aligned component from the reasoning gradients to update the base guardrail model, thereby enhancing streaming safety detection. Extensive experiments on various safety benchmarks demonstrate that FreoStream achieves lower over-refusal rates and better jailbreak defense compared to existing streaming guardrails.

cs.CR

RECITYGEN -- Interactive and Generative Participatory Urban Design Tool with Latent Diffusion and Segment Anything

Urban design profoundly impacts public spaces and community engagement. Traditional top-down methods often overlook public input, creating a gap in design aspirations and reality. Recent advancements in digital tools, like City Information Modelling and augmented reality, have enabled a more participatory process involving more stakeholders in urban design. Further, deep learning and latent diffusion models have lowered barriers for design generation, providing even more opportunities for participatory urban design. Combining state-of-the-art latent diffusion models with interactive semantic segmentation, we propose RECITYGEN, a novel tool that allows users to interactively create variational street view images of urban environments using text prompts. In a pilot project in Beijing, users employed RECITYGEN to suggest improvements for an ongoing Urban Regeneration project. Despite some limitations, RECITYGEN has shown significant potential in aligning with public preferences, indicating a shift towards more dynamic and inclusive urban planning methods. The source code for the project can be found at RECITYGEN GitHub.

cs.CV

ChemVTS-Bench: Evaluating Visual-Textual-Symbolic Reasoning of Multimodal Large Language Models in Chemistry

Chemical reasoning inherently integrates visual, textual, and symbolic modalities, yet existing benchmarks rarely capture this complexity, often relying on simple image-text pairs with limited chemical semantics. As a result, the actual ability of Multimodal Large Language Models (MLLMs) to process and integrate chemically meaningful information across modalities remains unclear. We introduce \textbf{ChemVTS-Bench}, a domain-authentic benchmark designed to systematically evaluate the Visual-Textual-Symbolic (VTS) reasoning abilities of MLLMs. ChemVTS-Bench contains diverse and challenging chemical problems spanning organic molecules, inorganic materials, and 3D crystal structures, with each task presented in three complementary input modes: (1) visual-only, (2) visual-text hybrid, and (3) SMILES-based symbolic input. This design enables fine-grained analysis of modality-dependent reasoning behaviors and cross-modal integration. To ensure rigorous and reproducible evaluation, we further develop an automated agent-based workflow that standardizes inference, verifies answers, and diagnoses failure modes. Extensive experiments on state-of-the-art MLLMs reveal that visual-only inputs remain challenging, structural chemistry is the hardest domain, and multimodal fusion mitigates but does not eliminate visual, knowledge-based, or logical errors, highlighting ChemVTS-Bench as a rigorous, domain-faithful testbed for advancing multimodal chemical reasoning. All data and code will be released to support future research.

cs.AI

PRevivor: Reviving Ancient Chinese Paintings using Prior-Guided Color Transformers

Ancient Chinese paintings are a valuable cultural heritage that is damaged by irreversible color degradation. Reviving color-degraded paintings is extraordinarily difficult due to the complex chemistry mechanism. Progress is further slowed by the lack of comprehensive, high-quality datasets, which hampers the creation of end-to-end digital restoration tools. To revive colors, we propose PRevivor, a prior-guided color transformer that learns from recent paintings (e.g., Ming and Qing Dynasty) to restore ancient ones (e.g., Tang and Song Dynasty). To develop PRevivor, we decompose color restoration into two sequential sub-tasks: luminance enhancement and hue correction. For luminance enhancement, we employ two variational U-Nets and a multi-scale mapping module to translate faded luminance into restored counterparts. For hue correction, we design a dual-branch color query module guided by localized hue priors extracted from faded paintings. Specifically, one branch focuses attention on regions guided by masked priors, enforcing localized hue correction, whereas the other branch remains unconstrained to maintain a global reasoning capability. To evaluate PRevivor, we conduct extensive experiments against state-of-the-art colorization methods. The results demonstrate superior performance both quantitatively and qualitatively.

cs.CV

Stability analysis of discrete Boltzmann simulation for supersonic flows: Influencing factors, coupling mechanisms and optimization strategies

Supersonic flow simulations face challenges in trans-scale modeling, numerical stability, and complex field analysis due to inherent nonlinear, nonequilibrium, and multiscale characteristics. The discrete Boltzmann method (DBM) provides a multiscale kinetic modeling framework and analysis tool to capture complex discrete/nonequilibrium effects. While the numerical scheme plays a fundamental role in DBM simulations, a comprehensive stability analysis remains lacking. Similar to LBM, complexity mainly lies in the intrinsic coupling between velocity and spatiotemporal discretizations, compared with CFD. This study conducts von Neumann stability analysis to investigate key factors influencing DBM simulation stability, including phase-space discretization, thermodynamic nonequilibrium (TNE) levels, spatiotemporal schemes, initial conditions, and model parameters. Key findings include: (i) the moment-matching approach outperforms the expansion- and weighting-based methods in the test simulations; (ii) increased TNE enhances system nonlinearity and the intrinsic nonlinearity embedded in the model equations, amplifying instabilities; (iii) additional viscous dissipation based on distribution functions improves stability but distorts flow fields and alters constitutive relations; (iv) larger CFL numbers and relative time steps degrade stability, necessitating appropriate time-stepping strategies. To assess the stability regulation capability of DBMs across TNE levels, stability-phase diagrams and probability curves are constructed via morphological analysis within the moment-matching framework. These diagrams identify common stable parameter regions across model orders. This study reveals key factors and coupling mechanisms affecting DBM stability and proposes strategies for optimizing equilibrium distribution discretization, velocity design, and parameter selection in supersonic regimes.

physics.flu-dyn

Thermodynamic nonequilibrium effects in three-dimensional high-speed compressible flows: Multiscale modeling and simulation via the discrete Boltzmann method

Three-dimensional (3D) high-speed compressible flow is a typical nonlinear, nonequilibrium, and multiscale complex flow. Traditional fluid mechanics models, based on the quasi-continuum assumption and near-equilibrium approximation, are insufficient to capture significant discrete effects and thermodynamic nonequilibrium effects (TNEs) as the Knudsen number increases. To overcome these limitations, a discrete Boltzmann modeling and simulation method, rooted in kinetic and mean-field theories, has been developed. By applying Chapman-Enskog multiscale analysis, the essential kinetic moment relations $\bmΦ$ for characterizing second-order TNEs are determined. These relations are invariants in coarse-grained physical modeling, providing a unique mesoscopic perspective for analyzing TNE behaviors. A discrete Boltzmann model, accurate to the second-order in the Knudsen number, is developed to enable multiscale simulations of 3D supersonic flows. As key TNE measures, nonlinear constitutive relations (NCRs), are theoretically derived for the 3D case, offering a constitutive foundation for improving macroscopic fluid modeling. The NCRs in three dimensions exhibit greater complexity than their two-dimensional counterparts. This complexity arises from increased degrees of freedom, which introduce additional kinds of nonequilibrium driving forces, stronger coupling between these forces, and a significant increase in nonequilibrium components. At the macroscopic level, the model is validated through several classical test cases, ranging from 1D to 3D scenarios, from subsonic to supersonic regimes. At the mesoscopic level, the model accurately captures typical TNEs, such as viscous stress and heat flux, around mesoscale structures, across various scales and orders. This work provides kinetic insights that advance multiscale simulation techniques for 3D high-speed compressible flows.

physics.flu-dyn

Supersonic flow kinetics: Mesoscale structures, thermodynamic nonequilibrium effects and entropy production mechanisms

Supersonic flow is a typical nonlinear, nonequilibrium, multiscale, and complex phenomenon. This paper applies discrete Boltzmann method/model (DBM) to simulate and analyze these characteristics. A Burnett-level DBM for supersonic flow is constructed based on the Shakhov-BGK model. Higher-order analytical expressions for thermodynamic nonequilibrium effects are derived, providing a constitutive basis for improving traditional macroscopic hydrodynamics modeling. Criteria for evaluating the validity of DBM are established by comparing numerical and analytical solutions of nonequilibrium measures. The multiscale DBM is used to investigate discrete/nonequilibrium characteristics and entropy production mechanisms in shock regular reflection. The findings include: (a) Compared to NS-level DBM, the Burnett-level DBM offers more accurate representations of viscous stress and heat flux, ensures non-negativity of entropy production in accordance with the second law of thermodynamics, and exhibits better numerical stability. (b) Near the interfaces of incident and reflected shock waves, strong nonequilibrium driving forces lead to prominent nonequilibrium effects. By monitoring the timing and location of peak nonequilibrium quantities, the evolution characteristics of incident and reflected shock waves can be accurately and dynamically tracked. (c) In the intermediate state, the bent reflected shock and incident shock interface are wider and exhibit lower nonequilibrium intensities compared to their final state. (d) The Mach number enhances various kinds of nonequilibrium intensities in a power-law manner $D_{mn} \sim \mathtt{Ma}^α$. The power exponent $α$ and kinetic modes of nonequilibrium effects $m$ follows a logarithmic relation $α\sim \ln (m - m_0)$. This research provides new perspectives and kinetic insights into supersonic flow studies.

physics.flu-dyn

RHDLPP: A multigroup radiation hydrodynamics code for laser-produced plasmas

We introduce the RHDLPP, a flux-limited multigroup radiation hydrodynamics numerical code designed for simulating laser-produced plasmas in diverse environments. The code bifurcates into two packages: RHDLPP-LTP for low-temperature plasmas generated by moderate-intensity nanosecond lasers, and RHDLPP-HTP for high-temperature, high-density plasmas formed by high-intensity laser pulses. The core radiation hydrodynamic equations are resolved in the Eulerian frame, employing an operator-split method. This method decomposes the solution into two substeps: first, the explicit resolution of the hyperbolic subsystems integrating radiation and fluid dynamics, and second, the implicit treatment of the parabolic part comprising stiff radiation diffusion, heat conduction, and energy exchange. Laser propagation and energy deposition are modeled through a hybrid approach, combining geometrical optics ray-tracing in sub-critical plasma regions with a one-dimensional solution of the Helmholtz wave equation in super-critical areas. The thermodynamic states are ascertained using an equation of state, based on either the real gas approximation or the quotidian equation of state (QEOS). Additionally, RHDLPP includes RHDLPP-SpeIma3D, a three-dimensional spectral simulation post-processing module, for generating both temporally-spatially resolved and time-integrated spectra and imaging, facilitating direct comparisons with experimental data. The paper showcases a series of verification tests to establish the code's accuracy and efficiency, followed by application cases, including simulations of laser-produced aluminum (Al) plasmas, pre-pulse-induced target deformation of tin (Sn) microdroplets relevant to extreme ultraviolet lithography light sources, and varied imaging and spectroscopic simulations.

physics.plasm-ph

DyFormer: A Scalable Dynamic Graph Transformer with Provable Benefits on Generalization Ability

Transformers have achieved great success in several domains, including Natural Language Processing and Computer Vision. However, its application to real-world graphs is less explored, mainly due to its high computation cost and its poor generalizability caused by the lack of enough training data in the graph domain. To fill in this gap, we propose a scalable Transformer-like dynamic graph learning method named Dynamic Graph Transformer (DyFormer) with spatial-temporal encoding to effectively learn graph topology and capture implicit links. To achieve efficient and scalable training, we propose temporal-union graph structure and its associated subgraph-based node sampling strategy. To improve the generalization ability, we introduce two complementary self-supervised pre-training tasks and show that jointly optimizing the two pre-training tasks results in a smaller Bayesian error rate via an information-theoretic analysis. Extensive experiments on the real-world datasets illustrate that DyFormer achieves a consistent 1%-3% AUC gain (averaged over all time steps) compared with baselines on all benchmarks.

cs.LG

Sequential Detection of Transient Signals with Exponential Family Distribution

We first consider the sequential detection of transient signals by generalizing the moving average chart to exponential family and study the false detection probability (FDP) and power of detection (POD) in the steady state. Then windowed adjusted signed (or modified directed) likelihood ratio chart is studied by treating it as normal random variable. In the multi-parameter exponential family, the detection of the transient change of one of the canonical parameters or a function of canonical parameters is considered by using the generalized adjusted signed likelihood ratio chart. Comparisons with window restricted CUSUM and Shiryayev-Roberts (S-R) procedures show that the generalized signed likelihood ratio chart performs quite well. Several important examples including the mean or variance change under normal model and a real example are used for illustration.

math.ST

Sequential Detection of Common Change in High-dimensional Data Stream

After obtaining an accurate approximation for $ARL_0$, we first consider the optimal design of weight parameter for a multivariate EWMA chart that minimizes the stationary average delay detection time (SADDT). Comparisons with moving average (MA), CUSUM, generalized likelihood ratio test (GLRT), and Shiryayev-Roberts (S-R) charts after obtaining their $ARL_0$ and SADDT's are conducted numerically. To detect the change with sparse signals, hard-threshold and soft-threshold EWMA charts are proposed. Comparisons with other charts including adaptive techniques show that the EWMA procedure should be recommended for its robust performance and easy design.

math.ST

Sequential Detection of Transient Signals in High Dimensional Data Stream

Motivated by sequential detection of transient signals in high dimensional data stream, we study the performance of EWMA, MA, CUSUM, and GLRT charts for detecting a transient signal in multivariate data streams in terms of the power of detection (POD) under the constraint of false detecting probability (FDP) at the stationary state. Approximations are given for FDP and POD. Comparisons show that the EWMA chart performs equally well as the GLRT chart when the signal strength is unknown, while its design is free of signal length and easy to update. In addition, the MEWMA chart with hard-threshold performs better when the signal only appears in a small portion of the channels. Dow Jones 30 industrial stock prices are used for illustration.

math.ST

VideoModerator: A Risk-aware Framework for Multimodal Video Moderation in E-Commerce

Video moderation, which refers to remove deviant or explicit content from e-commerce livestreams, has become prevalent owing to social and engaging features. However, this task is tedious and time consuming due to the difficulties associated with watching and reviewing multimodal video content, including video frames and audio clips. To ensure effective video moderation, we propose VideoModerator, a risk-aware framework that seamlessly integrates human knowledge with machine insights. This framework incorporates a set of advanced machine learning models to extract the risk-aware features from multimodal video content and discover potentially deviant videos. Moreover, this framework introduces an interactive visualization interface with three views, namely, a video view, a frame view, and an audio view. In the video view, we adopt a segmented timeline and highlight high-risk periods that may contain deviant information. In the frame view, we present a novel visual summarization method that combines risk-aware features and video context to enable quick video navigation. In the audio view, we employ a storyline-based design to provide a multi-faceted overview which can be used to explore audio content. Furthermore, we report the usage of VideoModerator through a case scenario and conduct experiments and a controlled user study to validate its effectiveness.

cs.HC

Beating Attackers At Their Own Games: Adversarial Example Detection Using Adversarial Gradient Directions

Adversarial examples are input examples that are specifically crafted to deceive machine learning classifiers. State-of-the-art adversarial example detection methods characterize an input example as adversarial either by quantifying the magnitude of feature variations under multiple perturbations or by measuring its distance from estimated benign example distribution. Instead of using such metrics, the proposed method is based on the observation that the directions of adversarial gradients when crafting (new) adversarial examples play a key role in characterizing the adversarial space. Compared to detection methods that use multiple perturbations, the proposed method is efficient as it only applies a single random perturbation on the input example. Experiments conducted on two different databases, CIFAR-10 and ImageNet, show that the proposed detection method achieves, respectively, 97.9% and 98.6% AUC-ROC (on average) on five different adversarial attacks, and outperforms multiple state-of-the-art detection methods. Results demonstrate the effectiveness of using adversarial gradient directions for adversarial example detection.

cs.CV

GroupIM: A Mutual Information Maximization Framework for Neural Group Recommendation

We study the problem of making item recommendations to ephemeral groups, which comprise users with limited or no historical activities together. Existing studies target persistent groups with substantial activity history, while ephemeral groups lack historical interactions. To overcome group interaction sparsity, we propose data-driven regularization strategies to exploit both the preference covariance amongst users who are in the same group, as well as the contextual relevance of users' individual preferences to each group. We make two contributions. First, we present a recommender architecture-agnostic framework GroupIM that can integrate arbitrary neural preference encoders and aggregators for ephemeral group recommendation. Second, we regularize the user-group latent space to overcome group interaction sparsity by: maximizing mutual information between representations of groups and group members; and dynamically prioritizing the preferences of highly informative members through contextual preference weighting. Our experimental results on several real-world datasets indicate significant performance improvements (31-62% relative NDCG@20) over state-of-the-art group recommendation techniques.

cs.IR

NTIRE 2020 Challenge on Real Image Denoising: Dataset, Methods and Results

This paper reviews the NTIRE 2020 challenge on real image denoising with focus on the newly introduced dataset, the proposed methods and their results. The challenge is a new version of the previous NTIRE 2019 challenge on real image denoising that was based on the SIDD benchmark. This challenge is based on a newly collected validation and testing image datasets, and hence, named SIDD+. This challenge has two tracks for quantitatively evaluating image denoising performance in (1) the Bayer-pattern rawRGB and (2) the standard RGB (sRGB) color spaces. Each track ~250 registered participants. A total of 22 teams, proposing 24 methods, competed in the final phase of the challenge. The proposed methods by the participating teams represent the current state-of-the-art performance in image denoising targeting real noisy images. The newly collected SIDD+ datasets are publicly available at: https://bit.ly/siddplus_data.

cs.CV

Nonuniform Timeslicing of Dynamic Graphs Based on Visual Complexity

Uniform timeslicing of dynamic graphs has been used due to its convenience and uniformity across the time dimension. However, uniform timeslicing does not take the data set into account, which can generate cluttered timeslices with edge bursts and empty timeslices with few interactions. The graph mining filed has explored nonuniform timeslicing methods specifically designed to preserve graph features for mining tasks. In this paper, we propose a nonuniform timeslicing approach for dynamic graph visualization. Our goal is to create timeslices of equal visual complexity. To this end, we adapt histogram equalization to create timeslices with a similar number of events, balancing the visual complexity across timeslices and conveying more important details of timeslices with bursting edges. A case study has been conducted, in comparison with uniform timeslicing, to demonstrate the effectiveness of our approach.

cs.SI

Estimation of common change point and isolation of changed panels after sequential detection

Quick detection of common changes is critical in sequential monitoring of multi-stream data where a common change is referred as a change that only occurs in a portion of panels. After a common change is detected by using a combined CUSUM-SR procedure, we first study the joint distribution for values of the CUSUM process and the estimated delay detection time for the unchanged panels. The BH method by using the asymptotic exponential property for the CUSUM process is developed to isolate the changed panels with the control on FDR. The common change point is then estimated based on the isolated changed panels. Simulation results show that the proposed method can also control the FNR by properly selecting FDR.

math.ST