SearcharxivSearch

arXiv subjects

Lipo Wang

Publications and source records attributed to Lipo Wang.

At least 19 recordsLinked to original sources

When is Test-Time Adaptation Identifiable From Unlabeled Evidence?

Test-time adaptation (TTA) offers many ways to update a deployed model without labels, but choosing the wrong update can make a strong source model worse. Recent methods therefore try to predict which adaptation will work from unlabeled test data. We ask a prior question: does the evidence given to the selector contain enough information to determine the best action at all? We show that this is not guaranteed, even with a perfect selector. If an observation channel makes two deployments look the same while their TTA rankings differ, reliable selection is impossible from that channel; richer evidence can restore the decision only when it resolves the relevant ambiguity. We make this boundary exact in a finite-batch Gaussian TTA model, where doing nothing beats mean recentering for small shifts, recentering wins beyond a unique critical shift, and the boundary shrinks as $1/\sqrt n$. Public benchmark studies on CIFAR-100-C and DomainNet-126 show the same failure mode with modern TTA methods: changing only deployment structure can reverse the oracle action while global order-blind evidence remains unchanged. The result is a practical way to separate two failure modes that are usually mixed together: a weak selector versus an information channel that cannot support the desired decision in the first place.

cs.CV

Filtered turbulent flame model with wrinkling correction on chemical source for nonpremixed combustion simulation

One of the most critical challenges in turbulent combustion modeling is the chemical source closure. In the recently developed filtered turbulent flame model (FTFM), a oneto-one correspondence between filtered scalar quantities and filtered chemical sources can be constructed by inversely solving the filtered flame equations, without the use of conventional presumed probability density functions (PDFs). However, the turbulence induced flame wrinkling, and thus the enhancement of the chemical source, has not been explicitly considered at the resolved scale. In the present study, the wrinkling effect is analytically quantified in a counterfow flame setup, from which FTFM is then further updated by incorporating such a physics grounded stretching correction on the chemical source. The satisfactory accuracy and robustness of the present model are justified from case tests of the non-premixed Sydney swirl flame and the Delft III flame.

physics.flu-dyn

FedTopo: Relation-Level Topology Sharing for Model-Heterogeneous Federated Learning

Federated learning (FL) enables collaborative learning over decentralized data silos without centralizing raw data. However, heterogeneous local architectures often induce non-aligned representation spaces, making it difficult to transfer global knowledge across silos. Existing paradigms share this knowledge as model parameters, distilled predictions, or class prototypes, yet all encode it in an absolute space that must be aligned across clients. Heterogeneous backbones break this alignment, so the shared knowledge becomes unreliable and misleads local training. We propose FedTopo, a relation-level framework that encodes global knowledge as class relation topology, capturing how classes relate within each client rather than where they lie in feature space. Each client builds its relation topology from local prototypes and uploads it with class statistics. The server then aggregates these relations in a reliability-aware manner that down-weights weakly supported ones, and broadcasts the global topology to clients. The global topology guides local training by emphasizing topology-similar negative classes. Experiments on three datasets under eight heterogeneous backbones show that FedTopo consistently outperforms parameter-, distillation-, and prototype-sharing baselines, with low communication and no inference overhead. Our code is available at https://github.com/Zhaoyang-Ma/FedTopo.

cs.LG

AURA: Active-Response Attribution under Treatment Ambiguity in Bacterial Cytological Profiling

When a bacterial sample is exposed to several antibiotics, not every applied drug necessarily acts: if the organism is resistant to one of them, that drug leaves no morphological trace. The clinically meaningful quantity is therefore not which antibiotics were applied, but which ones were active. We show that these two are sharply decoupled in real E. coli microscopy - naively assuming the applied combination equals the active one is correct only about 37% of the time - yet existing computational tools are ill-suited to recovering the active set. Forward perturbation models such as scGen, CPA, and IMPA are designed to predict appearance from treatment, not the reverse, and inverting them degrades sharply; discriminative image classifiers tend to memorise strain- and batch-specific texture and fail to transfer across experimental replicates. We introduce AURA, which reframes the task as constrained, energy-based inverse attribution. Its central inductive bias is that the active set must be a subset of the applied set; this collapses the candidate space and lets AURA infer the active subset of applied antibiotics by decomposing residual morphology into antibiotic response atoms and selecting the subset with the lowest reconstruction energy, using no strain label at test time. AURA-E adds evidence-aware abstention, withholding a prediction when candidate explanations remain near-equally plausible. On cross-replicate transfer in an E. coli cytological profiling dataset, AURA recovers the active antibiotic combination with 95.47% exact-match accuracy.

cs.CV

CHASE: Competing Hypotheses for Ambiguity-Aware Selective Prediction

Standard selective prediction methods typically estimate uncertainty from the output of a single predictive branch. While effective for general uncertainty estimation, these approaches often struggle under partial observability, where local temporal evidence can be contradictory and standard confidence scores become misleading. We introduce CHASE (Competing Hypotheses for Ambiguity-Aware Selective Prediction), a selective prediction framework that explicitly compares structured temporal explanations to determine whether to commit to a decision or abstain. Because genuine ambiguity causes the score gap between competing hypotheses to collapse, CHASE optimizes a ranking-aware selector over these hypothesis margins to globally separate safe commitments from fundamentally uncertain ones. We evaluate this framework on the problem of hidden connectivity inference, utilizing a controlled, physically grounded simulator inspired by the dynamics of giant unilamellar vesicles (GUVs), alongside zero-shot qualitative transfer (without retraining or fine tuning) to representative real GUV videos. Our experiments demonstrate that explicitly reasoning over competing hypotheses provides a superior balance of metrics. Compared to canonical uncertainty baselines, CHASE achieves statistically significant gains in overall no-abstain accuracy, three-way accuracy, and overall ambiguity-aligned abstention (at 80% coverage). Specifically, it yields up to an 11.0% relative mean improvement in overall alignment, alongside up to an 8.8% relative boost in three-way accuracy in the very-high ambiguity regime. By maintaining a selective risk boundary strictly at par with the best baselines at 80% coverage, and reducing overall risk by 9.9% at 90% coverage, this framework offers a more reliable approach to decision-making under structured ambiguity.

cs.CV

HD-TTA: Hypothesis-Driven Test-Time Adaptation for Safer Brain Tumor Segmentation

Standard Test-Time Adaptation (TTA) methods typically treat inference as a blind optimization task, applying generic objectives to all or filtered test samples. In safety-critical medical segmentation, this lack of selectivity often causes the tumor mask to spill into healthy brain tissue or degrades predictions that were already correct. We propose Hypothesis-Driven TTA, a novel framework that reformulates adaptation as a dynamic decision process. Rather than forcing a single optimization trajectory, our method generates intuitive competing geometric hypotheses: compaction (is the prediction noisy? trim artifacts) versus inflation (is the valid tumor under-segmented? safely inflate to recover). It then employs a representation-guided selector to autonomously identify the safest outcome based on intrinsic texture consistency. Additionally, a pre-screening Gatekeeper prevents negative transfer by skipping adaptation on confident cases. We validate this proof-of-concept on a cross-domain binary brain tumor segmentation task, applying a source model trained on adult BraTS gliomas to unseen pediatric and more challenging meningioma target domains. HD-TTA improves safety-oriented outcomes (Hausdorff Distance (HD95) and Precision) over several state-of-the-art representative baselines in the challenging safety regime, reducing the HD95 by approximately 6.4 mm and improving Precision by over 4%, while maintaining comparable Dice scores. These results demonstrate that resolving the safety-adaptation trade-off via explicit hypothesis selection is a viable, robust path for safe clinical model deployment. Code will be made publicly available upon acceptance.

cs.CV

Analysis of near wall flame and wall heat flux modeling in turbulent premixed combustion

Reactive flows in confined spaces involve complex flame-wall interaction (FWI). This work aims to gain more insights into the physics of the premixed near-wall flame and the wall heat flux as an important engineering relevant quantity. Two different flame configurations have been studied, including the normal flushing flame and inclined sweeping flame. By introducing the skin friction vector defined second-order tensor, direct numerical simulation (DNS) results of these two configurations show consistently that larger flame curvatures are associated with small vorticity magnitude under the influence of the vortex pair structure. Correlation of both the flame normal and tangential strain rates with the flame curvature has also been quantified. Alignment of the progress variable gradient with the most compressive eigenvector on the wall is similar to the boundary free behavior. To characterize the flame ordered structure, especially in the near-wall region, a species alignment index is proposed. The big difference in this index for flames in different regions suggests distinct flame structures. Building upon these fundamental insights, a predictive model for wall heat flux is proposed. For the purpose of applicability, realistic turbulent combustion situations need to be taken into account, for instance, flames with finite thickness, complex chemical kinetics, non-negligible near-wall reactions, and variable flame orientation relative to the wall. The model is first tested in an one-dimensional laminar flame and then validated against DNS datasets, justifying the model performance with satisfying agreement.

physics.flu-dyn

Harmonizing Intra-coherence and Inter-divergence in Ensemble Attacks for Adversarial Transferability

The development of model ensemble attacks has significantly improved the transferability of adversarial examples, but this progress also poses severe threats to the security of deep neural networks. Existing methods, however, face two critical challenges: insufficient capture of shared gradient directions across models and a lack of adaptive weight allocation mechanisms. To address these issues, we propose a novel method Harmonized Ensemble for Adversarial Transferability (HEAT), which introduces domain generalization into adversarial example generation for the first time. HEAT consists of two key modules: Consensus Gradient Direction Synthesizer, which uses Singular Value Decomposition to synthesize shared gradient directions; and Dual-Harmony Weight Orchestrator which dynamically balances intra-domain coherence, stabilizing gradients within individual models, and inter-domain diversity, enhancing transferability across models. Experimental results demonstrate that HEAT significantly outperforms existing methods across various datasets and settings, offering a new perspective and direction for adversarial attack research.

cs.LG

Local-peak scale-invariant feature transform for fast and random image stitching

Image stitching aims to construct a wide field of view with high spatial resolution, which cannot be achieved in a single exposure. Typically, conventional image stitching techniques, other than deep learning, require complex computation and thus computational pricy, especially for stitching large raw images. In this study, inspired by the multiscale feature of fluid turbulence, we developed a fast feature point detection algorithm named local-peak scale-invariant feature transform (LP-SIFT), based on the multiscale local peaks and scale-invariant feature transform method. By combining LP-SIFT and RANSAC in image stitching, the stitching speed can be improved by orders, compared with the original SIFT method. Nine large images (over 2600*1600 pixels), arranged randomly without prior knowledge, can be stitched within 158.94 s. The algorithm is highly practical for applications requiring a wide field of view in diverse application scenes, e.g., terrain mapping, biological analysis, and even criminal investigation.

cs.CV

Is a direct numerical simulation (DNS) of Navier-Stokes equations with small enough grid spacing and time-step definitely reliable/correct?

Traditionally, results given by the direct numerical simulation (DNS) of Navier-Stokes equations are widely regarded as reliable benchmark solutions of turbulence, as long as grid spacing is fine enough (i.e. less than the minimum Kolmogorov scale) and time-step is small enough, say, satisfying the Courant-Friedrichs-Lewy condition. Is this really true? In this paper a two-dimensional sustained turbulent Kolmogorov flow is investigated numerically by the two numerical methods with detailed comparisons: one is the traditional `direct numerical simulation' (DNS), the other is the `clean numerical simulation' (CNS). The results given by DNS are a kind of mixture of the false numerical noise and the true physical solution, which however are mostly at the same order of magnitude due to the butterfly-effect of chaos. On the contrary, the false numerical noise of the results given by CNS is much smaller than the true physical solution of turbulence in a long enough interval of time so that a CNS result is very close to the true physical solution and thus can be used as a benchmark solution. It is found that numerical noise as a kind of artificial tiny disturbances can lead to huge deviations at large scale on the two-dimensional Kolmogorov turbulence, not only quantitatively (even in statistics) but also qualitatively (such as symmetry of flow). Thus, fine enough spatial grid spacing with small enough time-step alone cannot guarantee the validity of the DNS: it is only a necessary condition but not sufficient. This finding might challenge some assumptions in investigation of turbulence. So, DNS results of a few sustained turbulent flows might have huge deviations on both of small and large scales from the true solution of Navier-Stokes equations even in statistics. Hopefully, CNS as a new tool to investigate turbulent flows more accurately than DNS could bring us some new discoveries.

physics.flu-dyn

Fast Blind Recovery of Linear Block Codes over Noisy Channels

This paper addresses the blind recovery of the parity check matrix of an (n,k) linear block code over noisy channels by proposing a fast recovery scheme consisting of 3 parts. Firstly, this scheme performs initial error position detection among the received codewords and selects the desirable codewords. Then, this scheme conducts Gaussian elimination (GE) on a k-by-k full-rank matrix and uses a threshold and the reliability associated to verify the recovered dual words, aiming to improve the reliability of recovery. Finally, it performs decoding on the received codewords with partially recovered dual words. These three parts can be combined into different schemes for different noise level scenarios. The GEV that combines Gaussian elimination and verification has a significantly lower recovery failure probability and a much lower computational complexity than an existing Canteaut-Chabaud-based algorithm, which relies on GE on n-by-n full-rank matrices. The decoding-aided recovery (DAR) and error-detection-&-codeword-selection-&-decoding-aided recovery (EDCSDAR) schemes can improve the code recovery performance over GEV for high noise level scenarios, and their computational complexities remain much lower than the Canteaut-Chabaud-based algorithm.

cs.IT

Subject-Independent Drowsiness Recognition from Single-Channel EEG with an Interpretable CNN-LSTM model

For EEG-based drowsiness recognition, it is desirable to use subject-independent recognition since conducting calibration on each subject is time-consuming. In this paper, we propose a novel Convolutional Neural Network (CNN)-Long Short-Term Memory (LSTM) model for subject-independent drowsiness recognition from single-channel EEG signals. Different from existing deep learning models that are mostly treated as black-box classifiers, the proposed model can explain its decisions for each input sample by revealing which parts of the sample contain important features identified by the model for classification. This is achieved by a visualization technique by taking advantage of the hidden states output by the LSTM layer. Results show that the model achieves an average accuracy of 72.97% on 11 subjects for leave-one-out subject-independent drowsiness recognition on a public dataset, which is higher than the conventional baseline methods of 55.42%-69.27%, and state-of-the-art deep learning methods. Visualization results show that the model has discovered meaningful patterns of EEG signals related to different mental states across different subjects.

cs.NE

BGaitR-Net: Occluded Gait Sequence reconstructionwith temporally constrained model for gait recognition

Recent advancements in computational resources and Deep Learning methodologies has significantly benefited development of intelligent vision-based surveillance applications. Gait recognition in the presence of occlusion is one of the challenging research topics in this area, and the solutions proposed by researchers to date lack in robustness and also dependent of several unrealistic constraints, which limits their practical applicability. We improve the state-of-the-art by developing novel deep learning-based algorithms to identify the occluded frames in an input sequence and next reconstruct these occluded frames by exploiting the spatio-temporal information present in the gait sequence. The multi-stage pipeline adopted in this work consists of key pose mapping, occlusion detection and reconstruction, and finally gait recognition. While the key pose mapping and occlusion detection phases are done %using Constrained KMeans Clustering and via a graph sorting algorithm, reconstruction of occluded frames is done by fusing the key pose-specific information derived in the previous step along with the spatio-temporal information contained in a gait sequence using a Bi-Directional Long Short Time Memory. This occlusion reconstruction model has been trained using synthetically occluded CASIA-B and OU-ISIR data, and the trained model is termed as Bidirectional Gait Reconstruction Network BGait-R-Net. Our LSTM-based model reconstructs occlusion and generates frames that are temporally consistent with the periodic pattern of a gait cycle, while simultaneously preserving the body structure.

cs.CV

A quadratic Reynolds stress development for the turbulent Kolmogorov flow

We study the three-dimensional turbulent Kolmogorov flow, i.e. the Navier-Stokes equations forced by a low-single-wave-number sinusoidal force in a periodic domain, by means of direct numerical simulations. This classical model system is a realization of anisotropic and non-homogeneous hydrodynamic turbulence. Boussinesq's eddy viscosity linear relation is checked and found to be approximately valid over half of the system volume. A more general nonlinear quadratic Reynolds stress development is proposed and its parameters estimated at varying the Taylor scale-based Reynolds number in the flow up to the value 200. The case of a forcing with a different shape, here chosen Gaussian, is considered and the differences with the sinusoidal forcing are emphasized.

physics.flu-dyn

3D Deep Learning on Medical Images: A Review

The rapid advancements in machine learning, graphics processing technologies and the availability of medical imaging data have led to a rapid increase in the use of deep learning models in the medical domain. This was exacerbated by the rapid advancements in convolutional neural network (CNN) based architectures, which were adopted by the medical imaging community to assist clinicians in disease diagnosis. Since the grand success of AlexNet in 2012, CNNs have been increasingly used in medical image analysis to improve the efficiency of human clinicians. In recent years, three-dimensional (3D) CNNs have been employed for the analysis of medical images. In this paper, we trace the history of how the 3D CNN was developed from its machine learning roots, we provide a brief mathematical description of 3D CNN and provide the preprocessing steps required for medical images before feeding them to 3D CNNs. We review the significant research in the field of 3D medical imaging analysis using 3D CNNs (and its variants) in different medical areas such as classification, segmentation, detection and localization. We conclude by discussing the challenges associated with the use of 3D CNNs in the medical imaging domain (and the use of deep learning models in general) and possible future trends in the field.

q-bio.QM

Fluctuations and correlations of reactive scalars near chemical equilibrium in incompressible turbulence

The statistical properties of species undergoing chemical reactions in a turbulent environment are studied. We focus on the case of reversible multi-component reactions of second and higher orders, in a condition close to chemical equilibrium sustained by random large-scale reactant sources, while the turbulent flow is highly developed. In such a state a competition exists between the chemical reaction that tends to dump reactant concentration fluctuations and enhance their correlation intensity and the turbulent mixing that on the contrary increases fluctuations and remove relative correlations. We show that a unique control parameter, the Damkhöler number ($Da_θ$) that can be constructed from the scalar Taylor micro-scale, the reactant diffusivity and the reaction rate characterises the functional dependence of fluctuations and correlations in a variety of conditions, i.e., at changing the reaction order, the Reynolds and the Schmidt numbers. The larger is such a Damkhöler number the more depleted are the scalar fluctuations as compared to the fluctuations of a passive scalar field in the same conditions, and vice-versa the more intense are the correlations. A saturation in this behaviour is observed beyond $Da_θ\simeq \mathcal{O}(10)$. We provide an analytical prediction for this phenomenon which is in excellent agreement with direct numerical simulation results.

physics.flu-dyn

Speech Fusion to Face: Bridging the Gap Between Human's Vocal Characteristics and Facial Imaging

While deep learning technologies are now capable of generating realistic images confusing humans, the research efforts are turning to the synthesis of images for more concrete and application-specific purposes. Facial image generation based on vocal characteristics from speech is one of such important yet challenging tasks. It is the key enabler to influential use cases of image generation, especially for business in public security and entertainment. Existing solutions to the problem of speech2face renders limited image quality and fails to preserve facial similarity due to the lack of quality dataset for training and appropriate integration of vocal features. In this paper, we investigate these key technical challenges and propose Speech Fusion to Face, or SF2F in short, attempting to address the issue of facial image quality and the poor connection between vocal feature domain and modern image generation models. By adopting new strategies on data model and training, we demonstrate dramatic performance boost over state-of-the-art solution, by doubling the recall of individual identity, and lifting the quality score from 15 to 19 based on the mutual information score with VGGFace classifier.

cs.CV

Multi-level scalar structure in complex system analyses

The geometrical structure is among the most fundamental ingredients in understanding complex systems. Is there any systematic approach in defining structures quantitatively, rather than illustratively? If yes, what are the basic principles to follow? By introducing the concept of extremal points at different scale levels, a multi-level dissipation element approach has been developed to define structures at different scale levels, in accordance with the concept of structure hierarchy. Each dissipation element can be characterized by the length scale and the scalar variance inside. Using the two-dimensional fractal Brownian motion as a benchmark case, the conditional mean of the scalar difference with respect to the length scale shows clearly a power law and the scaling exponent is in agreement with the Hurst number. For the 3D turbulence velocity component, the 1/3 scaling law can be represented. These results indicate the important linkage between the turbulence physics and ow structure, if well posed and defined. In principle, the multi-level dissipation element idea is generally applicable in analyzing other multiscale complex systems as well.

physics.flu-dyn