Searcharxiv⌕ Search

arXiv subjects

Hui Tang

Publications and source records attributed to Hui Tang.

At least 37 records · Page 2Linked to original sources

Efficient Masked Image Compression with Position-Indexed Self-Attention

In recent years, image compression for high-level vision tasks has attracted considerable attention from researchers. Given that object information in images plays a far more crucial role in downstream tasks than background information, some studies have proposed semantically structuring the bitstream to selectively transmit and reconstruct only the information required by these tasks. However, such methods structure the bitstream after encoding, meaning that the coding process still relies on the entire image, even though much of the encoded information will not be transmitted. This leads to redundant computations. Traditional image compression methods require a two-dimensional image as input, and even if the unimportant regions of the image are set to zero by applying a semantic mask, these regions still participate in subsequent computations as part of the image. To address such limitations, we propose an image compression method based on a position-indexed self-attention mechanism that encodes and decodes only the visible parts of the masked image. Compared to existing semantic-structured compression methods, our approach can significantly reduce computational costs.

cs.CV↗

Respiratory Differencing: Enhancing Pulmonary Thermal Ablation Evaluation Through Pre- and Intra-Operative Image Fusion

CT image-guided thermal ablation is widely used for lung cancer treatment; however, follow-up data indicate that physicians' subjective assessments of intraoperative images often overestimate the ablation effect, potentially leading to incomplete treatment. To address these challenges, we developed \textit{Respiratory Differencing}, a novel intraoperative CT image assistance system aimed at improving ablation evaluation. The system first segments tumor regions in preoperative CT images and then employs a multi-stage registration process to align these images with corresponding intraoperative or postoperative images, compensating for respiratory deformations and treatment-induced changes. This system provides two key outputs to help physicians evaluate intraoperative ablation. First, differential images are generated by subtracting the registered preoperative images from the intraoperative ones, allowing direct visualization and quantitative comparison of pre- and post-treatment differences. These differential images enable physicians to assess the relative positions of the tumor and ablation zones, even when the tumor is no longer visible in post-ablation images, thus improving the subjective evaluation of ablation effectiveness. Second, the system provides a quantitative metric that measures the discrepancies between the tumor area and the treatment zone, offering a numerical assessment of the overall efficacy of ablation.This pioneering system compensates for complex lung deformations and integrates pre- and intra-operative imaging data, enhancing quality control in cancer ablation treatments. A follow-up study involving 35 clinical cases demonstrated that our system significantly outperforms traditional subjective assessments in identifying under-ablation cases during or immediately after treatment, highlighting its potential to improve clinical decision-making and patient outcomes.

cs.CV↗

Goal-oriented Feature Extraction: a novel approach for enhancing data-driven surrogate model

Surrogate model can replace the parametric full-order model (FOM) by an approximation model, which can significantly improve the efficiency of optimization design and reduce the complexity of engineering systems. However, due to limitations in efficiency and accuracy, the applications of high-dimensional surrogate models are still challenging. In the present study, we propose a method for extracting hidden features to simplify high-dimensional problems, thereby improving the accuracy and robustness of surrogate models. We establish a goal-oriented feature extraction (GFE) neural network through indirect supervised learning. We constrained the distance between hidden features based on the differences in the target output. This means that in the hidden feature space, cases that are closer in distance output approximately the same, and vice versa. The proposed hidden feature learning method can significantly reduce the dimensionality and nonlinearity of the surrogate model, thereby improving modeling accuracy and generalization capability. To demonstrate the efficiency of our proposed ideas, We conducted numerical experiments on three popular surrogate models. The modeling results of typical high-dimensional mathematical cases and aerodynamic performance cases of ONERA M6 wings show that goal-oriented feature extraction significantly improves the modeling accuracy. Goal-oriented feature extraction can effectively reduce the error distribution of predicting cases and reduce the convergence and robustness differences caused by various data-driven surrogate models.

physics.flu-dyn↗

Adaptive Extensive Cancellation Algorithm and Harmonic Enhanced Heart Rate Estimation based on MMWave Radar

Heart rate (HR) monitoring is crucial for assessing physical fitness, cardiovascular health, and stress management. Millimeter-wave radar offers a promising noncontact solution for long-term monitoring. However, accurate HR estimation remains challenging in low signal-tonoise ratio (SNR) conditions. To deal with both respiration harmonics and intermodulation interference, this paper proposes a cancellation-before-estimation strategy. Firstly, we present the adaptive extensive cancellation algorithm (ECA) to suppress respiratory and its low-order harmonics. Then, we propose an adaptive harmonic enhanced trace (AHET) method to avoid intermodulation interference by refining the HR search region. Various experimental results validate the effectiveness of the proposed methods, demonstrating improvements in accuracy, robustness, and computational efficiency compared to conventional approaches based on the FMCW (Frequency Modulated Continuous Wave) system

eess.SP↗

Unveiling Discrete Clues: Superior Healthcare Predictions for Rare Diseases

Accurate healthcare prediction is essential for improving patient outcomes. Existing work primarily leverages advanced frameworks like attention or graph networks to capture the intricate collaborative (CO) signals in electronic health records. However, prediction for rare diseases remains challenging due to limited co-occurrence and inadequately tailored approaches. To address this issue, this paper proposes UDC, a novel method that unveils discrete clues to bridge consistent textual knowledge and CO signals within a unified semantic space, thereby enriching the representation semantics of rare diseases. Specifically, we focus on addressing two key sub-problems: (1) acquiring distinguishable discrete encodings for precise disease representation and (2) achieving semantic alignment between textual knowledge and the CO signals at the code level. For the first sub-problem, we refine the standard vector quantized process to include condition awareness. Additionally, we develop an advanced contrastive approach in the decoding stage, leveraging synthetic and mixed-domain targets as hard negatives to enrich the perceptibility of the reconstructed representation for downstream tasks. For the second sub-problem, we introduce a novel codebook update strategy using co-teacher distillation. This approach facilitates bidirectional supervision between textual knowledge and CO signals, thereby aligning semantically equivalent information in a shared discrete latent space. Extensive experiments on three datasets demonstrate our superiority.

cs.LG↗

Unconventional spin Hall effect in PT symmetric spin-orbit coupled quantum gases

We theoretically study the intrinsic spin Hall effect in PT symmetric, spin-orbit coupled quantum gases confined in an optical lattice. The interplay of the PT symmetry and the spin-orbit coupling leads to a doubly degenerate non-interacting band structure in which the spin polarization and the Berry curvature of any Bloch state are opposite to those of its degenerate partner. Using experimentally available systems as examples, we show that such a system with a two-component Fermi gas exhibits an intrinsic spin Hall effect akin to that found in the context of electronic materials. For a two-component Bose gas, however, an unconventional spin Hall effect emerges in which the spin polarization and the currents are coplanar and the spin Hall conductivity displays a characteristic anisotropy. We propose to detect such an unconventional spin Hall effect in harmonically trapped systems using dipole oscillations and perform extensive numerical simulations to validate the proposal. Our work paves the way for quantum simulation of the solid-state intrinsic spin Hall effect and experimental explorations of unconventional spin Hall effects in quantum gases.

cond-mat.quant-gas↗

Flow control-oriented coherent mode prediction via Grassmann-kNN manifold learning

A data-driven method using Grassmann manifold learning is proposed to identify a low-dimensional actuation manifold for flow-controlled fluid flows. The snapshot flow field are twice compressed using Proper Orthogonal Decomposition (POD) and a diffusion model. Key steps of the actuation manifold are Grassmann manifold-based Polynomial Chaos Expansion (PCE) as the encoder and K-nearest neighbor regression (kNN) as the decoder. This methodology is first tested on a simple dielectric cylinder in a homogeneous electric field to predict the out-of-sample electric field, demonstrating fast and accurate performance. Next, the present model is evaluated by predicting dynamic coherence modes of an oscillating-rotation cylinder. The cylinder's oscillating rotation amplitude and frequency are regarded as independent control parameters. The mean mode and the first dynamic mode are selected as the representative cases to test present model. For the mean mode, the Grassman manifold describes all parameterized modes with 8 latent variables. All the modes can be divided into four clusters, and they share similar features but with different wake length. For the dynamic mode, the Grassman manifold describes all modes with 12 latent variables. All the modes can be divided into three clusters. Intriguingly, each cluster is aligned with clear physical meanings. One describes the near-wake periodic vortex shedding resembling Karman vortices, one describes the far wake periodic vortex shedding, and one shows high-frequency K-H vortices shedding. Moreover, Grassmann-kNN manifold learning can accurately predict the modes. It is possible to estimate the full flow state with small reconstruction errors just by knowing the actuation parameters. This manifold learning model is demonstrated to be crucial for flow control-oriented flow estimation.

physics.flu-dyn↗

FITA: Fine-grained Image-Text Aligner for Radiology Report Generation

Radiology report generation aims to automatically generate detailed and coherent descriptive reports alongside radiology images. Previous work mainly focused on refining fine-grained image features or leveraging external knowledge. However, the precise alignment of fine-grained image features with corresponding text descriptions has not been considered. This paper presents a novel method called Fine-grained Image-Text Aligner (FITA) to construct fine-grained alignment for image and text features. It has three novel designs: Image Feature Refiner (IFR), Text Feature Refiner (TFR) and Contrastive Aligner (CA). IFR and TFR aim to learn fine-grained image and text features, respectively. We achieve this by leveraging saliency maps to effectively fuse symptoms with corresponding abnormal visual regions, and by utilizing a meticulously constructed triplet set for training. Finally, CA module aligns fine-grained image and text features using contrastive loss for precise alignment. Results show that our method surpasses existing methods on the widely used benchmark

cs.CV↗

Real-Time 4K Super-Resolution of Compressed AVIF Images. AIS 2024 Challenge Survey

This paper introduces a novel benchmark as part of the AIS 2024 Real-Time Image Super-Resolution (RTSR) Challenge, which aims to upscale compressed images from 540p to 4K resolution (4x factor) in real-time on commercial GPUs. For this, we use a diverse test set containing a variety of 4K images ranging from digital art to gaming and photography. The images are compressed using the modern AVIF codec, instead of JPEG. All the proposed methods improve PSNR fidelity over Lanczos interpolation, and process images under 10ms. Out of the 160 participants, 25 teams submitted their code and models. The solutions present novel designs tailored for memory-efficiency and runtime on edge devices. This survey describes the best solutions for real-time SR of compressed high-resolution images.

cs.CV↗

NTIRE 2024 Challenge on Low Light Image Enhancement: Methods and Results

This paper reviews the NTIRE 2024 low light image enhancement challenge, highlighting the proposed solutions and results. The aim of this challenge is to discover an effective network design or solution capable of generating brighter, clearer, and visually appealing results when dealing with a variety of conditions, including ultra-high resolution (4K and beyond), non-uniform illumination, backlighting, extreme darkness, and night scenes. A notable total of 428 participants registered for the challenge, with 22 teams ultimately making valid submissions. This paper meticulously evaluates the state-of-the-art advancements in enhancing low-light images, reflecting the significant progress and creativity in this field.

cs.CV↗

A multi-stage semi-supervised learning for ankle fracture classification on CT images

Because of the complicated mechanism of ankle injury, it is very difficult to diagnose ankle fracture in clinic. In order to simplify the process of fracture diagnosis, an automatic diagnosis model of ankle fracture was proposed. Firstly, a tibia-fibula segmentation network is proposed for the joint tibiofibular region of the ankle joint, and the corresponding segmentation dataset is established on the basis of fracture data. Secondly, the image registration method is used to register the bone segmentation mask with the normal bone mask. Finally, a semi-supervised classifier is constructed to make full use of a large number of unlabeled data to classify ankle fractures. Experiments show that the proposed method can segment fractures with fracture lines accurately and has better performance than the general method. At the same time, this method is superior to classification network in several indexes.

eess.IV↗

An enthalpy-based model for the physics of ice crystal icing

Ice crystal icing (ICI) in aircraft engines is a major threat to flight safety. Due to the complex thermodynamic and phase-change conditions involved in ICI, rigorous modelling of the accretion process remains limited. The present study proposes a novel modelling approach based on the physically-observed mixed-phase nature of the accretion layers. The mathematical model, which is derived from the enthalpy change after accretion (the enthalpy model), is compared to an existing pure-phase layer model (the three-layer model). Scaling laws and asymptotic solutions are developed for both models. The onset of ice accretion, the icing layer thickness, and solid ice fraction within the layer are determined by a set of non-dimensional parameters including the Peclet number, the Stefan number, the Biot number, the Melt Ratio, and the evaporative rate. Thresholds for freezing and non-freezing conditions are developed. The asymptotic solutions presents good agreement with numerical solutions at low Peclet numbers. Both the asymptotic and numerical solutions show that, when compared to the three-layer model, the enthalpy model presents a thicker icing layer and a thicker water layer above the substrate due to mixed-phased features and modified Stefan conditions. Modelling in terms of the enthalpy poses significant advantages in the development of numerical methods to complex three-dimensional geometrical and flow configurations. These results improve understanding of the accretion process and provide a novel, rigorous mathematical framework for accurate modelling of ICI.

physics.flu-dyn↗

Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models

With the emergence of pre-trained vision-language models like CLIP, how to adapt them to various downstream classification tasks has garnered significant attention in recent research. The adaptation strategies can be typically categorized into three paradigms: zero-shot adaptation, few-shot adaptation, and the recently-proposed training-free few-shot adaptation. Most existing approaches are tailored for a specific setting and can only cater to one or two of these paradigms. In this paper, we introduce a versatile adaptation approach that can effectively work under all three settings. Specifically, we propose the dual memory networks that comprise dynamic and static memory components. The static memory caches training data knowledge, enabling training-free few-shot adaptation, while the dynamic memory preserves historical test features online during the testing process, allowing for the exploration of additional data insights beyond the training set. This novel capability enhances model performance in the few-shot setting and enables model usability in the absence of training data. The two memory networks employ the same flexible memory interactive strategy, which can operate in a training-free mode and can be further enhanced by incorporating learnable projection layers. Our approach is tested across 11 datasets under the three task settings. Remarkably, in the zero-shot scenario, it outperforms existing methods by over 3\% and even shows superior results against methods utilizing external training data. Additionally, our method exhibits robust performance against natural distribution shifts. Codes are available at \url{https://github.com/YBZh/DMN}.

cs.CV↗

Dual-Decoder Consistency via Pseudo-Labels Guided Data Augmentation for Semi-Supervised Medical Image Segmentation

While supervised learning has achieved remarkable success, obtaining large-scale labeled datasets in biomedical imaging is often impractical due to high costs and the time-consuming annotations required from radiologists. Semi-supervised learning emerges as an effective strategy to overcome this limitation by leveraging useful information from unlabeled datasets. In this paper, we present a novel semi-supervised learning method, Dual-Decoder Consistency via Pseudo-Labels Guided Data Augmentation (DCPA), for medical image segmentation. We devise a consistency regularization to promote consistent representations during the training process. Specifically, we use distinct decoders for student and teacher networks while maintain the same encoder. Moreover, to learn from unlabeled data, we create pseudo-labels generated by the teacher networks and augment the training data with the pseudo-labels. Both techniques contribute to enhancing the performance of the proposed method. The method is evaluated on three representative medical image segmentation datasets. Comprehensive comparisons with state-of-the-art semi-supervised medical image segmentation methods were conducted under typical scenarios, utilizing 10% and 20% labeled data, as well as in the extreme scenario of only 5% labeled data. The experimental results consistently demonstrate the superior performance of our method compared to other methods across the three semi-supervised settings. The source code is publicly available at https://github.com/BinYCn/DCPA.git.

eess.IV↗

O(1) benchmarking of precise rotation in a spin-squeezed Bose-Einstein condensate

Benchmarking a high-precision quantum operation is a big challenge for many quantum systems in the presence of various noises as well as control errors. Here we propose an $O(1)$ benchmarking of a dynamically corrected rotation by taking the quantum advantage of a squeezed spin state in a spin-1 Bose-Einstein condensate. Our analytical and numerical results show that tiny rotation infidelity, defined by $1-F$ with $F$ the rotation fidelity, can be calibrated in the order of $1/N^2$ by only several measurements of the rotation error for $N$ atoms in an optimally squeezed spin state. Such an $O(1)$ benchmarking is possible not only in a spin-1 BEC but also in other many-spin or many-qubit systems if a squeezed or entangled state is available.

cond-mat.quant-gas↗

Simulation of magnetic hyperthermia cancer treatment near a blood vessel

In this study, we conduct a study on magnetic hyperthermia treatment when a vessel is located near the tumor. The holistic framework is established to solve the process of tumor treatment. The interstitial tissue fluid, MNP distribution, temperature profile, and nanofluids are involved in the simulation. The study evaluates the cancer treatment efficacy by cumulative-equivalent-minutes-at-43 centigrade (CEM43), a widely accepted thermal dose. The influence of the nearby blood vessel is investigated, and parameter studies about the distance to the tumor, the width of the blood vessel, and the vessel direction are also conducted. After that, the effects of the fluid-structure interaction of moving vessel boundaries and blood rheology are discussed. The results demonstrate the cooling effect of a nearby blood vessel, and such effect reduces with the augment of the distance between the tumor and blood vessel. The combination of downward gravity and the cool effect from the lower horizontal vessel leads to the best performance with 97.77% ablation in the tumor and 0.87% injury in healthy tissue at distance d = 4 mm, but the cases of the vertical vessel are relatively poor. The vessel width and blood rheology affect the treatment by velocity gradient near the vessel wall. Additionally, the moving boundary almost has no impact on treatment efficacy. The simulation tool has been well-validated and the outcomes can provide a useful reference to magnetic hyperthermia treatment.

physics.flu-dyn↗

Simulation of tumor ablation in hyperthermia cancer treatment: A parametric study

A holistic simulation framework is established on magnetic hyperthermia modeling to solve the treatment process of tumor, which is surrounded by a healthy tissue block. The interstitial tissue fluid, MNP distribution, temperature profile, and nanofluids are involved in the simulation. Study evaluates the cancer treatment efficacy by cumulative-equivalent-minutes-at-43 centigrade (CEM43), a widely accepted thermal dose coming from the cell death curve. Results are separated into the conditions of with or without gravity effect in the computational domain, where two baseline case are investigated and compared. An optimal treatment time 46.55 min happens in the baseline case without gravity, but the situation deteriorates with gravity effect where the time for totally killing tumor cells prolongs 36.11% and meanwhile causing 21.32% ablation in healthy tissue. For the cases without gravity, parameter study of Lewis number and Heat source number are conducted and the variation of optimal treatment time are both fitting to the inverse functions. For the case considering the gravity, parameters Buoyancy ratio and Darcy ratio are investigated and their influence on totally killing tumor cells and the injury on healthy tissue are matching with the parabolic functions. The results are beneficial to the prediction of various conditions, and provides useful guide to the magnetic hyperthermia treatment.

physics.med-ph↗

A New Benchmark: On the Utility of Synthetic Data with Blender for Bare Supervised Learning and Downstream Domain Adaptation

Deep learning in computer vision has achieved great success with the price of large-scale labeled training data. However, exhaustive data annotation is impracticable for each task of all domains of interest, due to high labor costs and unguaranteed labeling accuracy. Besides, the uncontrollable data collection process produces non-IID training and test data, where undesired duplication may exist. All these nuisances may hinder the verification of typical theories and exposure to new findings. To circumvent them, an alternative is to generate synthetic data via 3D rendering with domain randomization. We in this work push forward along this line by doing profound and extensive research on bare supervised learning and downstream domain adaptation. Specifically, under the well-controlled, IID data setting enabled by 3D rendering, we systematically verify the typical, important learning insights, e.g., shortcut learning, and discover the new laws of various data regimes and network architectures in generalization. We further investigate the effect of image formation factors on generalization, e.g., object scale, material texture, illumination, camera viewpoint, and background in a 3D scene. Moreover, we use the simulation-to-reality adaptation as a downstream task for comparing the transferability between synthetic and real data when used for pre-training, which demonstrates that synthetic data pre-training is also promising to improve real test results. Lastly, to promote future research, we develop a new large-scale synthetic-to-real benchmark for image classification, termed S2RDA, which provides more significant challenges for transfer from simulation to reality. The code and datasets are available at https://github.com/huitangtang/On_the_Utility_of_Synthetic_Data.

cs.CV↗