SearcharxivSearch

arXiv subjects

Hongtao Chen

Publications and source records attributed to Hongtao Chen.

12 recordsLinked to original sources

Hierarchical Semantic-Constrained Heterogeneous Graph for Audio-Visual Event Localization

Open-vocabulary audio-visual event localization (OV-AVEL) jointly models audio-visual cues to recognize and temporally localize events, including categories unseen during training. Existing methods primarily learn joint audio-visual representations in Euclidean space, but still face two significant challenges. First, the lack of supervision signals for unseen categories makes it difficult to maintain audio-visual consistency across multiple temporal scales. Second, the lack of hierarchical constraints between segment- and video-level semantics prevents the model from establishing semantic consistency across different levels. To address these challenges, we propose a hierarchical semantic constrained heterogeneous graph (HSCHG) for audio-visual event localization framework. We first construct a heterogeneous hierarchical graph in Euclidean space, which includes audio and visual segment nodes and their corresponding video-level nodes. We use multi-directional temporal edges to capture complete temporal information within each modality. Simultaneously, we employ a dual-threshold filtering gated fusion strategy, introducing cross-modal information only when the alignment confidence is high. Furthermore, we introduce bidirectional semantic constraints between segment- and video-level representations to achieve semantic consistency across different levels. Based on this, we map the multi-level audio-visual representations and text prototypes uniformly into hyperbolic space. We use a hierarchical entailment regularization loss to characterize the hierarchical relationships between videos and segments. Extensive experimental results show that our method outperforms existing methods on the OV-AVEL benchmark. Ablation studies further validate the effectiveness of our method.

cs.AI

Refined convergence structures of the rectangular Raviart-Thomas element

In this work, we fully explore three refined convergence structures of the lowest-order rectangular Raviart-Thomas element in solving the Laplace eigenvalue problem. Firstly, the scheme possesses a property of supercloseness between the discrete eigenfunctions and the interpolated ones, so that post-processing can be easily constructed to improve the accuracy at most by one order. The essentially skillful method is the integral expansion for interpolation terms. Secondly, based on the supercloseness property, we derive the error expansions for not only simple eigenvalues but also multiple eigenvalues, and provide a rigorous proof for them, based on which Richardson extrapolation can be performed. As a byproduct, we prove that all eigenvalues converge from above. Moreover, by utilizing the supercloseness property and Rayleigh quotient analysis, we give a rigorous proof for the convergence behavior for multiple eigenvalues on uniform meshes for the problem on the square domain. Thirdly, the equivalence between the lowest-order rectangular Raviart-Thomas element and the enriched rotated bilinear element is also indicated. At the last of this work, several numerical experiments are designed to demonstrate our theory.

math.NA

All-in-One Video Restoration under Smoothly Evolving Unknown Weather Degradations

All-in-one image restoration aims to recover clean images from diverse unknown degradations using a single model. But extending this task to videos faces unique challenges. Existing approaches primarily focus on frame-wise degradation variation, overlooking the temporal continuity that naturally exists in real-world degradation processes. In practice, degradation types and intensities evolve smoothly over time, and multiple degradations may coexist or transition gradually. In this paper, we introduce the Smoothly Evolving Unknown Degradations (SEUD) scenario, where both the active degradation set and degradation intensity change continuously over time. To support this scenario, we design a flexible synthesis pipeline that generates temporally coherent videos with single, compound, and evolving degradations. To address the challenges in the SEUD scenario, we propose an all-in-One Recurrent Conditional and Adaptive prompting Network (ORCANet). First, a Coarse Intensity Estimation Dehazing (CIED) module estimates haze intensity using physical priors and provides coarse dehazed features as initialization. Second, a Flow Prompt Generation (FPG) module extracts degradation features. FPG generates both static prompts that capture segment-level degradation types and dynamic prompts that adapt to frame-level intensity variations. Furthermore, a label-aware supervision mechanism improves the discriminability of static prompt representations under different degradations. Extensive experiments show that ORCANet achieves superior restoration quality, temporal consistency, and robustness over image and video-based baselines. Code is available at https://github.com/Friskknight/ORCANet-SEUD.

cs.CV

Adaptive Redundancy Regulation for Balanced Multimodal Information Refinement

Multimodal learning aims to improve performance by leveraging data from multiple sources. During joint multimodal training, due to modality bias, the advantaged modality often dominates backpropagation, leading to imbalanced optimization. Existing methods still face two problems: First, the long-term dominance of the dominant modality weakens representation-output coupling in the late stages of training, resulting in the accumulation of redundant information. Second, previous methods often directly and uniformly adjust the gradients of the advantaged modality, ignoring the semantics and directionality between modalities. To address these limitations, we propose Adaptive Redundancy Regulation for Balanced Multimodal Information Refinement (RedReg), which is inspired by information bottleneck principle. Specifically, we construct a redundancy phase monitor that uses a joint criterion of effective gain growth rate and redundancy to trigger intervention only when redundancy is high. Furthermore, we design a co-information gating mechanism to estimate the contribution of the current dominant modality based on cross-modal semantics. When the task primarily relies on a single modality, the suppression term is automatically disabled to preserve modality-specific information. Finally, we project the gradient of the dominant modality onto the orthogonal complement of the joint multimodal gradient subspace and suppress the gradient according to redundancy. Experiments show that our method demonstrates superiority among current major methods in most scenarios. Ablation experiments verify the effectiveness of our method. The code is available at https://github.com/xia-zhe/RedReg.git

cs.LG

A Dual-Modulation Framework for RGB-T Crowd Counting via Spatially Modulated Attention and Adaptive Fusion

Accurate RGB-Thermal (RGB-T) crowd counting is crucial for public safety in challenging conditions. While recent Transformer-based methods excel at capturing global context, their inherent lack of spatial inductive bias causes attention to spread to irrelevant background regions, compromising crowd localization precision. Furthermore, effectively bridging the gap between these distinct modalities remains a major hurdle. To tackle this, we propose the Dual Modulation Framework, comprising two modules: Spatially Modulated Attention (SMA), which improves crowd localization by using a learnable Spatial Decay Mask to penalize attention between distant tokens and prevent focus from spreading to the background; and Adaptive Fusion Modulation (AFM), which implements a dynamic gating mechanism to prioritize the most reliable modality for adaptive cross-modal fusion. Extensive experiments on RGB-T crowd counting datasets demonstrate the superior performance of our method compared to previous works. Code available at https://github.com/Cht2924/RGBT-Crowd-Counting.

cs.CV

SFN-YOLO: Towards Free-Range Poultry Detection via Scale-aware Fusion Networks

Detecting and localizing poultry is essential for advancing smart poultry farming. Despite the progress of detection-centric methods, challenges persist in free-range settings due to multiscale targets, obstructions, and complex or dynamic backgrounds. To tackle these challenges, we introduce an innovative poultry detection approach named SFN-YOLO that utilizes scale-aware fusion. This approach combines detailed local features with broader global context to improve detection in intricate environments. Furthermore, we have developed a new expansive dataset (M-SCOPE) tailored for varied free-range conditions. Comprehensive experiments demonstrate our model achieves an mAP of 80.7% with just 7.2M parameters, which is 35.1% fewer than the benchmark, while retaining strong generalization capability across different domains. The efficient and real-time detection capabilities of SFN-YOLO support automated smart poultry farming.

cs.CV

Global Rice Multi-Class Segmentation Dataset (RiceSEG): A Comprehensive and Diverse High-Resolution RGB-Annotated Images for the Development and Benchmarking of Rice Segmentation Algorithms

Developing computer vision-based rice phenotyping techniques is crucial for precision field management and accelerating breeding, thereby continuously advancing rice production. Among phenotyping tasks, distinguishing image components is a key prerequisite for characterizing plant growth and development at the organ scale, enabling deeper insights into eco-physiological processes. However, due to the fine structure of rice organs and complex illumination within the canopy, this task remains highly challenging, underscoring the need for a high-quality training dataset. Such datasets are scarce, both due to a lack of large, representative collections of rice field images and the time-intensive nature of annotation. To address this gap, we established the first comprehensive multi-class rice semantic segmentation dataset, RiceSEG. We gathered nearly 50,000 high-resolution, ground-based images from five major rice-growing countries (China, Japan, India, the Philippines, and Tanzania), encompassing over 6,000 genotypes across all growth stages. From these original images, 3,078 representative samples were selected and annotated with six classes (background, green vegetation, senescent vegetation, panicle, weeds, and duckweed) to form the RiceSEG dataset. Notably, the sub-dataset from China spans all major genotypes and rice-growing environments from the northeast to the south. Both state-of-the-art convolutional neural networks and transformer-based semantic segmentation models were used as baselines. While these models perform reasonably well in segmenting background and green vegetation, they face difficulties during the reproductive stage, when canopy structures are more complex and multiple classes are involved. These findings highlight the importance of our dataset for developing specialized segmentation models for rice and other crops.

eess.IV

Stability and error analysis of pressure-correction scheme for the Navier-Stokes-Planck-Nernst-Poisson equations

In this paper, we propose and analyze first-order time-stepping pressure-correction projection scheme for the Navier-Stokes-Planck-Nernst-Poisson equations. By introducing a governing equation for the auxiliary variable through the ionic concentration equations, we reconstruct the original equations into an equivalent system and develop a first-order decoupled and linearized scheme. This scheme preserves non-negativity and mass conservation of the concentration components and is unconditionally energy stable. We derive the rigorous error estimates in the two dimensional case for the ionic concentrations, electric potential, velocity and pressure in the $L^2$- and $H^1$-norms. Numerical examples are presented to validate the proposed scheme.

math.NA

Solving High-dimensional Parametric Elliptic Equation Using Tensor Neural Network

In this paper, we introduce a tensor neural network based machine learning method for solving the elliptic partial differential equations with random coefficients in a bounded physical domain. With the help of tensor product structure, we can transform the high-dimensional integrations of tensor neural network functions to one-dimensional integrations which can be computed with the classical quadrature schemes with high accuracy. The complexity of its calculation can be reduced from the exponential scale to a polynomial scale. The corresponding machine learning method is designed for solving high-dimensional parametric elliptic equations. Some numerical examples are provided to validate the accuracy and efficiency of the proposed algorithms.

math.NA

Precision and accuracy of single-molecule FRET measurements - a worldwide benchmark study

Single-molecule Förster resonance energy transfer (smFRET) is increasingly being used to determine distances, structures, and dynamics of biomolecules in vitro and in vivo. However, generalized protocols and FRET standards ensuring both the reproducibility and accuracy of measuring FRET efficiencies are currently lacking. Here we report the results of a worldwide, comparative, blind study, in which 20 labs determined the FRET efficiencies of several dye-labeled DNA duplexes. Using a unified and straightforward method, we show that FRET efficiencies can be obtained with a standard deviation between $Δ$E = +-0.02 and +-0.05. We further suggest an experimental and computational procedure for converting FRET efficiencies into accurate distances. We discuss potential uncertainties in the experiment and the modelling. Our extensive quantitative assessment of intensity-based smFRET measurements and correction procedures serve as an essential step towards validation of distance networks with the ultimate aim to archive reliable structural models of biomolecular systems obtained by smFRET-based hybrid methods.

q-bio.QM

A Multigrid Method Based On Shifted-Inverse Power Technique for Eigenvalue Problems

A multigrid method is proposed in this paper to solve eigenvalue problems by the finite element method based on the shifted-inverse power iteration technique. With this scheme, solving eigenvalue problem is transformed to a series of nonsingular solutions of boundary value problems on multilevel meshes. Since replacing the difficult eigenvalue solving by the easier solution of boundary value problems, the multigrid way can improve the overall efficiency of the eigenvalue problem solving. Some numerical experiments are presented to validate the efficiency of this new method.

math.NA

Finite element exterior calculus for parabolic problems

In this paper, we consider the extension of the finite element exterior calculus from elliptic problems, in which the Hodge Laplacian is an appropriate model problem, to parabolic problems, for which we take the Hodge heat equation as our model problem. The numerical method we study is a Galerkin method based on a mixed variational formulation and using as subspaces the same spaces of finite element differential forms which are used for elliptic problems. We analyze both the semidiscrete and a fully-discrete numerical scheme.

math.NA