SearcharxivSearch

arXiv subjects

Xiaotao Huang

Publications and source records attributed to Xiaotao Huang.

11 recordsLinked to original sources

VIG: Visual Information Gain as a Reward Signal for Multimodal Chain-of-Thought Compression

Multimodal large reasoning models often rely on long Chain-of-Thought (CoT) traces in which a substantial fraction of tokens, such as repeated visual descriptions, self-reflection, and other visually-disengaged filler, inflate inference cost without contributing to the answer. Existing CoT compression methods optimize output length but never measure whether a reasoning token is actually grounded in the image. We propose \textbf{VIG} (Visual Information Gain), an information-theoretic GRPO reward that scores each reasoning token by how much the image reduces its predictive uncertainty. VIG is computed online from two forward passes of the same policy, one with and one without the image, so no reference chains, external annotations, or auxiliary reward models are needed. Across six main multimodal reasoning benchmarks and three Qwen3-VL-Thinking model sizes (2B/4B/8B), plus an additional R1-Onevision-Bench evaluation on 8B, VIG consistently improves the accuracy--efficiency trade-off, supporting our central claim: \emph{efficient multimodal reasoning emerges from raising visual information density, where every reasoning token earns its place by anchoring to the image, rather than from imposing a length budget.} Our source code is available at https://github.com/chaser682/vig.

cs.CV

ViTCoP: Accelerating Large Vision-Language Models via Visual and Textual Semantic Collaborative Pruning

Large Vision-Language Models (LVLMs) incur high computational costs due to significant redundancy in their visual tokens. To effectively reduce this cost, researchers have proposed various visual token pruning methods. However, existing methods are generally limited, either losing critical visual information prematurely due to pruning in the vision encoder, or leading to information redundancy among the selected tokens due to pruning in the Large Language Models (LLMs). To address these challenges, we propose a Visual and Textual Semantic Collaborative Pruning framework (ViTCoP) that combines redundancy filtering in the vision encoder with step-wise co-pruning within the LLM based on its hierarchical characteristics, to efficiently preserve critical and informationally diverse visual tokens. Meanwhile, to ensure compatibility with acceleration techniques like FlashAttention, we introduce the L2 norm of K-vectors as the token saliency metric in the LLM. Extensive experiments on various Large Vision-Language Models demonstrate that ViTCoP not only achieves state-of-the-art performance surpassing existing methods on both image and video understanding tasks, but also significantly reduces model inference latency and GPU memory consumption. Notably, its performance advantage over other methods becomes even more pronounced under extreme pruning rates.

cs.CV

EDENet: Echo Direction Encoding Network for Place Recognition Based on Ground Penetrating Radar

Ground penetrating radar (GPR) based localization has gained significant recognition in robotics due to its ability to detect stable subsurface features, offering advantages in environments where traditional sensors like cameras and LiDAR may struggle. However, existing methods are primarily focused on small-scale place recognition (PR), leaving the challenges of PR in large-scale maps unaddressed. These challenges include the inherent sparsity of underground features and the variability in underground dielectric constants, which complicate robust localization. In this work, we investigate the geometric relationship between GPR echo sequences and underground scenes, leveraging the robustness of directional features to inform our network design. We introduce learnable Gabor filters for the precise extraction of directional responses, coupled with a direction-aware attention mechanism for effective geometric encoding. To further enhance performance, we incorporate a shift-invariant unit and a multi-scale aggregation strategy to better accommodate variations in di-electric constants. Experiments conducted on public datasets demonstrate that our proposed EDENet not only surpasses existing solutions in terms of PR performance but also offers advantages in model size and computational efficiency.

cs.CV

Improvement of $q^2$ resolution in semileptonic decays based on machine learning

The neutrino closure method is often used to obtain kinematics of semileptonic decays with one unreconstructed particle. The kinematics of decays can be deducted by a two-fold ambiguity with a quadratic equation. To resolve the two-fold ambiguity, a new method based on Machine Learning (ML) is proposed. We study the effect of different sets of features and regressors on the improvement of reconstructed invariant mass squared of $\ell ν$ system~($q^2$). The result shows that the best performance is obtained by using the flight vector as the features, and the multilayer perceptron (MLP) model as the regressor. Compared with the random choice, the MLP model improves the resolution of reconstructed $q^2$ by $\sim$40\%. Furthermore, the possibility of using this method on various semileptonic decays is shown.

hep-ph

Angular asymmetries in $B\toΛ\bar p M$ decays

The forward-backward angular asymmetry (${\cal A}_{FB}$) for $\bar B^0\to Λ\bar pπ^+$ measured by Belle has presented an experimental value in the range of $-30\%$ to $-50\%$. In our study, we find that ${\cal A}_{FB}[\bar B^0\toΛ\bar p π^+(B^-\to Λ\bar pπ^0)]$ can be as large as $(-14.6^{+0.9}_{-1.5}\pm 6.9)\%$. In addition, we present ${\cal A}_{FB}[\bar B^0\toΛ\bar p ρ^+(B^-\to Λ\bar pρ^0)] =(4.1^{+2.8}_{-0.7}\pm 2.0)\%$ as the first prediction involving a vector meson in the charmless $B\to{\bf B\bar B'}M$ decays. While ${\cal A}_{FB}(B\toΛ\bar p M)$ indicates an angular correlation caused by the rarely studied baryonic form factors in the timelike region, LHCb and Belle~II are capable of performing experimental examinations.

hep-ph

Baryonic $B$ meson decays

We review the two and three-body baryonic $B$ decays with the dibaryon (${\bf B\bar B'}$) as the final states. Accordingly, we summarize the experimental data of the branching fractions, angular asymmetries, and $CP$ asymmetries. Using the $W$-boson annihilation (exchange) mechanism, the branching fractions of $B\to {\bf B \bf \bar B'}$ are shown to be interpretable. In the approach of perturbative QCD counting rules, we study the three-body decay channels. In particular, we review the $CP$ asymmetries of $B\to {\bf B\bar B'}M$, which are promising to be measured by the LHCb and Belle~II experiments. Finally, we remark the theoretical challenges in interpreting ${\cal B}(B^-\to p\bar pρ^-)$ and ${\cal B}(B^-\to p\bar pμ^-\bar ν_μ)$.

hep-ph

MIMO Radar Waveform-Filter Design for Extended Target Detection from a View of Games

This paper studies the Two-Person Zero Sum(TPZS) game between a Multiple-Input Multiple-Output(MIMO) radar and an extended target with payoff function being the output Signal-to-Interference-pulse-Noise Ratio(SINR) at the radar receiver. The radar player wants to maximize SINR by adjusting its transmit waveform and receive filter. Conversely, the target player wants to minimize SINR by changing its Target Impulse Response(TIR) from a scaled sphere centered around a certain TIR. The interaction between them forms a Stackelberg game where the radar player acts as a leader. The Stackelberg equilibrium strategy of radar, namely robust or minimax waveform-filter pair, for three different cases are taken into consideration. In the first case, Energy Constraint(EC) on transmit waveform is introduced, where we theoretically prove that the Stackelberg equilibrium is also the Nash equilibrium of the game, and propose Algorithm 1 to solve the optimal waveform-filter pair through convex optimization. Note that the EC can't meet the demands of radar transmitter due to high Peak Average to power Ratio(PAR) of the transmit waveform, thus Constant Modulus and Similarity Constraint(CM-SC) on waveform is considered in the second case, and Algorithm 2 is proposed to solve this problem, where we theoretically prove the existence of Nash equilibrium for its Semi-Definite Programming(SDP) relaxation form. And the optimal waveform-filter pair is solved by calculating the Nash equilibrium followed by the randomization schemes. In the third case,...

eess.SP

Robust MIMO Radar Waveform-Filter Design for Extended Target Detection in the Presence of Multipath

The existence of multipath brings extra "looks" of targets. This paper considers the extended target detection problem with a narrow band Multiple-Input Multiple-Output(MIMO) radar in the presence of multipath from the view of waveform-filter design. The goal is to maximize the worst-case Signal-to-Interference-pulse-Noise Ratio(SINR) at the receiver against the uncertainties of the target and multipath reflection coefficients. Moreover, a Constant Modulus Constraint(CMC) is imposed on the transmit waveform to meet the actual demands of radar. Two types of uncertainty sets are taken into consideration. One is the spherical uncertainty set. In this case, the max-min waveform-filter design problem belongs to the non-convex concave minimax problems, and the inner minimization problem is converted to a maximization problem based on Lagrange duality with the strong duality property. Then the optimal waveform is optimized with Semi-Definite Relaxation(SDR) and randomization schemes. Therefore, we call the optimization algorithm Duality Maximization Semi-Definite Relaxation(DMSDR). Additionally, we further study the case of annular uncertainty set which belongs to non-convex non-concave minimax problems. In order to address it, the SDR is utilized to approximate the inner minimization problem with a convex problem, then the inner minimization problem is reformulated as a maximization problem based on Lagrange duality. We resort to a sequential optimization procedure alternating between two SDR problems to optimize the covariance matrix of transmit waveform and receive filter, so we call the algorithm Duality Maximization Double Semi-Definite Relaxation(DMDSDR). The convergences of DMDSDR are proved theoretically. Finally, numerical results highlight the effectiveness and competitiveness of the proposed algorithms as well as the optimized waveform-filter pair.

eess.SP

Multilevel Image Thresholding Using a Fully Informed Cuckoo Search Algorithm

Though effective in the segmentation, conventional multilevel thresholding methods are computationally expensive as exhaustive search are used for optimal thresholds to optimize the objective functions. To overcome this problem, population-based metaheuristic algorithms are widely used to improve the searching capacity. In this paper, we improve a popular metaheuristic called cuckoo search using a ring topology based fully informed strategy. In this strategy, each individual in the population learns from its neighborhoods to improve the cooperation of the population and the learning efficiency. Best solution or best fitness value can be obtained from the initial random threshold values, whose quality is evaluated by the correlation function. Experimental results have been examined on various numbers of thresholds. The results demonstrate that the proposed algorithm is more accurate and efficient than other four popular methods.

cs.NE

A Sparse Learning Approach to the Design of Radar Tunable Architectures with Enhanced Selectivity Properties

This paper considers the design of tunable decision schemes capable of rejecting with high probability mismatched signals embedded in Gaussian interference with unknown covariance matrix. To this end, a sparse recovery technique is exploited to enhance the resolution at which the target angle of arrival is estimated with the objective to obtain high-selective detectors. The outcomes of this estimation procedure are used to devise detection architectures relying on either the twostage design paradigm or heuristic design procedures based upon the generalized likelihood ratio test. Remarkably, the new decision rules exhibit a bounded-constant false alarm rate property and allow for a tradeoff between the matched detection performance and the rejection of undesired signals by tuning a design parameter. At the analysis stage, the performance of the newly proposed detectors is assessed also in comparison with existing selective competitors. The results show that the new detectors can outperform the considered counterparts in terms of rejection of unwanted signals, while retaining reasonable detection performance of matched signals.

eess.SP

Novel Co-variant Feature Point Matching Based on Gaussian Mixture Model

The feature frame is a key idea of feature matching problem between two images. However, most of the traditional matching methods only simply employ the spatial location information (the coordinates), which ignores the shape and orientation information of the local feature. Such additional information can be obtained along with coordinates using general co-variant detectors such as DOG, Hessian, Harris-Affine and MSER. In this paper, we develop a novel method considering all the feature center position coordinates, the local feature shape and orientation information based on Gaussian Mixture Model for co-variant feature matching. We proposed three sub-versions in our method for solving the matching problem in different conditions: rigid, affine and non-rigid, respectively, which all optimized by expectation maximization algorithm. Due to the effective utilization of the additional shape and orientation information, the proposed model can significantly improve the performance in terms of convergence speed and recall. Besides, it is more robust to the outliers.

cs.CV