SearcharxivSearch

arXiv subjects

Wenhui Hu

Publications and source records attributed to Wenhui Hu.

4 recordsLinked to original sources

Code-Vision: Evaluating Multimodal LLMs Logic Understanding and Code Generation Capabilities

This paper introduces Code-Vision, a benchmark designed to evaluate the logical understanding and code generation capabilities of Multimodal Large Language Models (MLLMs). It challenges MLLMs to generate a correct program that fulfills specific functionality requirements based on a given flowchart, which visually represents the desired algorithm or process. Code-Vision comprises three subsets: HumanEval-V, Algorithm, and MATH, which evaluate MLLMs' coding abilities across basic programming, algorithmic, and mathematical problem-solving domains. Our experiments evaluate 12 MLLMs on Code-Vision. Experimental results demonstrate that there is a large performance difference between proprietary and open-source models. On Hard problems, GPT-4o can achieve 79.3% pass@1, but the best open-source model only achieves 15%. Further experiments reveal that Code-Vision can pose unique challenges compared to other multimodal reasoning benchmarks MMCode and MathVista. We also explore the reason for the poor performance of the open-source models. All data and codes are available at https://github.com/wanghanbinpanda/CodeVision.

cs.CL

Correlation of the L-mode density limit with edge collisionality

The "density limit" is one of the fundamental bounds on tokamak operating space, and is commonly estimated via the empirical Greenwald scaling. This limit has garnered renewed interest in recent years as it has become clear that ITER and many tokamak pilot plant concepts must operate near or above the Greenwald limit to achieve their objectives. Evidence has also grown that the Greenwald scaling - in its remarkable simplicity - may not capture the full complexity of the density limit. In this study, we assemble a multi-machine database to quantify the effectiveness of the Greenwald limit as a predictor of the L-mode density limit and compare it with data-driven approaches. We find that a boundary in the plasma edge involving dimensionless collisionality and pressure, $\nu_{*\rm, edge}^{\rm limit} = 3.5 \beta_{T,{\rm edge}}^{-0.40}$, achieves significantly higher accuracy (false positive rate of 2.3% at a true positive rate of 95%) of predicting density limit disruptions than the Greenwald limit (false positive rate of 13.4% at a true positive rate of 95%) across a multi-machine dataset including metal- and carbon-wall tokamaks (AUG, C-Mod, DIII-D, and TCV). This two-parameter boundary succeeds at predicting L-mode density limits by robustly identifying the radiative state preceding the terminal MHD instability. This boundary can be applied for density limit avoidance in current devices and in ITER, where it can be measured and responded to in real time.

physics.plasm-ph

Bridging Modality Gap for Visual Grounding with Effecitve Cross-modal Distillation

Visual grounding aims to align visual information of specific regions of images with corresponding natural language expressions. Current visual grounding methods leverage pre-trained visual and language backbones independently to obtain visual features and linguistic features. Although these two types of features are then fused through elaborately designed networks, the heterogeneity of the features renders them unsuitable for multi-modal reasoning. This problem arises from the domain gap between the single-modal pre-training backbones used in current visual grounding methods, which can hardly be bridged by the traditional end-to-end training method. To alleviate this, our work proposes an Empowering Pre-trained Model for Visual Grounding (EpmVG) framework, which distills a multimodal pre-trained model to guide the visual grounding task. EpmVG relies on a novel cross-modal distillation mechanism that can effectively introduce the consistency information of images and texts from the pre-trained model, reducing the domain gap in the backbone networks, and thereby improving the performance of the model in the visual grounding task. Extensive experiments have been conducted on five conventionally used datasets, and the results demonstrate that our method achieves better performance than state-of-the-art methods.

cs.CV

Amplitude modulated Bloch oscillations of photon probability distribution in a cavity-atom system

We study the dynamics of the Rabi Hamiltonian in the medium coupling regime with $\left\vert g/ω\right\vert \sim 0.07$, where $g$ is atom-field coupling constant, $ω$ is the field frequency, for the quantum state with average photon number $\bar{n}\sim 10^{4}$. We map the original Hamiltonian to an effective one, which describes a tight-binding chain subjected to a staggered linear potential. It is shown that the photon probability distribution of a Gaussian-type state exhibits the amplitude modulated Bloch oscillation (BO), which is a superposition of two conventional BOs with a half-BO-period delay between them and is essentially another type of Bloch-Zener oscillation. The probability transition between the two BOs can be controlled and suppressed by the ratio $g\sqrt{\bar{n}}% /ω$, as well as in-phase resonant oscillating atomic frequency $Ω\left( t\right) $, leading to multiple zero-transition points.

quant-ph