SearcharxivSearch

arXiv subjects

Dongyang Chen

Publications and source records attributed to Dongyang Chen.

14 recordsLinked to original sources

Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents

Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexterous visual tool use: fine-grained, closed-loop parameterized visual action in which models infer tool parameters from visual evidence, and those parameters directly govern the final result. Existing benchmarks cover web navigation, GUI operation, and software engineering, but rarely target this coupling between visual evidence and execution precision. We propose EASEL, a benchmark evaluating a controlled instance of dexterous visual tool use that adopts reference-guided visual reconstruction as its primary proxy task: the agent incrementally paints a canvas to match a reference image. EASEL additionally includes semantic tasks spanning region annotation, handwriting, and path planning. We further provide EASEL-Data, a 440k-sample two-stage curriculum dataset for trajectory supervision, and EASEL-9B to investigate its effect on this capability. Evaluation of 25 models reveals that current multimodal agents systematically struggle on EASEL. Reconstruction similarity bottlenecks at low levels (0.40-0.54), while trajectory diagnostics expose severe closed-loop instability - models typically saturate early or degrade post-peak. Semantic tasks reveal sharp capability boundaries in precision annotation and path planning. EASEL-9B, trained on EASEL-Data, surpasses the base model by a relative 6.3%, ranking third among all evaluated models.

cs.AI

V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval

Multimodal Large Language Models (MLLMs) have recently been applied to universal multimodal retrieval, where Chain-of-Thought (CoT) reasoning improves candidate reranking. However, existing approaches remain largely language-driven, relying on static visual encodings and lacking the ability to actively verify fine-grained visual evidence, which often leads to speculative reasoning in visually ambiguous cases. We propose V-Retrver, an evidence-driven retrieval framework that reformulates multimodal retrieval as an agentic reasoning process grounded in visual inspection. V-Retrver enables an MLLM to selectively acquire visual evidence during reasoning via external visual tools, performing a multimodal interleaved reasoning process that alternates between hypothesis generation and targeted visual verification.To train such an evidence-gathering retrieval agent, we adopt a curriculum-based learning strategy combining supervised reasoning activation, rejection-based refinement, and reinforcement learning with an evidence-aligned objective. Experiments across multiple multimodal retrieval benchmarks demonstrate consistent improvements in retrieval accuracy (with 23.0% improvements on average), perception-driven reasoning reliability, and generalization.

cs.CV

QuantEval: A Benchmark for Financial Quantitative Tasks in Large Language Models

Large Language Models (LLMs) have shown strong capabilities across many domains, yet their evaluation in financial quantitative tasks remains fragmented and mostly limited to knowledge-centric question answering. We introduce QuantEval, a benchmark that evaluates LLMs across three essential dimensions of quantitative finance: knowledge-based QA, quantitative mathematical reasoning, and quantitative strategy coding. Unlike prior financial benchmarks, QuantEval integrates a CTA-style backtesting framework that executes model-generated strategies and evaluates them using financial performance metrics, enabling a more realistic assessment of quantitative coding ability. We evaluate some state-of-the-art open-source and proprietary LLMs and observe substantial gaps to human experts, particularly in reasoning and strategy coding. Finally, we conduct large-scale supervised fine-tuning and reinforcement learning experiments on domain-aligned data, demonstrating consistent improvements. We hope QuantEval will facilitate research on LLMs' quantitative finance capabilities and accelerate their practical adoption in real-world trading workflows. We additionally release the full deterministic backtesting configuration (asset universe, cost model, and metric definitions) to ensure strict reproducibility.

cs.CL

AdaTooler-V: Adaptive Tool-Use for Images and Videos

Recent advances have shown that multimodal large language models (MLLMs) benefit from multimodal interleaved chain-of-thought (CoT) with vision tool interactions. However, existing open-source models often exhibit blind tool-use reasoning patterns, invoking vision tools even when they are unnecessary, which significantly increases inference overhead and degrades model performance. To this end, we propose AdaTooler-V, an MLLM that performs adaptive tool-use by determining whether a visual problem truly requires tools. First, we introduce AT-GRPO, a reinforcement learning algorithm that adaptively adjusts reward scales based on the Tool Benefit Score of each sample, encouraging the model to invoke tools only when they provide genuine improvements. Moreover, we construct two datasets to support training: AdaTooler-V-CoT-100k for SFT cold start and AdaTooler-V-300k for RL with verifiable rewards across single-image, multi-image, and video data. Experiments across twelve benchmarks demonstrate the strong reasoning capability of AdaTooler-V, outperforming existing methods in diverse visual reasoning tasks. Notably, AdaTooler-V-7B achieves an accuracy of 89.8\% on the high-resolution benchmark V*, surpassing the commercial proprietary model GPT-4o and Gemini 1.5 Pro. All code, models, and data are released.

cs.CV

FeatureFool: Zero-Query Fooling of Video Models via Feature Map

The vulnerability of deep neural networks (DNNs) has been preliminarily verified. Existing black-box adversarial attacks usually require multi-round interaction with the model and consume numerous queries, which is impractical in the real-world and hard to scale to recently emerged Video-LLMs. Moreover, no attack in the video domain directly leverages feature maps to shift the clean-video feature space. We therefore propose FeatureFool, a stealthy, video-domain, zero-query black-box attack that utilizes information extracted from a DNN to alter the feature space of clean videos. Unlike query-based methods that rely on iterative interaction, FeatureFool performs a zero-query attack by directly exploiting DNN-extracted information. This efficient approach is unprecedented in the video domain. Experiments show that FeatureFool achieves an attack success rate above 70\% against traditional video classifiers without any queries. Benefiting from the transferability of the feature map, it can also craft harmful content and bypass Video-LLM recognition. Additionally, adversarial videos generated by FeatureFool exhibit high quality in terms of SSIM, PSNR, and Temporal-Inconsistency, making the attack barely perceptible. This paper may contain violent or explicit content.

cs.CV

Large discrepancy between observations and simulations: Implications for urban air quality in China

Chemical transport models (CTMs) have been widely used to provide instructions for the control of ozone (O3) pollution. However, we find large discrepancies between observation- and model-based urban O3 chemical regimes: volatile organic compound (VOC)-limited regimes over N. China and weak nitrogen oxides (NOx)-limited regimes over S. China in observations, in contrast to simulations with widespread distributions of strong NOx-limited regimes. The conflicting O3 evolutions are caused by underestimated urban NOx concentrations and the possible overestimation of biogenic VOC emissions. Reductions in NOx emissions, in response to regulations, have thus led to an unintended deterioration of O3 pollution over N. China provinces, for example, an increase in surface O3 by approximately 7 ppb over the Sichuan Basin (SCB) in 2014-2020. The NOx-induced urban O3 changes resulted in an increase in premature mortality by approximately 3000 cases in 2015-2020.

physics.ao-ph

$(1+)$-complemented, $(1+)$-isomorphic copies of $L_{1}$ in dual Banach spaces

The present paper contributes to the ongoing programme of quantification of isomorphic Banach space theory focusing on Pe{\l}czy\'nski's classical work on dual Banach spaces containing $L_{1}$ ($=L_{1}[0,1]$) and the Hagler--Stegall characterisation of dual spaces containing complemented copies of $L_{1}$. We prove the following quantitative version of the Hagler--Stegall theorem asserting that for a Banach space $X$ the following statements are equivalent: $\bullet$ $X$ contains almost isometric copies of $(\bigoplus_{n=1}^{\infty} \ell_{\infty}^{n})_{\ell_1}$, $\bullet$ for all $\varepsilon>0$, $X^{*}$ contains a $(1+\varepsilon)$-complemented, $(1+\varepsilon)$-isomorphic copy of $L_{1}$, $\bullet$ for all $\varepsilon>0$, $X^{*}$ contains a $(1+\varepsilon)$-complemented, $(1+\varepsilon)$-isomorphic copy of $C[0,1]^{*}$. Moreover, if $X$ is separable, one may add the following assertion: $\bullet$ for all $\varepsilon>0$, there exists a $(1+\varepsilon)$-quotient map $T\colon X\rightarrow C(\Delta)$ so that $T^{*}[C(\Delta)^{*}]$ is $(1+\varepsilon)$-complemented in $X^{*}$, where $\Delta$ is the Cantor set.

math.FA

Quantifying shrinking and boundedly complete bases

We investigate possible quantifications of R. C. James' classical work on bases and reflexivity of Banach spaces. By introducing new quantities measuring how far a basic sequence is from being shrinking and/or boundedly complete, we prove quantitative versions of James' famous characterisations of reflexivity in terms of bases. Furthermore, we establish quantitative versions of James' characterisations of reflexivity of Banach spaces with unconditional bases.

math.FA

Quantifying properties ($K$) and ($\mu^{s}$)

A Banach space $X$ has \textit{property $(K)$}, whenever every weak* null sequence in the dual space admits a convex block subsequence $(f_{n})_{n=1}^\infty$ so that $\langle f_{n},x_{n}\rangle\to 0$ as $n\to \infty$ for every weakly null sequence $(x_{n})_{n=1}^\infty$ in $X$; $X$ has \textit{property $(\mu^{s})$} if every weak$^{*}$ null sequence in $X^{*}$ admits a subsequence so that all of its subsequences are Ces\`{a}ro convergent to $0$ with respect to the Mackey topology. Both property $(\mu^{s})$ and reflexivity (or even the Grothendieck property) imply property $(K)$. In the present paper we propose natural ways for quantifying the aforementioned properties in the spirit of recent results concerning other familiar properties of Banach spaces.

math.FA

Positively $p$-nuclear operators, positively $p$-integral operators and approximation properties

In the present paper, we introduce and investigate a new class of positively $p$-nuclear operators that are positive analogues of right $p$-nuclear operators. One of our main results establishes an identification of the dual space of positively $p$-nuclear operators with the class of positive $p$-majorizing operators that is a dual notion of positive $p$-summing operators. As applications, we prove the duality relationships between latticially $p$-nuclear operators introduced by O. I. Zhukova and positively $p$-nuclear operators. We also introduce a new concept of positively $p$-integral operators via positively $p$-nuclear operators and prove that the inclusion map from $L_{p^{*}}(\mu)$ to $L_{1}(\mu)$($\mu$ finite) is positively $p$-integral. New characterizations of latticially $p$-integral operators by O. I. Zhukova and positively $p$-integral operators are presented and used to prove that an operator is latticially $p$-integral (resp. positively $p$-integral) precisely when its second adjoint is. Finally, we describe the space of positively $p^{*}$-integral operators as the dual of the $\|\cdot\|_{\Upsilon_{p}}$-closure of the subspace of finite rank operators in the space of positive $p$-majorizing operators. Approximation properties, even positive approximation properties, are needed in establishing main identifications.

math.FA

Quantifications of strictly singular operators and strictly cosingular operators

We investigate possible quantifications of strictly singular operators, $l_{p}$-strictly singular operators, $c_{0}$-strictly singular operators, strictly cosingular operators, $l_{p}$-strictly cosingular operators. We prove quantitative, even strengthening versions of well-known results about relationships of these five classes of operators and compact, weakly compact, unconditionally converging operators.

math.FA

Unconditionally $p$-converging operators and Dunford-Pettis Property of order $p$

In the present paper we study unconditionally $p$-converging operators and Dunford-Pettis property of order $p$. New characterizations of unconditionally $p$-converging operators and Dunford-Pettis property of order $p$ are established. Six quantities are defined to measure how far an operator is from being unconditionally $p$-converging. We prove quantitative versions of relationships of completely continuous operators,unconditionally $p$-converging operators and unconditionally converging operators. We further investigate possible quantifications of the Dunford-Pettis property of order $p$.

math.FA

Pełczyński's property ($V^{*}$) of order $p$ and its quantification

We introduce the concepts of Pełczyński's property ($V$) of order $p$ and Pełczyński's property ($V^{*}$) of order $p$. It is proved that, for each $1<p<\infty$, the James $p$-space $J_{p}$ enjoys Pełczyński's property ($V^{*}$) of order $p$ and the James $p^{*}$-space $J_{p^{*}}$ (where $p^{*}$ denotes the conjugate number of $p$) enjoys Pełczyński's property ($V$) of order $p$. We prove that both $L_{1}(μ)$ ($μ$ a finite positive measure) and $l_{1}$ enjoy the quantitative version of Pełczyński's property ($V^{*}$).

math.FA

Perturbations of frames

In this paper, we give some sufficient conditions under which perturbations preserve Hilbert frames and near-Riesz bases. Similar results are also extended to frame sequences, Riesz sequences and Schauder frames. It is worth mentioning that some of our perturbation conditions are quite different from those used in the previous literatures on this topic.

math.FA