SearcharxivSearch

arXiv subjects

Reina Ishikawa

Publications and source records attributed to Reina Ishikawa.

7 recordsLinked to original sources

AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models

Vision-Language Models (VLMs) have demonstrated immense promise in Spatio-Temporal Video Grounding (STVG). However, current evaluation protocols are largely confined to zero-shot assessments on general, daily-life benchmarks. This creates a critical disconnect from real-world applications in specialized fields, where models inevitably encounter rare visual concepts and complex spatio-temporal dynamics. Since exhaustive pre-training across infinite data distributions is infeasible, the ability to adapt to novel domains is essential. To bridge this gap, we introduce AnyGroundBench, a domain-adaptation benchmark designed to shift the STVG evaluation paradigm from static zero-shot testing to rigorous domain adaptation. Targeting five specialized domains (animal, industry, sports, surgery, and public security), AnyGroundBench pairs newly captured videos such as expert-annotated mouse behaviors with established datasets, unifying them through dense, high-fidelity spatio-temporal annotations. Crucially, the benchmark provides dedicated training subsets to systematically measure domain adaptability. We extensively evaluate 15 state-of-the-art VLMs, assessing their zero-shot generalization and In-Context Learning (ICL) capabilities under practical computational constraints. Ultimately, our findings reveal that current models fail in both zero-shot and ICL-based adaptation when confronted with specialized domains, exposing critical flaws in spatio-temporal reasoning that future research must address.

cs.CV

Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation

Evaluating concept customization is challenging, as it requires a comprehensive assessment of fidelity to generative prompts and concept images. Moreover, evaluating multiple concepts is considerably more difficult than evaluating a single concept, as it demands detailed assessment not only for each individual concept but also for the interactions among concepts. While humans can intuitively assess generated images, existing metrics often provide either overly narrow or overly generalized evaluations, resulting in misalignment with human preference. To address this, we propose Decomposed GPT Score (D-GPTScore), a novel human-aligned evaluation method that decomposes evaluation criteria into finer aspects and incorporates aspect-wise assessments using Multimodal Large Language Model (MLLM). Additionally, we release Human Preference-Aligned Concept Customization Benchmark (CC-AlignBench), a benchmark dataset containing both single- and multi-concept tasks, enabling stage-wise evaluation across a wide difficulty range -- from individual actions to multi-person interactions. Our method significantly outperforms existing approaches on this benchmark, exhibiting higher correlation with human preferences. This work establishes a new standard for evaluating concept customization and highlights key challenges for future research. The benchmark and associated materials are available at https://github.com/ReinaIshikawa/D-GPTScore.

cs.CV

Harmonic Tutte polynomials of matroids II

In this work, we introduce the harmonic generalization of the $m$-tuple weight enumerators of codes over finite Frobenius rings. A harmonic version of the MacWilliams-type identity for $m$-tuple weight enumerators of codes over finite Frobenius ring is also given. Moreover, we define the demi-matroid analogue of well-known polynomials from matroid theory, namely Tutte polynomials and coboundary polynomials, and associate them with a harmonic function. We also prove the Greene-type identity relating these polynomials to the harmonic $m$-tuple weight enumerators of codes over finite Frobenius rings. As an application of this Greene-type identity, we provide a simple combinatorial proof of the MacWilliams-type identity for harmonic $m$-tuple weight enumerators over finite Frobenius rings. Finally, we provide the structure of the relative invariant spaces containing the harmonic $m$-tuple weight enumerators of self-dual codes over finite fields.

math.CO

Jacobi polynomials and design theory II

In this paper, we introduce some new polynomials associated to linear codes over $\mathbb{F}_{q}$. In particular, we introduce the notion of split complete Jacobi polynomials attached to multiple sets of coordinate places of a linear code over $\mathbb{F}_{q}$, and give the MacWilliams type identity for it. We also give the notion of generalized $q$-colored $t$-designs. As an application of the generalized $q$-colored $t$-designs, we derive a formula that obtains the split complete Jacobi polynomials of a linear code over $\mathbb{F}_{q}$.Moreover, we define the concept of colored packing (resp. covering) designs. Finally, we give some coding theoretical applications of the colored designs for Type~III and Type~IV codes.

math.CO

The CORSMAL benchmark for the prediction of the properties of containers

The contactless estimation of the weight of a container and the amount of its content manipulated by a person are key pre-requisites for safe human-to-robot handovers. However, opaqueness and transparencies of the container and the content, and variability of materials, shapes, and sizes, make this estimation difficult. In this paper, we present a range of methods and an open framework to benchmark acoustic and visual perception for the estimation of the capacity of a container, and the type, mass, and amount of its content. The framework includes a dataset, specific tasks and performance measures. We conduct an in-depth comparative analysis of methods that used this framework and audio-only or vision-only baselines designed from related works. Based on this analysis, we can conclude that audio-only and audio-visual classifiers are suitable for the estimation of the type and amount of the content using different types of convolutional neural networks, combined with either recurrent neural networks or a majority voting strategy, whereas computer vision methods are suitable to determine the capacity of the container using regression and geometric approaches. Classifying the content type and level using only audio achieves a weighted average F1-score up to 81% and 97%, respectively. Estimating the container capacity with vision-only approaches and estimating the filling mass with audio-visual multi-stage approaches reach up to 65% weighted average capacity and mass scores. These results show that there is still room for improvement on the design of new methods. These new methods can be ranked and compared on the individual leaderboards provided by our open framework.

cs.MM