SearcharxivSearch

arXiv subjects

Haobo Cheng

Publications and source records attributed to Haobo Cheng.

2 recordsLinked to original sources

MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook

This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We hope it better allows researchers to follow the state-of-the-art in this very dynamic area. Meanwhile, a growing number of testbeds have boosted the evolution of general-purpose large language models. Thus, this year's MARS2 focuses on real-world and specialized scenarios to broaden the multimodal reasoning applications of MLLMs. Our organizing team released two tailored datasets Lens and AdsQA as test sets, which support general reasoning in 12 daily scenarios and domain-specific reasoning in advertisement videos, respectively. We evaluated 40+ baselines that include both generalist MLLMs and task-specific models, and opened up three competition tracks, i.e., Visual Grounding in Real-world Scenarios (VG-RS), Visual Question Answering with Spatial Awareness (VQA-SA), and Visual Reasoning in Creative Advertisement Videos (VR-Ads). Finally, 76 teams from the renowned academic and industrial institutions have registered and 40+ valid submissions (out of 1200+) have been included in our ranking lists. Our datasets, code sets (40+ baselines and 15+ participants' methods), and rankings are publicly available on the MARS2 workshop website and our GitHub organization page https://github.com/mars2workshop/, where our updates and announcements of upcoming events will be continuously provided.

cs.CV

Bayesian Analysis of Multivariate Matched Proportions with Sparse Response

Multivariate matched proportions (MMP) data appears in a variety of contexts including post-market surveillance of adverse events in pharmaceuticals, disease classification, and agreement between care providers. It consists of multiple sets of paired binary measurements taken on the same subject. While recent work proposes non-Bayesian methods to address the complexities of MMP data, the issue of sparse response, where no or very few "yes" responses are recorded for one or more sets, is unaddressed. The presence of sparse response sets results in underestimates of variance, loss of coverage, and lowered power in existing methods. Bayesian methods have not previously been considered for MMP data but provide a useful framework when sparse responses are present. In particular, the Bayesian probit model provides an elegant solution to the problem of variance underestimation. We examine three approaches built on that model: a naive analysis with flat priors, a penalized analysis using half-Cauchy priors on the mean model variances, and a multivariate analysis with a Bayesian functional principal component analysis (FPCA) to model the latent covariance. We show that the multivariate analysis performs well on MMP data with sparse responses and outperforms existing non-Bayesian methods. In a re-analysis of data from a study of the system of care (SOC) framework for children with mental and behavioral disorders, we are able to provide a more complete picture of the relationships in the data. Our analysis provides additional insights into the functioning on the SOC that a previous univariate analysis missed.

stat.ME