SearcharxivSearch

arXiv subjects

Harish Krishnakumar

Publications and source records attributed to Harish Krishnakumar.

2 recordsLinked to original sources

WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark

In real-world applications, models are expected to perform reliably across diverse settings. Yet, many existing multimodal benchmarks expand task types without capturing the visual diversity needed to handle open-ended visual inputs. We present WorldBench, a challenging and visually diverse reasoning benchmark to evaluate Multimodal Large Language Models (MLLMs). We build a taxonomy of thousands of visual concepts across multiple domains (e.g., living things). Guided by this taxonomy, we curate a broad collection of images from search engines and existing datasets to comprehensively represent the visual world. Through structured trial-and-error, we manually design challenging questions that frontier MLLMs fail to answer. On quantitative and human evaluations, WorldBench achieves higher visual diversity than any existing diverse benchmark. Evaluating 15 MLLMs on WorldBench reveals weaknesses in visual understanding: even the strongest model reaches only 64.0% accuracy, while some models perform marginally above chance-level. We hope our work highlights the importance of visual diversity in building multimodal benchmarks.

cs.CV

Analysis of Ring Galaxies Detected Using Deep Learning with Real and Simulated Data

Understanding the formation and evolution of ring galaxies, which possess an atypical ring-like structure, is crucial for advancing knowledge of black holes and galaxy dynamics. However, current catalogs of ring galaxies are limited, as manual analysis takes months to accumulate an appreciable sample of rings. This paper presents a convolutional neural network (CNN) to identify ring galaxies from unclassified samples. A CNN was trained on 100,000 simulated galaxies, transfer learned to a sample of real galaxies, and applied to a previously unclassified dataset to generate a catalog of rings which was then manually verified. Data augmentation with a generative adversarial network (GAN) to simulate images of galaxies was also employed. The resulting catalog contains 1967 ring galaxies. The properties of these galaxies were then estimated from their photometry and compared to the Galaxy Zoo 2 catalog of rings. However, the model's precision is currently limited due to a severe imbalance of rings in real datasets, leading to a significant false-positive rate of 41.1%, which poses challenges for large-scale application in surveys imaging billions of galaxies. This study demonstrates the potential of optimizing ML pipelines with low training data for rare morphologies and underscores the need for further refinements to enhance precision for extensive surveys like the Vera Rubin Observatory Legacy Survey of Space and Time.

astro-ph.GA