Searcharxiv⌕ Search

arXiv subjects

Md Tanvir Hassan

Publications and source records attributed to Md Tanvir Hassan.

2 recordsLinked to original sources

How Semantically Stable Are LLM Refusals? Measuring Confusion in Local Safety Boundaries

As safety alignment becomes standard in large language models, refusal behavior has become an important part of model reliability. However, models may still reject benign prompts, especially when the wording resembles risky content. Existing evaluations usually report global scores, such as false rejection rate or compliance rate. These scores are useful, but they treat each prompt independently. As a result, they miss local inconsistency, where a model accepts one phrasing of an intent but rejects a close paraphrase. This makes it difficult to understand whether refusals are only frequent or also semantically unstable. We address this gap by introducing Semantic Confusion, a failure mode that captures contradictory refusal decisions across meaning-preserving paraphrases. We build ParaGuard, a 10k-prompt corpus of controlled paraphrase clusters that keep intent fixed while varying surface form. We also propose three model-agnostic token-level metrics: Confusion Index, Confusion Rate, and Confusion Depth. These metrics compare each rejected prompt with its nearest accepted neighbors using token embeddings, next-token probabilities, and perplexity signals. Experiments across diverse model families show that global false rejection rate can hide important structure in the refusal boundary. Our metrics reveal cases where confusion is spread broadly, cases where it appears only in specific semantic regions, and cases where stricter refusal does not lead to more local inconsistency. These findings show that refusal evaluation should measure not only how often a model refuses, but also how consistently it refuses across nearby paraphrases. This gives developers a practical signal for reducing false refusals while preserving safety.

cs.CL↗

MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models

Recent advancements in foundation models have improved autonomous tool usage and reasoning, but their capabilities in map-based reasoning remain underexplored. To address this, we introduce MapEval, a benchmark designed to assess foundation models across three distinct tasks - textual, API-based, and visual reasoning - through 700 multiple-choice questions spanning 180 cities and 54 countries, covering spatial relationships, navigation, travel planning, and real-world map interactions. Unlike prior benchmarks that focus on simple location queries, MapEval requires models to handle long-context reasoning, API interactions, and visual map analysis, making it the most comprehensive evaluation framework for geospatial AI. On evaluation of 30 foundation models, including Claude-3.5-Sonnet, GPT-4o, and Gemini-1.5-Pro, none surpass 67% accuracy, with open-source models performing significantly worse and all models lagging over 20% behind human performance. These results expose critical gaps in spatial inference, as models struggle with distances, directions, route planning, and place-specific reasoning, highlighting the need for better geospatial AI to bridge the gap between foundation models and real-world navigation. All the resources are available at: https://mapeval.github.io/.

cs.CL↗