SearcharxivSearch

arXiv subjects

Xiaohe Zhou

Publications and source records attributed to Xiaohe Zhou.

3 recordsLinked to original sources

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by reporting precise coordinates or discrete textual symbols. Yet existing spatial reasoning benchmarks usually require coordinates, options, or text, creating an answer-interface mismatch for image-generation models. This makes it difficult to evaluate image-generation models under the same task semantics as text-output VLMs, despite their ability to externalize spatial judgments directly in pixel space. We propose ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics. ProVisE also includes an Agentic builder that constructs and validates task-specific protocols for new benchmarks. We further introduce SpatialGen-Bench, a curated diagnostic benchmark of 470 samples across 14 spatial subtasks, four capability levels, and diverse answer forms. We evaluate representative text-output VLMs and image-generation models in a unified setting and validate Agentic protocol construction on six external spatial benchmarks. Results show that image-generation models are competitive when spatial answers can be externalized directly in pixel space, while text-output VLMs retain a clear advantage in compositional spatial reasoning. These findings reveal complementary strengths of pixel-space expression and text-based reasoning and establish a metric-compatible testbed for studying spatial cognition in image-generation models.

cs.CV

Harnessing self-sensitized scintillation by supramolecular engineering of CsPbBr3 nanocrystals in dense mesoporous template nanospheres

Perovskite-based nanoscintillators, such as CsPbBr3 nanocrystals (NCs), are emerging as promising candidates for ionizing radiation detection, thanks to their high emission efficiency, rapid response, and facile synthesis. However, their nanoscale dimensions - smaller than the mean free path of secondary carriers - and relatively low emitter density per unit volume, limited by their high molecular weight and reabsorption losses, restrict efficient secondary carrier conversion and hamper their practical deployment. In this work, we introduce a strategy to enhance scintillation performance by organizing NCs into densely packed domains within porous SiO2 mesospheres (MSNs). This engineered architecture achieves up to a 40-fold increase in radioluminescence intensity compared to colloidal NCs, driven by improved retention and conversion of secondary charges, as corroborated by electron release measurements. This approach offers a promising pathway toward developing next-generation nanoscintillators with enhanced performance, with potential applications in high-energy physics, medical imaging, and space technologies.

cond-mat.mtrl-sci

NavMarkAR: A Landmark-based Augmented Reality (AR) Wayfinding System for Enhancing Spatial Learning of Older Adults

Wayfinding in complex indoor environments is often challenging for older adults due to declines in navigational and spatial-cognition abilities. This paper introduces NavMarkAR, an augmented reality navigation system designed for smart-glasses to provide landmark-based guidance, aiming to enhance older adults' spatial navigation skills. This work addresses a significant gap in design research, with limited prior studies evaluating cognitive impacts of AR navigation systems. An initial usability test involved 6 participants, leading to prototype refinements, followed by a comprehensive study with 32 participants in a university setting. Results indicate improved wayfinding efficiency and cognitive map accuracy when using NavMarkAR. Future research will explore long-term cognitive skill retention with such navigational aids.

cs.HC