SearcharxivSearch

arXiv subjects

Yuanyuan Mao

Publications and source records attributed to Yuanyuan Mao.

6 recordsLinked to original sources

Metabook: A Mobile-to-Headset Pipeline for 3D Story Book Creation in Augmented Reality

The AR 3D book has shown significant potential in enhancing students' learning outcomes. However, the creation process of 3D books requires a significant investment of time, effort, and specialized skills. Thus, in this paper, we first conduct a three-day workshop investigating how AI can support the automated creation of 3D books. Informed by the design insights derived from the workshop, we developed Metabook, a system that enables even novice users to create 3D books from text automatically. To our knowledge, Metabook is the first system to offer end-to-end 3D book generation. A follow-up study with adult users indicates that Metabook enables inexperienced users to create 3D books, achieving reduced efforts and shortened preparation time. We subsequently recruited 22 children to examine the effects of AR 3D books on children's learning compared with paper-based books. The findings indicate that 3D books significantly enhance children's interest, improve memory retention, and reduce cognitive load, though no significant improvement was observed in comprehension. We conclude by discussing strategies for more effectively leveraging 3D books to support children's learning and offer practical recommendations for educators.

cs.HC

Pinning "Reflection" on the Agenda: Investigating Reflection in Human-LLM Co-Creation for Creative Coding

Large language models (LLMs) are increasingly integrated into creative coding, yet how users reflect, and how different co-creation conditions influence reflective behavior, remains underexplored. This study investigates situated, moment-to-moment reflection in creative coding under two prompting strategies: the entire task invocation (T1) and decomposed subtask invocation (T2), to examine their effects on reflective behavior. Our mixed-method results reveal three distinct reflection types and show that T2 encourages more frequent, strategic, and generative reflection, fostering diagnostic reasoning and goal redefinition. These findings offer insights into how LLM-based tools foster deeper creative engagement through structured, behaviorally grounded reflection support.

cs.HC

BDIQA: A New Dataset for Video Question Answering to Explore Cognitive Reasoning through Theory of Mind

As a foundational component of cognitive intelligence, theory of mind (ToM) can make AI more closely resemble human thought processes, thereby enhancing their interaction and collaboration with human. In particular, it can significantly improve a model's comprehension of videos in complex scenes. However, current video question answer (VideoQA) datasets focus on studying causal reasoning within events few of them genuinely incorporating human ToM. Consequently, there is a lack of development in ToM reasoning tasks within the area of VideoQA. This paper presents BDIQA, the first benchmark to explore the cognitive reasoning capabilities of VideoQA models in the context of ToM. BDIQA is inspired by the cognitive development of children's ToM and addresses the current deficiencies in machine ToM within datasets and tasks. Specifically, it offers tasks at two difficulty levels, assessing Belief, Desire and Intention (BDI) reasoning in both simple and complex scenarios. We conduct evaluations on several mainstream methods of VideoQA and diagnose their capabilities with zero shot, few shot and supervised learning. We find that the performance of pre-trained models on cognitive reasoning tasks remains unsatisfactory. To counter this challenge, we undertake thorough analysis and experimentation, ultimately presenting two guidelines to enhance cognitive reasoning derived from ablation analysis.

cs.MM

A Review on Machine Theory of Mind

Theory of Mind (ToM) is the ability to attribute mental states to others, the basis of human cognition. At present, there has been growing interest in the AI with cognitive abilities, for example in healthcare and the motoring industry. Beliefs, desires, and intentions are the early abilities of infants and the foundation of human cognitive ability, as well as for machine with ToM. In this paper, we review recent progress in machine ToM on beliefs, desires, and intentions. And we shall introduce the experiments, datasets and methods of machine ToM on these three aspects, summarize the development of different tasks and datasets in recent years, and compare well-behaved models in aspects of advantages, limitations and applicable conditions, hoping that this study can guide researchers to quickly keep up with latest trend in this field. Unlike other domains with a specific task and resolution framework, machine ToM lacks a unified instruction and a series of standard evaluation tasks, which make it difficult to formally compare the proposed models. We argue that, one method to address this difficulty is now to present a standard assessment criteria and dataset, better a large-scale dataset covered multiple aspects of ToM.

cs.AI

Structure of dimension-bounded temporal correlations

We analyze the structure of the space of temporal correlations generated by quantum systems. We show that the temporal correlation space under dimension constraints can be nonconvex. For the general case, we provide the necessary and sufficient dimension of a quantum system needed to generate a convex correlation space for a given scenario. We further prove that this dimension coincides with the dimension necessary to generate any point in the temporal correlation polytope. As an application of our results, we derive nonlinear inequalities to witness the nonconvexity for qubits and qutrits in the simplest scenario, and present an algorithm which can help to find the minimum for a certain type of nonlinear expressions under dimension constraints.

quant-ph

Geometry of faithful entanglement

A typical concept in quantum state analysis is based on the idea that states in the vicinity of some pure entangled state share the same properties; implying that states with a high fidelity must be entangled. States whose entanglement can be detected in this way are also called faithful. We prove a structural result on the corresponding fidelity-based entanglement witnesses, resulting in a simple condition for faithfulness of a two-party state. For the simplest case of two qubits faithfulness can directly be decided and for higher dimensions accurate analytical criteria are given. Finally, our results show that faithful entanglement is, in a certain sense, useful entanglement; moreover, they establish connections to computational complexity and simplify several results in entanglement theory.

quant-ph