SearcharxivSearch

arXiv subjects

Hexu Liu

Publications and source records attributed to Hexu Liu.

4 recordsLinked to original sources

Commonsense Reasoning in Computer Vision: Foundations, Recent Advancements, and Future Directions

Commonsense reasoning in computer vision encompasses integrating visual data and contextual knowledge, crucial for enhancing AI's understanding of everyday scenarios. This understanding not only improves machine learning models but also enhances their ability to interact meaningfully with humans and the environment. Unlike CNN-based conventional vision models, which are designed to identify objects within a specific image, incorporating commonsense knowledge enables models to interpret scenes in a more holistic manner, thereby improving their spatial ability to reason about relationships among objects and actions. This integration not only enhances object recognition but also facilitates a deeper understanding of the contextual factors, ultimately leading to more precise predictions and interactions in real-world applications. This paper presents a comprehensive survey of recent developments that integrate commonsense knowledge into computer vision tasks. We systematically review approaches based on knowledge graphs, scene graphs, neuro-symbolic models, and commonsense-augmented transformers. We also outline current limitations related to dataset bias, knowledge incompleteness, and integration challenges. Finally, we highlight prospective research trajectories in cross-modal reasoning, scalable commonsense knowledge injection, and neuro-symbolic hybrid architectures to develop truly intelligent visual systems.

cs.AI

The Role of Instructional Guidance in Generative AI-Assisted Learning: Empirical Evidence from Construction Engineering Education

Generative artificial intelligence (AI) is increasingly used to support self-directed learning, yet student interaction with such systems often remains unstructured, limiting engagement in deeper cognitive processes. This study examines how instructional guidance shapes student and AI interaction in construction education. A five-step prompting framework grounded in Generative Learning Theory (GLT) is introduced to guide learner interaction during review activities. A controlled experiment compares three learning conditions: slide-based learning, unprompted AI-supported learning, and prompted AI-supported learning. Learning performance is assessed using multiple-choice and open-ended tasks, and user experience is measured using the User Experience Questionnaire (UEQ). Performance differences are concentrated on tasks requiring explanation and reasoning. The prompted condition achieves higher open-ended scores, with an improvement of approximately 2 or 3 points on a scale of 18 (p < 0.01), while no significant differences are observed in multiple-choice performance. The unprompted condition remains comparable to slide-based learning. These findings indicate that the effectiveness of AI-supported learning depends on how interaction is structured. The proposed framework provides a basis for integrating learning science principles into generative AI systems for construction education.

cs.HC

A lifting principle for canonical stability indices of varieties of general type

For any integer $n>0$, the $n$th canonical stability index $r_n$ is defined to be the smallest positive integer so that the $r_n$-canonical map $Φ_{r_n}$ is stably birational onto its image for all smooth projective $n$-folds of general type. We prove the lifting principle for $\{r_n\}$ as follows: $r_n$ equals to the maximum of the set of those canonical stability indices of smooth projective $(n+1)$-folds with sufficiently large canonical volumes. Equivalently, there exists a constant $\mathfrak{V}(n)>0$ such that, for any smooth projective $n$-fold $X$ with the canonical volume $\textrm{vol}(X)>{\mathfrak V}(n)$, the pluricanonical map $φ_{m,X}$ is birational onto the image for all $m\geq r_{n-1}$. The ''lifting principle'' was first put forward by James McKernan in Mathematics Review (MR2339333).

math.AG