Searcharxiv⌕ Search

arXiv subjects

Mingwei S. G. Li

Publications and source records attributed to Mingwei S. G. Li.

2 recordsLinked to original sources

A Design Space for Visual Interfaces for Generative Image Models

Interactive visual interfaces have become an important means of controlling generative image models, enabling users to manipulate generation through prompts, direct manipulation, and a range of interactions. However, existing techniques are typically presented as independent systems, making it difficult to understand how they relate, compare their interaction mechanisms, or identify opportunities for new interface designs. We introduce a design space for interactive visual interfaces for generative image models derived from 51 research systems and practitioner tools. The framework decomposes each system into three complementary components: the user interface (U), the controllable model objects (Z), and the mapping function ($ϕ$) that translates user interaction into model operations. This decomposition provides a common representation for analyzing heterogeneous interaction techniques across model families, revealing recurring design patterns and underexplored regions of the design space. We further present an interactive corpus explorer that support comparative analysis, and discuss usage scenarios for both educational settings and HCI/AI practitioners identifying research and design opportunities.

cs.HC↗

LatentGandr: Visual Exploration of Generative AI Latent Space via Local Embeddings

Generative AI has demonstrated significant potential in creative design, enabling the rapid generation of visual content and imaginative concepts. Although deep AI models achieve effective featurization in the latent space, navigating the space remains a challenge. Current techniques, such as GANSlider and SliderSpace, use multiple sliders to generate high-dimensional vectors in generative AI's latent space. Despite applying (global) PCA to reduce the number of sliders, these approaches struggle with scalability and usability as the number of control dimensions increases. In this paper, we introduce LatentGandr, a visual analytics technique that facilitates latent space exploration by extracting locally linear dimensions from embeddings in high-dimensional latent spaces. By analyzing the topology and local curvature of the embeddings, LatentGandr automatically identifies local neighborhoods and computes their principal components using localized PCA. These local principal components are visualized as interactive image grids, allowing users to efficiently explore and control the generative process, providing an intuitive means to refine the generation of novel content and concepts. To evaluate the effectiveness of LatentGandr, we conducted a study comparing it to GANSlider, the current state-of-the-art visualization interface for generative AI models. The results offer insights into how localized exploration techniques can enhance user interaction with these models.

cs.HC↗