arXiv · 2608.07243
Recipes for Creativity: Iterative Generation and Evaluation in Large Language Models
Abstract
Generative models are often evaluated through singular artifacts, whereas human creativity typically emerges through iterative generation, appraisal, and refinement. This pilot study examines whether iterative search improves LLM creativity by adapting FunSearch to recipe generation for the 2024 Pillsbury Bake-Off and evaluating outputs against human benchmarks using TTCT-based LLM evaluation. Across two experiments, we test iteration count, generator temperature, and in-loop selection-scorer model size. Results show that iterative generation-selection can produce recipes with creativity scores comparable to human benchmarks, but additional iterations alone do not improve creativity. The in-loop evaluator matters most: a smaller selection scorer yields significantly higher scores across most TTCT dimensions, while temperature has limited effects except for originality. These findings suggest that evaluator design is a first-order design variable in subjective creative search.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Rens Anderson, Tessa Verhoef, Amirhossein Zohrehvand. 2026-08-07. Recipes for Creativity: Iterative Generation and Evaluation in Large Language Models. https://arxiv.org/abs/2608.07243
Cite the original work for its findings. Save a collection to share your selection of sources.