Text-to-Image Ai Update
Prompt: "line drawing of ghostly hands typing on a computer keyboard, writing a blog post." (Generated Feb 2024; Dall-E)
When we began this research project last summer, we wanted to try out any publicly available (and usually free) AI tools to help us with the various aspects of conducting an experiment, producing a podcast, and writing a horror screenplay. One we ended up using often are text-to-image tools to create cover images for our episodes, social media posts, and more.
Our first use case - as you'll remember from one of our first blog posts - was to create our Mad Scary Podcast logo and image. Some takeaways there was that text-to-image tools (at that time, ~9 months ago) struggled with generating human features (like hands and faces), symbols (words, letters, numbers - even crucifixes on gravestones), and contextualizing certain figurative descriptors (like "eerie" or"spooky"). While some of the results were interesting or to our liking, we were new at writing prompts (specifically prompts for image generations vs text generation) and didn't actually have a clear idea of what we wanted.
Prompt: "full moon at night over new york city, spooky, in the style of a graphic novel." (Generated Jun 2023; Canva Magic Studio)
Eight months later, we're still struggling with some of the same things, but have learned (surprisingly) that less is more. Giving ChatGPT or Dall-E a more detailed prompt, rife with notes on style or tone may work sometimes, but I've noticed that they often ignore details that appear later in the prompt text. I suspect this is because the later details may "contradict" the earlier ones, or may seem redundant. For example, if you included "realistic lighting" and then "high contrast" later in the prompt - perhaps the LLM disregards the second lighting cue since it already has one? Or, maybe it's more simple than that; it doesn't understand what the ask is, or the punctuation used might be off.
Below are 2 images produced (this month, March 2024) from the same prompt, generated in Canva Magic Studio (available with a Canva pro subscription):
The prompt for both was the following: "2 ghost girls, sitting at a desk writing on paper or typing on a computer; one girl has long brown hair, one girl has medium length black hair; ambient lighting, realistic lighting."
The only difference I introduced was selecting from the list of possible themes available -- sort of like IG or Snapchat "filters." I chose the "Dreamy" theme for the image on the left, and "Oil Painting" theme for the one on the right. Though both are interesting to me, the image on the right seems to incorporate more of my prompt cues than the left (differences in hair color, both writing tools). Another interesting thing is that it ignores the "ghost" element, while the left image goes hard on it. It also makes me wonder if the data sets it pulls from for the "Oil Painting" theme includes more variety in terms of skin and hair color, objects (different iterations of a desk), maybe even race? If we were to reduce it down to the "idea" of a girl, or a ghost, or a desk -- like Plato's forms -- is there more/less variability for the output for one theme/style over another? Some questions to keep in mind in our continued experimentation with these tools, for sure!










