×

Some text in the Modal..

Ant Smith
Creativity is an action not an attribute... Photographic Skill: The Book - now available HERE

Artificial Intelligence

Paint Along With Frida #3: Creative Expression | Artificial Intelligence | Articles | Ant Smith

Paint Along With Frida #3: Creative Expression

Stable Diffusion plays a little game of tussel between where the random noise from the seed wants to go versus the demands of the prompt. If the noise happens to contain a lot of red and you are asking for a banana the results will likely be somewhat screwy. The CFG arbitrates between these two sides, lower values favour the noise direction, higher values favour the prompt - but the game is not always harmonious; which is why people create images in batches, looking for something where the seed and prompt work well together.

Oft times people write prompts in full flowing sentences, some swear by this approach. It can be interesting and at times I will throw a full stanza from a poem at Frida just to see what she comes up with; and when I'm just looking to illustrate some other creative endeavour it can work fine - i.e. if I don't have much of a preconception of what it is I want to see, exactly.

But if you have a picture in mind of what you are trying to create, quite often less can be more

Less IS More

The AI esentially uses the prompt to match against the descriptions it found of the images it has been trained on. That might be the associated text on the website where it was found (e.g. the ALT text) or it could be that the image was tagged specifically for the training. So the prompt isn't about reasoning with the AI (which has neither reason nor understanding) about what it should draw. The prompt is trying to match up with how other people might have described things similar to what you are looking to create. And in fact the prompt will only work if other people have described images in that way.

You are also limited in how much you can ask for. Technically there is no hard limit, but the prompt gets converted to 'tokens' which are sent for processing in batches of 77. The AI pays progressively less attention to the tokens as you ramble on. If you chatter away and cause 2 or more batches of tokens to be processed, things can get very confusing for the AI. There's always the temptation to add a little more to a prompt when things aren't quite going your way, but in so doing you're increasing the 'tussel' that the AI is going through to try and create what you want.

So, you don't need to write in full sentences, and it can end up causing a lot of pain if you do so! Elaborating rarely helps when talking to an AI, stripping the prompt back to its bare essentials (perhaps substituting some words for synonymns) is a much better strategy.

Consider these two prompts:

  • a nice bright yellow banana sitting on a countertop in the midday sun

  • banana, countertop, midday sun

which gave rise to these images:

Two images of bananas from different prompts

The longer prompt gave us a bit more detail in the scene (a kettle in the background), which we may or may not like. The shorter prompt was just as effective in regards of what we wanted, and we have lots of 'token space' remaining if we were to decide to add more details (e.g. banana, countertop, kettle, midday sun)

With shorter prompts there is also less chance of creating unintentional contradictions, which you might not even spot when typing some flowery prose but which will further aggravate the whole image generation for the AI.

The AI gives most attention to the first word of the prompt, and will fight hardest to ensure that concept appears in the image. Attention wanes as the prompt progresses. For this reason its a good idea to structure your prompt so that the most important concerns are listed first.

A common pattern for prompting is: Subject, Scene, Style

E.g. Woman, Rice Field, Hokusai

Image from prompt: Woman, Rice Field, Hokusai

I don't often cite individual artists because, you know, who's art is it anyway? But that's not really a judgemental stance and if you want to lean on the styles of known artists you can find a list of them here: https://stablediffusion.fr/artists

This is a very simple prompt format that's easy to remember, and which takes account of the AI's diminishing attention - i.e. subject is most important and comes first.

But each of the 3 parts would typically be much more complex, and then the distinction between subject and scene, or between scene and style, becomes harder to draw. So I can offer a much fuller format that covers all concerns you might want to describe. I still split the prompt into 3: The scene (including subject), The techniques to be emulated, and The styles to be applied. For each of these I include a list of things to be considered. The full prompt structure looks like this:

  • Mise-en-scène

    • subject

    • physical attributes

    • emotional traits

    • environment

    • place

    • setting

    • time

    • lighting

    • tone

    • mood

    • patterns, symmetry

  • Technique

    • medium

    • depth-of-field

    • point-of-view

    • engle of view

    • resolution

  • Style

    • genre

    • artist references

    • meaning

    • reception

Here's a video that gives some examples of this prompt structure in use:

First of all,notice how I don't feel the need to offer something for every category; I still aim for the simplest prompt possible. It should be clear that this format works very well in terms of delivering images that are close to what I was aiming for.

Some of the concepts (mood, meaning, etc...) are pretty abstract but their effect can be seen in the female construction worker example. The final, firefighter, example also shows how this scheme is every bit as effective as more traditional approaches to prompts - whilst avoiding that 'battle of wills' feeling that can come from working with AIs!

There are lots of weird terms in the firefighter prompt: hyperdetailed, cel-shaded, precisionism, 8k resolution, etc... People try stuff out (whilst fighting with their AIs) and come to believe such terms are useful. Whether they are or not isn't at all clear! '8k resolution' doesn't instruct the AI to create that kind of detail - what it does is to instruct the AI to consider its training arising from images that were so tagged. And if you look at such images you will see that they are mostly just charts and logos, not at all helpful in getting more detail into you image! This link: https://rom1504.github.io/clip-retrieval/ lets you search for images matching a given term, so you can test how useful such a term might be (you may want to turn safe off and set a high aesthetic weight if using this tool). These terms proliferate because people copy prompts without thinking. You can ask the AI to consider what it learnt from images tagged '8K resolution', but that will not cause the AI to draw an image at 8k resolution!

There are three key summary points:

  • Start with a simple, rough cut, prompt and generate many images to find a seed that works well with it.

  • Don't get into a fight with the AI, forcing extra elaborations down its throat; be precise and measured with your prompts.

  • Don't blindly copy other people's prompts, they will be full of garbage!

Negative Prompting

The AI progressively refines random noise (from the seed) into pixel visualisations of concepts that it has learnt - i.e. it doesn't understand concepts but it has learnt how concepts are visualised. Without knowing what 'a bucket' is, it has learnt what 'a bucket' looks like. If the pixels emerging from the refinement of the noise start to look like 'a bucket' then that is what the AI will end up drawing. Which is all well and good if you happened to ask for a bucket...

...but, the refinement is an iterative process. A bunch or emerging pixels may at first look as though they might end up being a bucket, or could equally well look as though they might end up being a hat. If the prompt mentions 'hat', then the AI will progress more towards that than towards 'bucket'...

...but, if the emerging pixels don't look like they may represent anything in the prompt then something seemingly random will be added to the image (like maybe a kettle in the background).

The one thing the AI must do is fill the canvas. If there are patches of emerging pixels that just don't relate to anything in the prompt then something will be drawn there; which could well be something you just don't want in the scene. So, using a negative prompt, you can tell the AI expressly not to draw such things. Looking again at the prompt 'a nice bright yellow banana sitting on a countertop in the midday sun' I can supress the AI's desire to draw a kettle in the background by adding 'kettle' to the negative prompt:

First image of bananas with kettle suppressed

And it looks like that patch of pixels that were going to be a kettle have now become a pepper mill, much better!

This is a very effective use of negative prompting - generate an image and then with all the same settings add unwanted elements to the negative prompt before regenerating.

But you need to add the visual concepts ('kettle') using terminology the AI has learnt. That's not too difficult with concrete concepts like 'kettle' (or 'watermark' if you like), but more abstract ideas can also be suppressed; like 'blurry' or 'grainy. It is very likely that some of the images in the AI's training data were blurry or grainy, so it has probably been trained to recognise such characteristics; allowing such results to be suppressed by the negative prompt.

You often also find people adding things like 'extra limbs' to the negative prompt because that is a common failing of the AIs - but it seems to me unlikely that the AI will have been trained to recognise 'extra limbs'. The raw images used in the training typically wouldn't have extra limbs so the probability the AI would recognise such a 'visual concept' seems pretty low. It ought to recognise the visual concept 'limbs', but extra limbs? I doubt it.

Just like prompts, negative prompts get shared and copied all the time. 'Extra limbs' probably seems like a good thing to try and suppress, so people stick it in the negative prompt just because they've seen it before - irrespective of whether the AI was trying to draw such or not in this case. And irrespective of whether it's doing any good or not!

Here's the default negative prompt used on the NightCafe online service:

ugly, tiling, poorly drawn hands, poorly drawn feet, poorly drawn face, out of frame, extra limbs, disfigured, deformed, body out of frame, blurry, bad anatomy, blurred, watermark, grainy, signature, cut off, draft

First of all, it seems unlikely that the AI has been trained to recognise a number of these visual concepts. How many images do we think it will have been trained on that were tagged 'poorly drawn feet'? As an experiment I asked for 'a man's bare foot' (at cfg 7) and then I regenerated with the negative prompt 'badly drawn feet'. Then, I regenerated the original (without the negative prompt) at a higher CFG setting (12):

comparisson of negative prompt versus CFG to improve image

The original image has typical bad feet with too many toes. The prompt was really simple and not enough to fill the canvas, so the AI decided to draw 2 feet standing on sand (I can imagine it has seen a lot of barefeet on beaches). The negative prompt has moved us closer to having 'a foot' (rather than feet) but it has still drawn that foot rather badly (insufficient toes). The third image shows we can more successfully improve how the feet are drawn by increasing the CFG, rather than trying to suppress a visual concept the AI probably doesn't know ('poorly drawn feet').

It is possible to add to the AI's training so that negative prompts like 'badly drawn hands' do some good using textual inversions which I'll look at in a later article. This works by explicitly showing the AI what badly drawn hands look like; but, 'out of the box' such negative prompts simply do not work.

But secondly, and more importantly, look again at that default prompt. There are terms like 'ugly', 'deformed', 'disfigured'. Again, people stick these in the negative prompt and then they get shared around, because they think they are telling the AI not to create disfigured images (since with low CFG or step count the AI often will). But the prompts do not tell the AI how to draw, only what to draw. By putting 'ugly' in the negative prompt you're deflecting the AI from considering the training that has arisen from images other people have tagged 'ugly' - so it won't attempt to draw 'ugly' people, but it will still draw all the 'beautiful' people in an ugly manner if the CFG or step count are too low. By adding 'ugly' to the negative prompt all you are doing is restricting the diversity that the AI could draw upon. This is a bad idea.

The key summary points for negative prompting are:

  • See what gets drawn first, then add any unwanted visual concepts to the negative prompt.

  • Remember you cannot prompt the AI on how to draw, only on things it should or shouldn't draw.

  • Mangled limbs are better fixed by adjusting CFG, Step count or Seed than by adding to the negative prompt.

  • Diversity matters, do not restrict the AI's attention from so called 'ugly', disfigured or deformed things.


More in Artificial Intelligence

Artificial Intelligence