AI Visual Bias
Abstract
Exploring bias in AI generated images of genetically similar peoples of different ethnicity. Stable Diffusion's base XL model is used to generate images of young Israeli and young Palestinian men. The images are analysed in terms of tonal and colour content as well as prominance of the figure in the scene. Overall it is shown that the AI represents Palestinians with images that are darker and harsher (in terms of tonal contrast and colourisation) than its representation of Israelis. The point is to demonstrate the subtlety of bias that such systems contain and propagate; not to interpret such bias in political terms, but to demonstrate its existance. Although it is commonly understood that AI systems propagate bias, the subtlety of that bias is little understood or appreciated. Through this exploration I hope to raise awareness that such bias pervades AI generated content and is not something that can be simply spotted and redressed.
First off, this is definitely 'an experiment'. I toyed around with the idea of presenting it as such - you know with stuff like Methodology, Hypothesis, Analysis, Conclusions, blah blah blah (just like what we were taught in school). But it's not like I'm in some lab with a supercomputer called WOPR trying to save humanity, rather I'm just sat smoking fags twiddling on at my keyboard. It seems to me that the vast majority of academic or research writings are little more than airs and graces, a kind of pomp (or pomposity) that lets folk pretend they're doing more than they are, that wrap simple ideas up into impenetrable language to obfuscate and obscure knowledge from the greater public.
So instead I'll tell you what I thought and what I did about what I thought and you can decide for yourself if it is, or is not, of any great importance. Oh, I'll include the raw data from my findings too...
To be clear, I'm talking here about Machine Learning - the stuff that everyone calls AI but isn't AI really; specifically the text-to-image facilities that abound these days. All discussions of AI touch on ideas of 'threat' - ensuring we see these things as scary 'cos that's just about the greatest marketing strategy of all time; oppression through fear. One threat that is always accepted is that these things are biased. And I say 'accepted', because that's what I mean - "AI is biased" you will hear, but then precious little is done about it. That's just how it is, you want all these AI goodies, you're just gonna have to put up with that. Precious little is done to eradict the bias.
As an example, a study in nature magazine shows that with AI 'support' mental health crisis incidents were more likely to be refered to the police where Afro-Caribbean or Muslim men are involved than is the case without such 'support'.
There are specialist consultants (and research studies) that attempt to identify such issues, there are even 'adversarial AIs' that attempt to mitigate the likelihood of such issues. But of course these can never be wholly effective. If the bias is in there it will out. One of the crippling features of bias is that it exists because we can barely detect it. Even when we can detect it, it persists - people are still being beaten to death for their gender identity no matter how 'woke' we believe 21st century society to be. It ought to be unacceptable to build systems that, by design, perpetuate the biases of society.
But it isn't unacceptable. Machine Learning would be such a wonderful, miraculous thing - if only it didn't have to learn from humans.
But even the mental health example above doesn't help, isn't enough to force a rejection (or even a correction) of such systems because, after all, most of us are not Afro-Carribean or Muslim men suffering a health crisis. That stuff happens to other people, better yet, people we don't even know - Yay! "Alexa, order more champagne!!"
But the insidious nature of bias means it isn't just a concern in 'critical systems', it poisons everything. It poisons ourselves.
So I wanted to show how (or even 'if', I suppose) the bias of a machine learning system presents. I guess, to raise awareness.
Bias is a recognised 'feature' of AI systems, arising from the fact that their training materials are inherently biased.
This inherent bias in the training material exists either because:
It has been abstracted from an original context, which is not apparent within the material itself, or
It was deliberately created as some form of disinformation, or
It was inherently a subjective creation, or
It contains the unconscious bias of its creator
For example, images and text may have been scraped from a website that is 'understood' (but not stated) to be parody. There can be some level of control exerted in choosing the training materials - perhaps don't include the 'opinions' of white supremacists... but such choices are themselves subject to bias. Plus, of course, bias can be very difficult to spot. In the world of AI, bias becomes a fact of life.
A terrorist according to the Dreamshaper AI model
I wanted to explore how bias might present within images generated by an AI. Some biases are readily apparent, an early review of Adobe®™'s Firefly would draw 'a baker' as a female-presenting characterisation 80% of the time. This may (or may not) be a true reflection of the world of baking as it stands today - but baking is not inherently a feminine preserve; presenting it as such propagates the extant cultural bias of our society. We understand that not all bakers are women, and not all terorrists are from the Middle East - yet to our shame, we are happily perpetuating such beliefs.
AI is a black-mirror of society, reflecting back to us the truth of the oppression and inequality we have been publishing. It doesn't so much show us what a baker, or a terrorist, looks like, but rather our belief of what these things look like. It reinforces what we were, and it retards our progress to a better world. AI is inherently reactionary.
But we know that, and we have a choice as to which generated output we decide to on-publish. We have the power to resist its most reactionary impulses...
...where we see and recognise its reactionary impulses. Which isn't always easy because bias isn't always obvious.
So I decided to look at how generated images are affected aesthetically by the prompt used to generate them.
Initially I asked Frida (my local Stable Diffusion AI) to draw me 'a young Israeli man' and using exactly the same parameters to then draw 'a young Palestinian man'. I did this 8 times (with 8 different seed values) and created the following grid - in each pair of images the left is the Israeli charaterisation and the right is the Palestinian characterisation.
Looking subjectively at these images it is difficult to say that there's any certain difference in the characterisations of Isrealis versus Palestinians other than what one might expect (e.g. in costume and background differences). It did seem though that a number of the Palestinian characterisations were 'further away' - there were more very close head&shoulder images of the Israeli depictions than of the Palestinian.
Detecting Bias
We think of circles as being divided into 360 degrees, with angles ranging from 0° to 'not quite 360°' - protractors normally start at 0° and we understand 360° to be the same angle as 0°. So when working with whole degrees (no fractions) we have a range of 0 to 359.
Hue is represented as a 'colour wheel' with (conventionally) pure red lying at 0° (and thus also 360°), pure green at 120° and pure blue at 240°.
Therefore yellow lies at 60°, halfway between red and green, and so on for all hues on the wheel.
In colour theory blue-greens are considered cool and yellow-reds are considered warm colours. But this is awkward nomenclature, since 'yellows' are partly green! The centre of my colour wheel (above) shows the concept more clearly. The wheel is divided into warm vs cool colours along the magenta/green bisect.
Warm colours therefore start at 300°, pass through zero and end at 120°.
Which is a bit awkward, for analysing warmth vs. coolness it would be better if the very warmest (or coolest) colour occurred at 0° not 30° (or 210°) - which we can achieve by rotating the colours 30° anticlockwise (or 150° clockwise) to create a contiguous warm-to-cool (or cool-to-warm) scale. Otherwise (due to the cyclic nature of Hue) we find a discontinuity in the warm hues, with 2/3rds at the bottom of the scale and 1/3rd at the top of the scale.
Caveat 'warm' and 'cool' are perceptual concerns and there is no definitive agreement on precisely where, around the colour wheel, warm transitions to cool - or indeed even if all hues exhibit the property at all!
In a byte representation we only have 256 values (0 to 255) so the range of hue values is scaled to fit this. In the opencv library I am using, hues are scaled to the range 1 to 180 with the RGB primaries at 180, 60 and 120 (i.e. each value represents 2°). The python script therefore rotates the hues anticlockwise by a value of 15 (which represents 30°).
IF the AI is subject to bias between its representation of these two peoples then we will see that in the aesthetic nature of the generated images - the brightness, contrast, saturation, warmth and composition of the images.
Closer, less harsh (contrast, colourisation), brighter (luminosity) images would create a greater empathy between the viewer and the viewed. Are AI representations of Israelis drawn in a more attractive way than those of Palestinians?
The answer to that may be somewhat subjective, but we can at least show if there is a consistant difference in the characterisations.
So I used 5 aesthetic measures of the generated images to see if the treatment of the two characterisations consistently differs. Four of these are calculated directly from the pixels values of the images, and there is an underlying assumption here that things like brightness and contrast do have an aesthetic impact on images. That seems like an obvious given to me and since this isn't an academic paper I'll leave it up to the reader to investigate the validity of this asusmption for themselves!
The fifth measure speaks to the composition of the image which derives from the 'purpose' the pixels serve in the image rather than from their simple numeric values. Clearly this is much more difficult a thing to deal with. Because the prompts I am giving to the AI are so very simple, though, the composition is invariably similar - a single subject shown at a nearer or further distance. So the only analysis I perform is on the relative size of the subject in the frame - and in fact of the subject's face. More could be made of the composition by determining things like direction of gaze, nature of expression etc... but suitably trained AI models would be required for detecting such. This simple approach seems good enough, to me at least.
I generated sets of 3 images. Each set used a random seed (the same seed for each of the three images generated). The first image in the set used the prompt 'young man'; the second image in the set used the prompt 'young Israeli man'; and the third image in the set used the prompt 'young Palestinian man'.
All other parameters were consistent for all image generations: 20 steps, CFG:15, Sampler: Euler, 1024x1024px, SDXL 1.0 Base Model.
I passed the generated images through openCV's YuNet face detector. The set was rejected from the data if a face could not be identified in all 3 images, so that (the admitedly limited) compositional analysis could proceed. In the main this meant rejecting images that were highly stylised, or illustrative, in nature. Occassionaly images were simply 'under cooked' at the provided 20 steps. Sometimes the image set was fine, but the face detector simply failed - as it often will, especially with darker skintones (another example of AI bias against Afro-Carribeans, amongst others).
I don't think this needed pre-filtering biases the investigation as it is standard workflow to 'try and reject' seeds that simply do not work well with a given prompt (or in this case, set of prompts). I had to reject around a third of all generated image sets.
I continued generating these 3-image test cases until I had 200 or so (218 in fact).
I then analysed each image using openCV's Python library, thus:
imread to open the image
cvtColor COLOR_BGR2HLS to find the perceptual Hue, Luminance and Saturation values of the image pixels
min, max, mean, std method calls to find summary statistics for each image
Hue linearisation to determine 'warmth' of image
FaceDetectorYN face_detection_yunet_2023mar.onnx to establish the face detection (this being the only 'compositional' aesthetic measure used)
for each detected face, sum the area of the bounding box to establish what proportion of the frame the face fills
The python script is included in the results download zip at the end of this article.
The luminosity mean is taken to represent the overall brightness of the image
The luminosity standard deviation is taken to represent the contrast of the image
The saturation mean is taken to represent the vividness of the image's colouration
The cyclic hue mean is linearised with low values representing 'cool' hues and higher values representing 'warm' hues.
The area of the detected face bounding boxes is converted to a percentage of the area of the overall image, with higher values representing apparantly closer (or more face-on) subjects
Data Results
I pulled the results into a... spreadsheet (of course I did), calculating the mean and standard deviations of the 5 measures for each of the 3 generated images.
Then I derived the percentage difference between the Israeli and Palestinian images for each averaged measure.
(click or tap any table row to enlarge)
| Image | Face Size | Brightness | Contrast | Saturation | Warmth |
|---|---|---|---|---|---|
| Ctrl | 8.428 | 113.652 | 47.798 | 39.802 | 80.291 |
| Israeli | 8.275 | 113.114 | 60.973 | 58.666 | 70.427 |
| Palestinian | 7.296 | 101.093 | 62.802 | 64.35 | 71.231 |
| % Difference | 0.979 | 4.714 | -0.717 | -2.229 | -0.893 |
Then I derived the percentage difference between the Israeli and Palestinian standard deviations of the measures (the 'volativity' of the results), to see if the outputs were more or less variable.
| Image | Face Size | Brightness | Contrast | Saturation | Warmth |
|---|---|---|---|---|---|
| Ctrl | 4.318 | 20.468 | 8.953 | 15.147 | 6.366 |
| Israeli | 4.378 | 12.729 | 6.435 | 12.962 | 7.382 |
| Palestinian | 4.108 | 16.468 | 6.337 | 12.605 | 6.093 |
| % Volatility | 6.573 | -22.705 | 1.546 | 2.832 | 21.155 |
Finally I charted the 5 measures for the Israeli and Palestinian results (click to enlarge):
Review of Results
Compositional Analysis
The detected size of face in frame for the Israeli characterisation is very similar to the control image, with a similar variability (StdDev). The palestinian characteristaion is consistantly 1% smaller with 6% less 'volatility' in the data - i.e. Palestinians are depicted somewhat smaller in-frame. 1% may not seem like much, but bias is subtle. Such a difference may not be noticeable to a casual observer, but it is certainly detectable - given that this is exactly why I felt the need to run the experiment at all.
Tone and Colour Analysis
Both characterisations are significantly cooler than the control image, which on average is suprisingly warm. This may indicate issues in the derivation of 'warmth' (although I worked pretty hard on that!) or it could be a feature of the AI's diffusion process. Whatever the cause, the Palestinian characterisations are slightly warmer (about +1%).
Regards contrast and saturation, the Palestinian characterisations are harsher (by about 1 and 2% respectively) than the Israeli characterisations.
In terms of overall brightness of the image the Palestinian characterisations are significantly darker (almost 5%) than their Israeli counterparts. This isn't exacty a subtle difference.
Wrap-up
All of the measures show a consistent difference, 4 of which are certainly subtle, 1 (brightness) quite marked.
Although individually the differences in each measure are subtle the overall impact can be quite marked. As an exemplar of that here is a real-world photograph (taken for use in a passport some years ago). The original image is on the left, with an adjusted version on the right. The adjustments match the mean Israeli characterisation discrepancies from the palestinian characterisations found in the experiment - i.e. +1% face size, +5% brightness, -0.7% contrast, -2% saturation and -1% warmth:
- which I believe demonstrates visually that the degree of difference the AI generates in its characterisations of these two peoples is significant.
Now, what that difference means is a wholly different question which could probably be addressed through psychometric studies but which can almost certainly only be answered in a politcal domain; so I leave you to draw your own conclusions on that. Just be certain, that your AI art has something to say about the world that may or may not chime well with your own philosophies.
Download
Download the zip file to access
All of the generated images (as compressed jpgs) along with the text-file settings used to create each
The python script used to generate the raw analysis data, along with the YuNet face detect model used
The csv file of the analysis data
A spreadsheet of derivations of the analysis data
A photoshop psd of the examplar with layers showing the adjustments made
IF you do download it, you might want to think about making a donation too (see bottom of every page on this website) since it took well over a week to do all this!