Skip to content

LLM: Synonym for Bias

Most people do not know how many of the technologies they use work. I do not have specific numbers, but I’d guess that if I asked people how pictures and sound get on their TVs or mobile devices they’d be hard pressed to come up with any kind of technical knowledge of how that happens. I’m also guessing it’s true for things we’ve had for quite some time now, like plumbing, electricity, or even how planes fly. The same is true for generative AI applications like Large Language Models.

By now, we’re all familiar with GPT. It’s right in the name of the most popular AI application. Some may even be aware that generative pre-trained transformers spit out text based on their training to sound like natural speech. My concern is that they are not aware of how that very process is essentially a bias engine. It has to be. If it’s going to sound like us then it has to sound like us.



We know that as more and more people use these tools for assistants in legal, medical, educational, marketing and political fields and more that they will soon end up doing the bulk of the content generation. Right now, many of the obvious responses and errors are silly and rightly called out for ridicule. Soon, the output will achieve the “good enough” stage where it just becomes part of the background noise we access.

My concerns about this in the world of education are well-documented here, but we have to consider the damage this causes in ways beyond how we educate future generations. If judges are using AI to assist with sentencing guidelines and the training these tools have have patterns of racial bias then we will see that reflected in the suggestions (particularly if the information that is entered about those offenders contains information about names, race, socio-economic status, geography, etc.)

It happens in medicine as well. We know that people from marginalized communities and women have worse healthcare outcomes. If LLM are being employed to make diagnoses, then they are likely to perpetuate those worse outcomes. In education, many African American students on IEPs have their race identified within the first two lines of their assessment. Those same students are also disproportionately labeled as having emotional behavioral disorders when white students with the same type of assessment receive the diagnosis of having a learning disability. If LLM are used to assist students with special needs (which has the potential to be an incredible use of them), are we accounting for those discrepancies in the prompting and reviewing the output from those tools for potential biases?

Because most people do not really know how LLM work, they may not know that it requires continual pushback in prompting and response review to determine where some of those implicit biases may have worked their way in. We trained these tools on all of our (even if well-intentioned) insights, research, practices, media, reporting, our everything! We coded into these tools the directive to sound like us. Even as we evolve as a society, we are going to start sounding more and more like LLM than they will start sounding like us because that’s what humans do, they adapt to their surroundings.

Like I said, LLM are at their very core, bias engines because they have to be if they are going to produce human-style writing. Making sure that users understand the complexities of how responses are generated and strategies to mitigate the impact of bias in those results is going to be an essential skill in this new world of AI integration into all aspects of our lives.

Tags: