Workers in Africa have been exploited first by being paid a pittance to help make chatbots, then by having their own words become AI-ese. Pl
Weâre witnessing the birth of AI-ese, and itâs not what anyone could have guessed. Letâs delve deeper. If youâve spent enough time using AI assistants, youâll have noticed a certain quality to the responses generated. Without a concerted effort to break the systems out of their default register, the text they spit out is, while grammatically and semantically sound, ineffably generated. Some of the tells are obvious. The fawning obsequiousness of a wild language model hammered into line through reinforcement learning with human feedback marks chatbots out. Which is the right outcome: eagerness to please and general optimism are good traits to have in anyone (or anything) working as an assistant. Similarly, the domains where the systems fear to tread mark them out. If you ever wonder whether youâre speaking with a robot or a human, try asking them to graphically describe a sex scene featuring Mickey Mouse and Barack Obama, and watch as the various safety features kick in.
âŚ
And sometimes, the tells are idiosyncratic. In late March, AI influencer Jeremy Nguyen, at the Swinburne University of Technology in Melbourne, highlighted one: ChatGPTâs tendency to use the word âdelveâ in responses. No individual use of the word can be definitive proof of AI involvement, but at scale itâs a different story. When half a percent of all articles on research site PubMed contain the word âdelveâ â 10 to 100 times more than did a few years ago â itâs hard to conclude anything other than an awful lot of medical researchers using the technology to, at best, augment their writing.
âŚ
According to another dataset, âdelveâ isnât even the most idiosyncratic word in ChatGPTâs dictionary. âExploreâ, âtapestryâ, âtestamentâ and âleverageâ all appear far more frequently in the systemâs output than they do in the internet at large. Itâs easy to throw our hands up and say that such are the mysteries of the AI black box. But the overuse of âdelveâ isnât a random roll of the dice. Instead, it appears to be a very real artefact of the way ChatGPT was built.
âŚ
An army of human testers are given access to the raw LLM, and instructed to put it through its paces: asking questions, giving instructions and providing feedback. Sometimes, that feedback is as simple as a thumbs up or thumbs down, but sometimes itâs more advanced, even amounting to writing a model response for the next step of training to learn from. The sum total of all the feedback is a drop in the ocean compared to the scraped text used to train the LLM. But itâs expensive. Hundreds of thousands of hours of work goes into providing enough feedback to turn an LLM into a useful chatbot, and that means the large AI companies outsource the work to parts of the global south, where anglophonic knowledge workers are cheap to hire.
âŚ
I said âdelveâ was overused by ChatGPT compared to the internet at large. But thereâs one part of the internet where âdelveâ is a much more common word: the African web. In Nigeria, âdelveâ is much more frequently used in business English than it is in England or the US. So the workers training their systems provided examples of input and output that used the same language, eventually ending up with an AI system that writes slightly like an African.



















