I once conducted a simple thought experiment on epidemics and globalisation: if the COVID-19 pandemic had replaced the outbreak of the Bubonic Plague—otherwise known as the Black Death—in the 14th century, would it have become a simultaneous, planetary event of the kind the world experienced in 2020? Of course, the circumstances were vastly different in both advantageous and disadvantageous ways. People had much less knowledge of basic hygiene or medical practices, but populations were far smaller and more sparse. Population mobility was also significantly lower, as air travel did not exist and maritime transport was more modest. While long-distance movement existed at the time, such as overland caravans, Black Sea commerce, and Mediterranean shipping, it was slow and could not have crossed vast oceans. Thus, I believe it is not hard to imagine that the virus simply would not have spread far enough to have been a global epidemic1. Everyday life was far more localised.
The Black Death probably started in Central Asia around 1338–39; from there it moved west along trans-Asian trade networks to the Black Sea, and later into the Mediterranean. It then tore through Europe, the Middle East, and North Africa. However, it never crossed the oceans that still closed the Americas off from Afro-Eurasia, nor did it ever reach the Australian continent—though it did eventually spread centuries later during the Third Plague Pandemic. By the same token, the COVID-19 pandemic proved uniquely devastating not only because the virus spread rapidly, but also because the world had globalised to levels unprecedented in human history. Without these two factors, it would not have landed everywhere at once, nor would it have disrupted life on every continent as it did a few years ago.
I believe this thought experiment is analogous to what we are currently seeing with AI, big data, and globalisation. In the same way that 21st-century mobility turned a contagious virus into a worldwide catastrophe, modern hyper-connectivity ensures that whoever commands the large language models and the data that powers them wields unprecedented, borderless influence over collective human thought and behaviour.
Whoever controls the media, controls the mind.
— Jim Morrison (1969)2
Put simply, Large Language Models (LLMs)—commonly sold under the trade names of OpenAI’s GPT, Anthropic’s Claude, and Google’s Gemini models—are statistical models that map a given text input to a probability distribution of possible continuing sequences. LLMs are able to estimate this probability by ’learning’ from their training corpus, i.e., human-generated content, which can range from traditional media such as books and film all the way to casual social media posts.
As with most human-derived data, these corpora are inherently full of historical biases, ideological skew, cultural blind spots, and deliberate misinformation. Take the example of social media posts, which comprise an important part of AI training data. Group polarisation tells us that collective deliberation among like-minded individuals drives its members toward more radical stances. On modern platforms engineered to maximise user engagement, this creates echo chambers that disproportionately amplify volatile and sensational content. Consequently, datasets scraped from these environments are saturated with radical material that is ill-suited for training language models. Lastly, the Western-skewed nature of the internet means that training corpora are also culturally biased. Over 90% of the training data in models like GPT-3 was English3, drawing heavily from web scrapes like Common Crawl. As a result, the models internalise and reinforce Western socio-political norms and prejudices as universal truths.
Even if we discount the inherent, unintentional biases that come with user data, the training corpora are frequently proprietary and undisclosed. This means that the companies using this data are able to curate and align it with agendas that have nothing to do with making the model smarter. They can do this because there is a lack of checks-and-balances. State actors are aware of this and recognise that data control equates to ideological leverage. For example, China recently released a comprehensive implementation plan to transform itself into a leading data supplier for AI training, proposing the construction of “high-quality datasets in industries.” While this positions state-backed brokers to supply the growing global demand for training data, it also ensures that datasets are rigorously filtered and aligned with state-sanctioned narratives. In an unregulated data market, controlling the raw training data that AI is trained on is no longer a simple matter of creating the ‘smartest’ models possible, but a high-stakes geopolitical race to take control of global thought.
There are currently active efforts in mitigating the language models’ inherent biases, most notably Reinforcement Learning from Human Feedback (RLHF)4. This is a technique used during the LLM post-training pipeline, after the base model has been trained. Human annotators are shown a prompt together with several candidate responses generated by the model, and are asked to rank them from best to worst. These rankings are then used to train a separate ‘reward model’ — essentially a stand-in for human judgement that learns to predict which responses people will prefer. Finally, the language model is fine-tuned through reinforcement learning to maximise the score it receives from this reward model. It is expected that, in this way, human evaluators can indirectly ‘punish’ biased or harmful behaviour present in the system and ‘reward’ it for producing answers that people find helpful and appropriate. In practice, this shapes the model into an assistant that is helpful, polite, and neutral in tone, and acts as a key safeguard to proactively prevent the model from generating radical and prejudiced outputs. For example, if a user prompts an unaligned base model with racially inflammatory remarks or demands for actionable instructions on illegal acts, a model refined through RLHF will typically decline to comply and instead offer a polite, de-escalating alternative. While this creates the appearance of a well-calibrated moral compass, the process in itself also relies on several problematic assumptions.
RLHF relies on the assumption that a general, universally agreed-upon definition of neutrality exists, as well as on knowing what to do when it does not. While definitions of what might count as ‘fair’ and ’non-discriminatory’ might be relatively consistent across people from the same race, religion, and community, they diverge sharply across the diverse global population using these tools. For instance, prompting frontier image generation models to depict a company’s CEO will most likely result in an image depicting a white, male executive in a business suit. Conversely, companies have also tried to overcompensate for systemic biases by artificially forcing diversity. Google’s image models infamously generated historically inaccurate figures of the past, such as racially diverse German soldiers from 1943 and non-white American Founding Fathers. When alignment attempts to correct real-world demographic skews, it often swings from replicating historical prejudice to engineering an artificial consensus. RLHF merely decides whose worldview gets to define “neutrality.”
Alignment efforts assume the AI company’s and the government’s good faith. While companies are generally incentivised to make sure that alignment is done to mitigate biases as much as possible to avoid potential lawsuits, the same mechanism can easily be weaponised as well. For example, under political or state pressure, mechanisms such as RLHF can cease to be a simple, objective safety filter and become an ideological steering wheel. Governments utilise fiscal policies such as mass subsidies under the pretence of advancing local markets. These state-backed AI companies can then—with their massive funding—ensure that they will have an edge in price-to-performance ratio over their unsubsidised peers. However, this comes with a caveat in that it indirectly allows state actors to have economic leverage over them, as they would normally be their largest source of capital. In the case of AI, this would take the form of key architectural AI decisions, one that includes training data and the post-training pipeline. The reality of this is already visible: querying popular Chinese models like DeepSeek or Qwen about politically sensitive topics such as the Tiananmen Square Massacre will result in a generic “Sorry, that’s beyond my current scope. Let’s talk about something else.”
Daniel Kahneman defines System 2 in his book Thinking, Fast and Slow as the slow, analytical part of the mind responsible for human agency and critical thinking. In theory, this mechanism is the internal, built-in system to filter radical thoughts and false premises. However, the very nature of LLMs and their use cases act as a cognitive immunosuppressant. Because System 2 is inherently “lazy” and governed by the law of least effort, the polished, instantaneous fluency of an LLM satisfies our immediate search for answers without demanding much thinking. While transformative inventions of the past, from the printing press to the mechanical calculator, exponentially boosted performance and speed of executing specific tasks, LLMs represent a universal paradigm: a single technology applied across nearly every task, delivering fully packaged, end-to-end cognitive work from prompt to conclusion. The aggressive integration of LLM technology into everyday tools, such as Google’s AI Overviews, routinely exposes users to synthetic text, subtly yet inevitably engineering trust in its outputs, even if that trust is not deserved. This misplaced trust builds complacency, nudging users to default to System 1—the brain’s rapid, intuitive, and subconscious mode of thinking. With the lack of critical evaluation, users stop interrogating the information they receive and simply treat these potentially malicious and problematic outputs as a universal ground truth.
Modern society also heavily rewards those who use LLMs over those who do not. AI technology is marketed under the promise that it can produce passing undergraduate coursework, screen through thousands of job applicants’ CVs instantaneously, and as the model behind credit-scoring decisions5. It is undeniable that AI has made a lot of work easier and more accurate, from automated data extraction and entry processes to lowering the barrier to entry for learning and acquiring new knowledge. However, it is these same merits that also make it such a tempting yet lethal weapon. In today’s capitalist, hyper-competitive landscape, LLMs offer a decisive productivity advantage to those who use them over those who do not, making such exponential gains almost impossible to resist. For instance, in a study conducted by Microsoft, a knowledge worker noted that they needed to use AI to “reach a certain quota daily or risk losing my job.” And even if they do have the time, the same study also showed that there are knowledge barriers required to verify and refine AI outputs. Hence, while the decision to adopt LLMs appears voluntary, that agency can become a little more than an illusion. When an entire industry shifts toward an accelerated output, opting out is no longer a matter of personal preference, but a decision to forfeit a competitive advantage—one that society actively recognises and rewards across many fields.
Ultimately, an LLM’s training pipeline is reflected in the model’s output—responses that are then delivered to millions of users daily. Systemic economic pressures and cognitive laziness then trap individuals into relying on these tools, nudging us to absorb algorithmic conclusions without question. The inherent biases and misinformation that these LLMs carry become the underlying ‘pathogen,’ and the ubiquitous integration of LLMs in daily life becomes the perfect, efficient vector through which this contagion is codified, automated, and scaled across the globe.
This is not to minimise the lethality of either pandemic. Both were profoundly devastating in their own way. The Black Death wiped out over 25 million people in Europe alone—about a third of the continent’s population. While COVID-19 recorded far more total infections (almost 800 million confirmed cases), the fatality rate was much lower, and occurred against a modern global population nearly twenty times larger than that of 14th-century Afro-Eurasia. The main point of the experiment is thus to compare the speed and vectors of transmission of the disease, and not the impact of it. ↩︎
The quote is most commonly attributed to Jim Morrison, but recent analysis has shown that there is a lack of primary reference material to prove this. Similar ideas along the same lines have been said in the past, such as the quote “Control of the media of communication and information means the control of the mind” by U.S. Congressman Francis E. Walter of Pennsylvania, during his speech on safeguarding the media against communism. ↩︎
GPT-3 is a rather outdated example. This is because frontier labs treat their exact training dataset as proprietary trade secrets. More recent studies have shown that modern LLMs perform reasonably on dominant global languages (e.g., Chinese, Hindi, and Arabic), but still perform considerably worse on low-resource languages. The problem therefore is still the same—it creates severe disparities in highly multilingual nations such as Indonesia and Nigeria, where linguistic diversity is fragmented across dozens of regional languages, and English is also not their primary language spoken in daily life. ↩︎
RLHF is chosen as the primary example because it pioneered the entire family of preference-tuning techniques. It has since been succeeded by more lightweight, stable algorithms such as Direct Preference Optimisation (DPO), which eliminates the need for a separate reward model and unstable RL loops by optimising the policy directly, and Kahneman-Tversky Optimisation (KTO), which aligns models directly from binary feedback rather than costly paired comparisons. Despite these algorithmic efficiencies, the key problem mentioned earlier remains the same: it still relies on maximising human preference ‘scores’ as the primary objective, which inherits the same tendencies towards partiality, sycophancy, and superficial agreeableness. ↩︎
Although credit-scoring algorithms in the past utilise more ’traditional’ ML techniques, there have been studies and proposals on the feasibility of integrating LLMs as part of a more ‘hybrid’ strategy in determining credit default risk. ↩︎