Module 5 - AI Risk
Quick Access to Introduction to AI Risk | Bias in Automated Decision-Making Systems | AI Value Alignment | Self-Test
Instructed by Prof. Adam Bradley & Prof. Andre Curtis-Trudel
Introduction to AI Risk
Increasingly, critical decisions about our lives are made by advanced AI systems. These include decisions about finance (will my loan be approved?), criminal justice (should a suspect make bail?), health (does this X-ray indicate lung cancer?), and mass communication (which news stories should be promoted?). Soon, autonomous AI systems may take direct control of vital military systems, sensitive parts of the energy grid, and risky biotechnology projects. In performing these tasks, AI systems often operate in ways that even human experts cannot understand or explain. Developing these AI systems and using them to perform these important tasks poses several serious risks. In this module, we discuss some of the most serious risks from a philosophical perspective. This module begins with a gentle introduction to contemporary AI systems for non-experts. Then, it examines AI risk by looking in depth at two topics.
1. Automated bias: how can we ensure that decisions made with AI systems are fair?
2. The value alignment problem: how can we ensure that AI systems act in accordance with human values?

General introduction to AI Risk
Learning Outcomes
This component of the module will equip students to:
- Identify the main kinds of AI systems and techniques
- Explain the typical development cycle of contemporary AI systems.
- Explain some main commercial, legal, governmental, medical, and scientific uses of AI.
At the broadest level, Artificial Intelligence (AI) is the project of designing systems which exhibit human-level performance on certain tasks – tasks which, if performed by a human, would be deemed to require ‘intelligence’. Interest in AI largely began in the 1950s, with the construction of the first digital computers. Although primitive by today’s standards, researchers immediately appreciated the enormous power of these machines to solve complex problems. Today, roughly seventy years later, AI systems exhibit human-level or super-human level performance on a wide variety of tasks. They can play chess, Go, pass the bar exam, write essays, generate images, and much, much more.
Although these developments are exciting and often beneficial, they also present humanity with new challenges. What if someone uses an AI system to spread political misinformation? What if a military AI decides to target civilians or, worse, use nuclear weapons? And could a superintelligent AI somehow take control and enslave humanity? These are just a few of the potential risks raised by AI technologies.
This module will help you better understand these risks. In it, we discuss a variety of specific risks raised by the widespread use of AI, as well as potential future risks arising from this usage. You do not need any background knowledge of AI to complete this module. In the rest of this introduction, we provide a brief overview of the relevant aspects of AI. If you feel comfortable with AI, you should feel free to skip this video. But if you’re new to AI, it may help you to understand the potential risks AI poses. In the rest of this module, we discuss specific AI risks in more detail.

Bias in Automated Decision-Making Systems
Learning Outcomes This component will equip students to:
1. Identify some main sources of bias in automated decision-making systems.
2. Explain some main risks associated with biased decision-making systems.
3. Evaluate different ways of responding to automated bias from a variety of moral, legal, and epistemological perspectives.

Introducing Bias in Automated Decision-Making Systems
What is automated bias, and where does it come from?
Automated decision-making systems, often driven by advanced AI technology, are now used to obtain predictions in a variety of domains: predictive policing technologies are used to determine bail and sentencing guidelines, based on an individual’s predicted recidivism risk (Lum and Isaac 2016; ProPublica 2016); mortgage and loan applications increasingly rely on machine learning to determine an individual’s loan-worthiness (Singh et al. 2022); medical decisions making is increasingly aided by black-box AI systems (London 2019); and high-stakes military decisions are often informed by automated decision-making systems.
One common concern about this development is that automated decision-making systems may be biased in certain respects. But what does this mean?
Definition: A system is biased when its behavior systematically deviates from some relevant standard. This standard can be moral, legal, epistemological, etc. (Epistemology is the branch of philosophy concerned with knowledge and justification: what can you know, and what is your justification for it?)
For example, a moral bias could exist when a system systematically deviates from some moral standard, while a statistical bias could exist when a system systematically deviates from statistical regularities observed in training data. Different kinds of bias can overlap, so that systems can be biased in some ways but not others, and can exhibit multiple different kinds of biases at once.
In general, automated decision-making systems have three main parts (diagram 1). They take some input (diagram 2), use sophisticated algorithms to process that data (diagram 3), and produce a decision about that input as an output (diagram 4). For example, an insurance company might use AI to determine whether to grant someone a loan. The system might take information about their salary, current job, and credit history, and their loan application, and output ‘yes’ or ‘no’ depending on whether they get the loan or not.
Where do biases in automated decision-making systems come from? There are three main sources of bias in automated decision-making systems
- Input-side bias concerns biases in the data used by an automated system. This includes both training data and data encountered when the system is used in the real world. Bias can arise here because of the way data is produced or measured.
- Processing bias concerns biases in the way the algorithm or decision-making procedure used by a system.
- Output-side bias concerns the outputs depending on the way the outputs are used or deployed, especially when a system is used outside its intended use-case or outside of ordinary contexts
You might be tempted to think that all automated bias is bad. But the situation is more complicated than this. Sometimes we can use one kind of bias to correct for another. In particular, because different biases can exist in a system at once, we can use good biases to compensate for harmful or problematic biases. For example, a credit card company may bias their algorithms to increase the probability that minority groups receive credit opportunities that they would otherwise not be eligible for, given historically disadvantaged circumstances such groups have faced. In this case, the use of one kind of bias helps limit or prevent a potential moral harm.
When considering the biases present in an automated-decision system, we should ask two questions: (1) where in the system do biases occur (input-side, processing, or output-side?), and (2) how do these biases overlap and relate to each other? Considering both of these aspects will help us arrive at a holistic understanding of a system’s benefits and harms.

We are currently living through an era of rapid development in AI technology. Just a few years ago, the idea of AI systems coding, writing poetry, and performing standardized tests at human levels was almost unthinkable. Now, with the release of programs like GPT-4, it is reality. But there is no reason to think the capabilities of AI systems will stop at their current levels. Future—even near-future—AI systems may far exceed the abilities of anything we currently know of. With the development of advanced AI systems, however, we face new risks. In this section, we discuss one type of risk which arises due to the value alignment problem, or the problem of ensuring that AI systems act in accordance with widely shared human values.
The development of advanced AI poses many types of risk. Some of these risks may be very serious, but not catastrophic. But others could be catastrophic. By catastrophic risk, we mean a risk that could threaten the lives or well-being of a significant proportion of humanity. So, for instance, there is a risk that millions of knowledge workers in advanced economies will lose their jobs due to AI. This is a very serious risk, but it is not necessarily catastrophic. On the other hand, the risk that AI systems in control of nuclear weapons might start a worldwide nuclear war is a catastrophic risk. So is the risk that AI undertakes some action which destroys advanced industrial civilization, throwing us back into the Stone Age.

Why might an AI system undertake an action with such catastrophic consequences? One potential reason is that its values do not align with ours. By values we mean the ends or goals that something pursues. So, for example, common human values include things like physical comfort, safety, meaningful relationships, etc. Humans also value the means to these ends, things like money or status, which can be used to attain these valuable goods. Value alignment occurs when two or more agents share the same values. So, for example, human beings in general tend to share many values, even as they disagree about others. Most human beings value the continuation of the human species, though they may disagree about what direction that future should take. But in other cases, our values may fail to align. This can create conflicts. At a fundamental level, wars are waged because of failures of value alignment: the two countries value different things and see military conflict as the only means of resolving the disagreement.
An influential group of AI researchers is concerned with the catastrophic risks posed by AI due to the value alignment problem. The value alignment problem is the technical problem of ensuring that future AI systems act in accordance with human values, broadly understood. The concern is that in the near-to-mid-term future, we may develop very powerful AI systems that do not share human values. This includes the possibility of superintelligent AI systems that outperform human beings on essentially every task. The worry is that these advanced AI systems may act in ways that are contrary to human values, with potentially catastrophic results. In the following video, I will explain the contours of this problem by discussing a well-known thought experiment: the paperclip maximizer, a thought experiment discussed by Nick Bostrom and Eliezier Yudkowsky (Bostrom 2003)
Solving the Alignment Problem
The obvious proposal for avoiding such outcomes is to ensure that AI systems share human values. Why not simply make the ultimate goal of an AI system something general like improving human welfare? Unfortunately, this proposal—that we simply build in human values—faces a number of serious problems. Here I outline three:
1. Specifying human values: we do not know what the ultimate human values are and there is deep and persistent disagreement about them.
2. Ensuring uptake: even if we knew what the ultimate human values are, we have no way of knowing whether we have succeeded in programming them into an AI system.
3. Unintended consequences: even if we settle disagreements about human values and program these values into AI, there might be unintended consequences.

I will consider these issues in order.
First, the term ‘human values’ is unclear, and there is no obvious way of making it precise. There is no consensus about many of the fundamental moral values that would be relevant for specifying the goals of an AI system. For instance, some moral philosophers endorse radical versions of consequentialism, or the ethical position that only the consequences of an action—and not other factors such as its motivation, or whether it violates intuitive principles of justice—are morally relevant. For a consequentialist, sacrificing the welfare of some to provide a greater benefit to others is not just morally acceptable, but is in fact morally obligatory. Many others reject this consequentialist viewpoint, holding instead that each person has certain unalienable rights that cannot be interfered with even for some ‘greater good.’
Second, even if we came to an agreement about the nature of value, we would still face the difficult practical challenge of instilling these values in AI systems. This is easier said than done. The current machine learning strategies for creating AI systems are infamously opaque. It is currently an unsolved problem knowing what rules or goals AI systems created by machine learning methods are in fact employing. A system trained to sort resumes based on the projected aptitude of the applicants may learn problematic rules such as favoring applicants with the types of names which have traditionally been hired for these positions in the past. Only a careful review of the output of such systems will reveal that they are using inappropriate procedures. Likewise, a system may appear to act in ways that conform to human values in training but later reveal itself to be acting in ways that are contrary to human values when it is released into the world.
Third, even if we determine the ‘correct’ values and figure out how to program them into an AI system, there are potential unintended consequences. Suppose the ultimate moral goal is to reduce suffering and we give this goal to an AI system. It might reason that the best way to achieve this goal is simply eliminate all conscious beings, since without consciousness there can be no suffering. Or perhaps the ultimate moral goal is to produce as much pleasure as possible, but this is most easily done by hooking human beings up to a device which artificially stimulates pleasure centres in the brain. These scenarios may seem outlandish, but arguably they are logical extensions of widely accepted ethical theories. We cannot easily know in advance whether a highly intelligent and capable AI system will share our ‘common sense’ aversion to these ideas.
For these reasons, the value alignment problem poses a serious—even potentially catastrophic—risk. Without some way of reliably instilling appropriate values in AI systems, we are at the mercy of whatever values they happen to develop. Given the potentially superhuman capabilities of these systems, this poses a significant risk for humanity.
Solving the Value Alignment Problem
Having explored the alignment problem in this section, I now want to consider potential solutions to it. How might we solve the alignment problem? Broadly, work falls into two categories. First, many AI researchers are looking into technical solutions, for instance by trying to make the behaviour of AI systems more intelligible to their designers. But these technical problems currently remain unsolved. Second, some AI researchers advocate for a slowdown or even moratorium on AI research. A recent open letter, signed by nearly 1,000 academics and researchers, proposes a 6-month stoppage on advancing AI research. Whether such policies will be adopted by society is a collective decision that we must make.
We do not want to give off the impression that the value alignment problem is hopeless. Many very clever and good-intentioned people are hard at work trying to prevent potential future AI systems from leading to catastrophes like the one discussed in the video above. But the issue is real. For it is hard to see all the potential ways in which the values of future AI systems might diverge from our own. On the hypothesis that these are just as intelligent as we are---perhaps much more so---then we need to ensure that their values align with our own.
Bostrom, Nick. (2003). ‘Ethical issues in Advanced Artificial Intelligence.’ URL: https://nickbostrom.com/ethics/ai.
Bostrom, Nick. Superintelligence: paths, dangers, strategies. Oxford University Press.
Q1) What does it mean for an automated decision making system to be biased? What are the main sources of bias in automated decision-making systems?
(Incorrect) The correct answer is B.
Q2) What is the orthogonality thesis? How does it bear on the potential risks posed to humans by AI systems?
(Incorrect) The correct answer is A.
(Incorrect) The correct answer is A.
Return to the Module List






