Blue Dots and the Gradual Delegation of Judgement to AI
How shifting standards shape our sense of harm, and what they might mean for the decisions we hand to machines.
Moving standards
Imagine a society that succeeds in reducing a serious problem. There is less cruelty, less discrimination, less danger. What should happen next?
We might expect people to recognise the improvement, become less alarmed and redirect their efforts towards the remaining work. But there is another possibility: their threshold for recognising the problem changes. Behaviour that once seemed peripheral becomes central. Ambiguity becomes evidence. The language of emergency survives the decline of the emergency itself.
The blue-dot experiments suggest that our standards can move with our surroundings. They also raise a question about our relationship with AI: could the boundary between assistance and dependence shift until decisions we once reserved for ourselves seem ordinary to delegate? The experiments establish a change in classification. The second possibility is the subject of this essay.
In a 2018 paper in Science, David Levari and colleagues asked participants to classify individual dots on a blue–purple spectrum. In the first experiment, the probability of drawing from the blue half fell from 50% to 6%. Participants increasingly called previously excluded colours blue. Blue had become a broader category. Related experiments found shifts in judgements of threatening faces and the acceptability of research proposals. The researchers called the phenomenon "prevalence-induced concept change".[1]
Blue dots became rare, not extinct. Participants admitted colours they had previously excluded from the category, and the effect appeared after abrupt as well as gradual changes. The result concerns a shifting threshold, not an inability to see that the world has changed.[1]
This offers a way to think about public anxiety, provided we keep the limits of a laboratory task in view. The frequency of a problem and the threshold for recognising it are separate variables. If the threshold falls while the problem becomes less common, the number of things attracting concern need not fall with it.
Tocqueville recognised a related paradox in Democracy in America: "the slightest dissimilarity is odious in the midst of general uniformity". Greater equality can make the remaining inequalities more conspicuous. Progress changes the background against which we judge what remains.[2]
What counts as harm?
Psychologist Nick Haslam's idea of 'concept creep' describes a related change over longer periods. In 2016, he examined the expanding meanings of abuse, bullying, trauma, mental disorder, addiction and prejudice. He distinguished horizontal expansion, which includes different kinds of experience, from vertical expansion, which includes milder examples. Recognising workplace bullying alongside school bullying illustrates the first; applying the label to less severe behaviour illustrates the second.[3]
That distinction matters because expanding a category can be an achievement. An older definition may have excluded serious suffering simply because it was poorly understood or socially tolerated. Haslam explicitly resisted treating all conceptual expansion as either good or bad. We need to ask what an enlarged definition helps us recognise and what distinctions it erases.[3]
An exploratory experiment published online in 2024 brought the question closer to everyday language. Sven Speerforck and colleagues asked 138 students to classify descriptions as indicating mental illness or not. When clear examples became rarer, participants broadened their judgements. The effects were small, and the study concerned classification in a task, not the prevalence of illness in society.[4]
Politics and proportion
Progressive politics is one setting in which this tension matters. A movement concerned with protecting people from harm has good reasons to search for suffering others overlook. How does it distinguish a neglected injustice from an increasingly expansive interpretation of an existing category?
Research reviewed by Haslam and colleagues links broader harm concepts with political liberalism, empathy and concern about injustice towards others. These associations help explain why the tension is relevant to progressive politics; they do not establish that any particular complaint is exaggerated.[5]
The strongest criticism concerns proportion. If an awkward remark, persistent harassment and deliberate exclusion are all described with the same language, the label may stop communicating how much harm occurred. If disagreement automatically becomes evidence of hostility, the possibility of resolving a misunderstanding shrinks. A response can be sincere and still be disproportionate.
The pattern also has a counterpart on the right. Harper, Purser and Baguley found broader definitions aligned with ideological concerns across the political spectrum, including liberal definitions of prejudice and conservative definitions in areas such as personal responsibility. These were differences between groups, rather than direct measurements of change within individuals. The wider lesson is that the concerns we prioritise help shape the categories we use.[6]
What, then, of the intuition that comfortable lives leave people searching for smaller problems? Ronald Inglehart's theory of postmaterialism offers a more careful starting point: material security can change people's priorities, increasing the importance of concerns beyond immediate survival. A 2017 study using World Values Survey data from 59 countries found support for links with individual socioeconomic circumstances, but weaker support for the expected national economic explanation.[7]
Material security can give people room to pursue concerns beyond immediate survival. That is not evidence that those concerns are invented. The question is when changing priorities and greater sensitivity produce a proportionate response.
Progress need not mean that hardship has disappeared. The World Bank's June 2025 estimates put the reduction in extreme poverty between 1990 and 2022 at about 1.5 billion people, while still estimating 838 million below its updated extreme-poverty line in 2022. Both facts belong in the picture.[8]
What the feed rewards
Social media adds an incentive to express concern forcefully. Brady and colleagues studied 12.7 million tweets from 7,331 users and conducted two behavioural experiments. Positive feedback encouraged subsequent expressions of moral outrage, while users also adapted to the expressive norms of their networks. Approval helps teach people how to speak.[9]
Public discourse is therefore an imperfect emotional census. A platform can reward a way of speaking without revealing how intensely each speaker feels it. People may learn which accusations attract approval without consciously deciding to exaggerate.
Rathje and colleagues found a complementary pattern in roughly 2.7 million posts from news organisations and US congressional accounts: posts referring to political opponents were shared or retweeted about twice as often as posts referring to the author's political group. Antagonistic framing attracted attention.[10]
Put these findings together and a possible feedback loop emerges. Changing standards affect what gets labelled harmful; social rewards affect what gets expressed; attention to opponents affects which examples circulate. Observers may then infer that the underlying problem is becoming more common. This is a synthesis of separate findings, rather than a single mechanism established by the studies. A feed can become more alarming without providing a reliable measure of conditions outside it.
Mastroianni and Gilbert's 2023 research on the illusion of moral decline adds another piece. People repeatedly reported that morality was deteriorating, while repeated assessments of contemporary morality did not show the corresponding downward trend. Their explanation centred on biased exposure and memory. The feeling that things are getting worse is itself something to examine.[11]
Together, these findings suggest that complaint is an unreliable standalone measure of harm. More complaints can reflect more harm, greater willingness to report it, broader definitions, increased exposure or stronger incentives to publicise it. These explanations may operate simultaneously. Counting accusations cannot tell us which explanation dominates.
Recognising improvement
The strongest objection to the blue dot argument is also its most useful safeguard: smaller problems can still be real problems. Reducing severe harassment does not make a persistent pattern of milder harassment acceptable. Conduct once treated as harmless may prove damaging when investigated. Historical normality is no guarantee of moral correctness.
But recognising this does not require treating every harm as equally severe. We can distinguish discomfort from intimidation, an isolated mistake from a pattern, and disagreement from coercion. The challenge is to keep our responses proportionate while remaining willing to revise our understanding when evidence warrants it.
These judgements also respond to how a task is designed. Lyu and colleagues found that feedback could reverse the direction of prevalence effects in their perceptual tasks. Levari later replicated the colour effect and eliminated it by extending the range of colours presented. Reference points and feedback matter; a moving threshold is not an unavoidable law of judgement.[12][13]
The practical lesson is to keep our standards visible. Measure severity as well as frequency. Explain when definitions change. Ask what evidence would count as improvement. A society needs the ability to recognise success as well as the vigilance to identify failure.
The quiet transfer
The AI parallel begins with a different boundary: which parts of thinking and deciding do we regard as ours to do? Research on cognitive offloading and reliance on automation helps us examine that question. The blue-dot experiment supplies an analogy of a moving standard; it does not establish that AI delegation follows the same psychological mechanism.
Consider a possible progression. First, an AI corrects an email. Then it drafts the email from notes. Later, it decides which points to include, how to interpret the recipient's motives and which outcome to pursue. Eventually, it handles the exchange and tells us what was agreed. Each step can feel like a modest extension of the previous one. Taken together, they transfer a substantial part of our judgement.
The same progression becomes more intimate when the subject is a relationship, a child's education or a major life choice. Asking for help expressing an apology is one thing. Letting a system decide whether an apology is owed reaches further into our values. A person who would once have recoiled from delegating that judgement could come to experience it as ordinary assistance.
The question is how that change becomes normal. If we compare each new delegation with yesterday's practice, rather than with our original boundary, the cumulative transfer can become difficult to see. What once demanded justification becomes the default from which the next request begins.
Cognitive offloading is already an established part of human behaviour. Risko and Gilbert describe how we reduce cognitive demands by using our bodies and external tools, from reminders to search engines. Such practices can improve performance and extend what we can accomplish. The relevant question is which mental work the tool takes over, and what work remains with us.[14]
AI makes the prospect of delegating judgement more immediate because it can supply an interpretation as well as information. Consider the difference between consulting a diary and asking an assistant which obligations deserve your time. The first supports a decision; the second participates in deciding what matters. The interface may look much the same, even as the authority being exercised changes.
In this scenario, growing capability drives growing reliance for understandable reasons. A system that repeatedly saves time and produces useful results earns another task. Success at that task earns more discretion. Eventually, asking it to act can feel as routine as asking it to advise. The temptation is to carry confidence earned in one kind of task into another with different stakes.
A good summary of an email does not establish good judgement about the relationship behind it. A persuasive analysis of an investment does not establish that the system should choose your tolerance for risk. Technical competence, authority to act and legitimacy to determine priorities are different things. A smooth user experience can obscure those differences.
Assistance and ability
A 2025 survey by researchers at Microsoft Research and Carnegie Mellon University examined 319 knowledge workers. Higher confidence in generative AI was associated with less self-reported critical thinking; greater confidence in one's own ability was associated with more. Participants also described changes in the effort involved. These are associations in reported behaviour, not evidence that AI use had caused a lasting decline in their abilities.[15]
A further risk is that delegation could weaken the conditions for independent judgement. We hand over a task because the AI makes it easier, then practise it less. If our ability or confidence falls, doing it unaided becomes harder, giving us another reason to delegate. That cycle is a possibility to investigate, rather than an outcome established by the survey.
An educational experiment makes the difference between assisted performance and independent capability concrete. Bastani and colleagues studied nearly a thousand mathematics students at a high school in Turkey. Access to a general-purpose GPT-based assistant improved performance during practice, but students subsequently performed worse without it than the control group. A version designed to support learning largely removed that penalty, without establishing an improvement in unaided exam performance. The design of the assistance mattered.[16]
One classroom trial cannot establish what all AI use will do to learning. It does, however, give us a reason to ask whether a tool helps us develop competence or simply supplies answers. Both can feel productive while we are using them. What we can do afterwards is a separate test.[16]
Approval and control
Now extend the scenario to increasingly capable systems operating across work and private life. Their value could make opting out expensive. A manager who insists on independent review might be slower than competitors. A person who makes every decision unaided might feel overwhelmed beside someone with an effective assistant. Delegation could become a social expectation as well as a personal preference.
At that point, the question could change from "Why would you let a machine decide this?" to "Why would you insist on deciding it yourself?" That reversal is the heart of the parallel. The boundary of acceptable delegation moves, and the burden of justification moves with it.
Human control could survive in form while thinning out in practice. We would still approve the recommendation, but the system might have selected the evidence, defined the options, ranked the trade-offs and drafted the explanation. Our final click would authorise a decision whose structure we had barely examined.
Meaningful oversight requires understanding, time and freedom to disagree. If the system supplies both a recommendation and all the reasons for accepting it, review risks becoming circular. Independent checks matter most where a mistaken decision would be difficult to undo.
This possibility does not require a conscious machine, a hidden agenda or a dramatic seizure of power. Ordinary incentives could be enough: users seek relief from effort, organisations seek speed, and providers seek greater adoption. A consequential transfer of authority could emerge from useful services that people repeatedly choose.
Increasing capability also makes some additional delegation entirely sensible. The standard should be evidence of fitness for a particular task, together with clear limits on authority. Research on trust in automation frames the objective as appropriate reliance: confidence that matches what a system can actually do.[17]
The danger would be allowing familiarity to stand in for that judgement. "It has become normal" is a description of acceptance. It does not tell us whether we still understand the decision, can reverse it, or have knowingly authorised the power being exercised.
What we choose to retain
A useful starting point is to distinguish help with expression, help with reasoning, a recommendation and permission to act. Moving from one to the next should be an explicit choice. Organisations need to evaluate the scope of authority as carefully as the quality of output. A review of what a system is allowed to decide should look at the whole arrangement, including permissions added gradually.
We should also ask what abilities a system helps its users retain. Can they explain the decision in their own words? Can they recognise a bad recommendation? Can they function when the service is unavailable? Can they reject its advice without losing the practical ability to get things done? Those questions make human agency something we can examine.
The two halves of the argument meet here. In public debate, changing thresholds can alter what counts as a serious problem. In our relationship with AI, changing expectations could alter what counts as acceptable dependence. The directions differ. Both require us to notice when the standard itself has moved.
A society can grow more capable while its citizens become less practised at exercising judgement. It can become more efficient while decisions become less intelligible to the people who authorise them. The prospect worth confronting is that these changes could feel like progress at every individual step.
The question is larger than 'Can the AI do this?' What are we handing over, what are we retaining, and would we choose the whole arrangement if we encountered it for the first time today?
Sources
The texts and research referred to in this essay.
- Prevalence-induced concept change in human judgment
Original experiments and supplementary methods; PDF.
- Democracy in America, Volume II, Part IV, Chapter III
The observation about inequality quoted in the essay. A historical argument, not experimental evidence.
- Concept Creep: Psychology's Expanding Concepts of Harm and Pathology
The distinction between horizontal and vertical expansion, and its possible benefits and costs.
- 'Concept creep' in perceptions of mental illness: an experimental examination of prevalence-induced concept change
Exploratory classification experiment with 138 students; the reported effects were small.
- Concept Creep and Psychiatrization
Review of concept expansion and individual differences in the breadth of harm concepts.
- Do Concepts Creep to the Left and the Right? Evidence for Ideologically Salient Concept Breadth Judgments Across the Political Spectrum
Author manuscript hosted by Nottingham Trent University; group differences do not demonstrate change within individuals. PDF.
- Inglehart's scarcity hypothesis revisited: Is postmaterialism a macro- or micro-level phenomenon around the world?
University publication record and abstract for the study using World Values Survey data from 59 countries.
- June 2025 Update to Global Poverty Lines
The quoted estimates refer to 1990–2022 and this dated release, rather than current-year poverty.
- How social learning amplifies moral outrage expression in online social networks
Observational Twitter studies and behavioural experiments; PubMed record with links to the full paper.
- Out-group animosity drives engagement on social media
Published paper and abstract in the University of Cambridge repository.
- The illusion of moral decline
Research on perceived moral decline, contemporary assessments, exposure and memory; PDF.
- Feedback moderates the effect of prevalence on perceptual decisions
Experiments comparing prevalence effects with and without feedback; PDF.
- Range-frequency effects can explain and eliminate prevalence-induced concept change
Replication, modelling and a colour-range intervention; PDF.
- Cognitive Offloading
Review of using physical actions and external tools to reduce cognitive demands; PDF.
- The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers
Survey of 319 knowledge workers. Reports associations and perceived effort, rather than a causal measure of long-term skill loss. PDF.
- Generative AI without guardrails can harm learning: Evidence from high school mathematics
Randomised classroom study comparing two GPT-based assistants and a control group; PDF.
- Trust in Automation: Designing for Appropriate Reliance
Research review on trust and reliance, with links to the publisher.