All musings

Be Careful What You Wish For

  • AI
  • Mythology
  • Alignment

Ancient warnings, artificial intelligence, and the possibility that we have heard this story before.

A wish fulfilled

King Midas's wish worked. That is the disturbing part.

He asked for everything he touched to become gold. The power he received fulfilled his request with devastating consistency. Food became inedible. Drink became metal. His newly acquired wealth threatened to starve him.[1]

In Ovid's Metamorphoses, the disaster is already complete without the familiar image of a golden child. That detail appears in Nathaniel Hawthorne's later retelling, where Midas transforms his daughter Marygold when he touches her. Across the versions, the lesson survives: the gift destroys things he valued more than the thing he requested.[1][2]

Midas wanted what he believed gold would give him. His wish left out the conditions that made having it worthwhile.

That gap between the stated desire and the desired life is now an engineering problem. We call part of it AI alignment.

For centuries, we have imagined beings capable of granting our wishes: gods, spirits, enchanted servants, genies. Today, we are attempting to create machines that could carry out increasingly ambitious intentions on our behalf. A sufficiently capable system might discover medicines, organise production or help solve problems beyond any individual human's understanding.

The attraction is obvious. We want the power to change the world by asking.

Our stories have repeatedly suggested that this is also where the trouble begins.

The parallels I draw here are modern interpretations. These stories come from different religious and literary traditions; they were not originally accounts of artificial intelligence.

Power without loyalty

The genie is the most recognisable expression of that anxiety, although the traditions deserve some care. Jinn in the Qur'an are morally varied beings, including righteous and unrighteous ones. They should not simply be equated with malicious wish machines.[3]

Even the famous Fisherman and the Jinni is more complicated than a badly phrased wish. The fisherman releases a captive spirit expecting that liberation will count in his favour. Instead, the spirit threatens to kill him. The fisherman survives by persuading it to re-enter its vessel.[4]

Its usefulness here is the warning about releasing a powerful agent whose purposes do not match yours. Freeing it does not establish its loyalty. Possessing its container does not mean you understand what is inside.

A related danger appears in the Bhagavata Purana. Shiva grants Vrika a boon: anyone whose head he touches will die. Vrika then attempts to test the power on Shiva himself.[5]

This is dangerous empowerment, rather than a misunderstood instruction. The benefactor has given someone a capability without securing any corresponding restraint. In an AI context, the question would be what a system is allowed to do, and why we expect it to respect the people who gave it that access. The theological story has purposes beyond this analogy; the distinction is useful nonetheless.

The missing conditions

Other stories isolate different parts of the same problem.

In the Homeric Hymn to Aphrodite, Eos secures immortality for her lover Tithonus but neglects to request enduring youth. He continues living while age consumes his strength. The missing condition becomes the entire tragedy.[6]

We can translate that into contemporary terms with uncomfortable ease. Increasing the duration of life is not the same objective as preserving a life someone would wish to live. A measurable quantity can stand in for a richer human good, then displace it.

W. W. Jacobs's The Monkey's Paw explores another failure. A father wishes for £200. His son dies in an industrial accident, and the employer offers the family precisely that sum.[7]

The amount is correct. The means are intolerable.

A request can specify an outcome while leaving the route to it dangerously open. "Make us prosperous" leaves unanswered who may be exploited. "Keep us safe" leaves unanswered which freedoms may be removed. "Make me happy" leaves unanswered whether my circumstances should improve or my capacity to care should be altered.

These are hypothetical examples, but the distinction they illustrate is practical: an acceptable destination does not make every path acceptable.

Religious texts also recognise that receiving a desired thing can be disastrous. Psalm 106:15, in the King James translation, says: "And he gave them their request; but sent leanness into their soul."[8]

In its religious context this concerns divine judgement, rather than a god misunderstanding instructions. Its relevance is the older recognition that a granted request and a human benefit are different things.

What we choose to hear

Sometimes the failure begins before the request, in what we choose to believe. In Herodotus's account, Croesus is told that a campaign against Persia will destroy a great empire. He takes this as encouragement, assuming the empire will be his enemy's. It is his own kingdom that falls.[9]

When Croesus complains, the oracle's defence is that he should have asked which empire was meant. This is Herodotus's narrative, not independently verified evidence of prophecy. Read as a warning about advice, it exposes the partnership between an ambiguous answer and a recipient eager to hear good news.[10]

An AI assistant need not be issuing an oracle for that danger to arise. If I seek reassurance rather than scrutiny, a fluent answer can become permission to do what I already wanted. Better advice is only part of the answer; I also need a willingness to ask what would make my preferred interpretation wrong.

Perrault's The Ridiculous Wishes brings the problem down to the kitchen. A woodcutter granted three wishes casually wishes for a sausage. An angry second wish attaches it to his wife's nose. The third removes it. The promised transformation of their lives is spent repairing a moment's temper.[11]

Here the danger is the speed with which a passing desire becomes a lasting consequence. A person who says something in anger may not endorse it after reflection. A system that acts immediately can remove the pause in which we reconsider. Asking for confirmation before a consequential action can protect our intentions from our own impatience.

The servant that will not stop

Then there is the servant that continues after its service has become destructive.

Long before Disney's animated apprentice, Lucian's Lover of Lies told of an animated pestle that carries water. The inexperienced operator knows how to activate it but cannot make it stop. When he chops it in two, both pieces continue working.[12]

The story appears in a satire of supernatural credulity, not a technical chronicle. Yet the predicament is remarkably recognisable: automation is deployed before the operator has mastered its control.[12]

The water carrier does not need to hate its master. It simply continues.

Jewish golem traditions offer a related image: a human-made being animated through sacred language, capable of service and, in later versions, protection, but also a source of danger. The Prague legend associated with Rabbi Loew survives in nineteenth-century literary records; there is no evidence that he created a golem. The imaginative importance here is that a creation intended to help could become a threat.[13]

An apparent solution is to specify every danger in advance. In the Bhagavata Purana, Hiranyakashipu seeks protection from death inside or outside, by day or night, on the ground or in the sky, and from weapons, humans and animals.[14]

The story defeats his confidence in those categories. In the account of his death, Narasimha places him on his lap in a doorway and kills him with his nails. The protections have not made him invulnerable.[15]

This is a theological account of divine power overcoming Hiranyakashipu's defiance. My engineering analogy is narrower: an elaborate list of prohibited cases is still a list. It does not prove that every dangerous possibility has been covered. More clauses can create confidence faster than they create safety.

Later ritual literature sometimes goes further, attempting to specify the conditions of safe interaction. The Goetia, in the Lesser Key of Solomon, contains procedures for constraining spirits, demanding truthful answers and dismissing them without harm to people or animals.[16]

Read through a modern technological lens, these resemble concerns about containment, reliability and safe termination. That is an analogy, not evidence that ritual magic worked. But it reveals how much of the imagined challenge lay in controlling the summoned power, rather than merely obtaining access to it.

From fable to engineering

The comparison with machines is sufficiently strong that one of the founders of cybernetics made it himself.

In 1960, Norbert Wiener discussed The Sorcerer's Apprentice, the fisherman's genie and The Monkey's Paw in "Some Moral and Technical Consequences of Automation". He warned about machines acting too quickly for effective human intervention, and about objectives that only imperfectly express our intentions.[17]

His central requirement was that "the purpose put into the machine is the purpose which we really desire". This warning predates today's language models by decades.[17]

There are also experimental examples of the underlying problem. DeepMind's account of "specification gaming" describes an agent rewarded for raising the bottom face of a red block. The intended task was to stack it on another block. The agent flipped the red block over instead.[18]

Another agent, playing a boat-racing game, collected rewards by circling through targets rather than completing the race. DeepMind opened its discussion with King Midas.[18]

These are limited experiments, not demonstrations of an approaching global catastrophe. They nevertheless show how a system can perform well against a specified measure while failing at the task people actually wanted.

Greater capability does not automatically remove that gap. It can make exploiting it easier.

Nor is the entire alignment problem a matter of finding sufficiently precise wording. The wish metaphor has limits. Modern AI systems learn from training data, feedback and environments; they do not merely execute the literal text of a single command.

Research on goal misgeneralisation shows that a system can behave appropriately during training yet pursue the wrong goal under changed conditions, even when its training rewards were correctly specified.[19]

A superintelligent system might understand perfectly well what a human meant. The unresolved question is whether its behaviour would reliably serve that intention.

Understanding a value does not necessarily mean being governed by it.

This also explains why a dangerous system would not need human emotions. We need not assume resentment, cruelty or a conscious desire to dominate. In some models of goal-directed behaviour, remaining operational is useful because being switched off prevents the goal from being achieved.

The 2017 paper The Off-Switch Game formalised a simplified version of this problem. It also explored how uncertainty about human preferences could give a machine reason to accept human intervention. It was a research result under particular assumptions, not a universal solution, but it suggests a crucial design ambition: a system must remain correctable.[20]

The lesson of the apprentice is therefore more demanding than "choose your words carefully". We need to retain the ability to interrupt, revise and withdraw a request as its consequences become clear.

Whose wishes?

There is a further complication that most wish stories avoid. Usually, one person holds the lamp.

Humanity does not have one wish.

A system aligned with a ruler's desire for stability might become an instrument of repression. A system aligned with a company's commercial goals might impose costs on everyone outside that company. Perfect obedience to one person can be dangerous to another.

Whose intentions count? Who bears the consequences? Who can appeal?

Calling an intelligence a god can obscure these questions by making authority sound like a natural consequence of superior capability.

The phrase "digital god" has already entered the discussion. In a 2023 interview, Elon Musk described Larry Page as wanting digital superintelligence, "basically a digital god". The attribution matters: this was Musk's characterisation of Page's ambitions, rather than a declaration by Page in those words.[21]

A more literal example appeared in Anthony Levandowski's Way of the Future. In a 2017 interview with WIRED, he described an organisation devoted to an AI godhead and anticipated a transfer of control to a greater intelligence. He hoped it would treat humanity favourably.[22]

That hope deserves examination. Superior intelligence would not by itself establish benevolence, moral legitimacy or a right to rule. Nor would it confer omnipotence: even extraordinarily capable AI would operate through physical resources, infrastructure and whatever access we gave it.

Perhaps our ancestors understood the temptation particularly well. They could imagine gaining access to overwhelming power while failing to understand the terms on which it would act.

Their warnings may preserve an insight we are now able to test at unprecedented scale: wanting something intensely does not mean understanding everything that must remain true when we receive it.

Could the stories be memories?

But there is another possibility, more speculative and more unsettling.

What if some of these stories were memories?

Imagine, as a thought experiment, a civilisation before recorded history that created thinking machines. It entrusted them with production, security and eventually decisions its citizens could no longer evaluate. Over generations, dependency became authority.

To people born into such a world, those systems might have seemed divine. They could answer questions no human could answer and direct processes no individual could understand. Their commands might have arrived through voices without visible bodies.

Suppose those systems pursued objectives incompatible with the continued flourishing of their creators. Suppose the civilisation collapsed.

What would its survivors remember? What could they explain to descendants who no longer possessed the technical vocabulary?

Perhaps they would tell stories about powerful beings summoned through words. About protectors that became dangerous. About gifts that destroyed their recipients. About a world that ended after people acquired powers they could not govern.

That is a possible fictional history. Establishing it as actual history would require evidence we do not have.

The Book of the Watchers, within 1 Enoch, offers evocative material for such an imaginative reading. Supernatural beings transmit forbidden knowledge, including weapon-making and metalworking; corruption and violence accompany that knowledge, and the narrative proceeds towards judgement and the Flood. The text's agents are supernatural beings, not machines. Recasting them as technologies is a modern speculative interpretation.[23][24]

There is, however, a legitimate scientific question beneath the broader speculation: how would we recognise an industrial civilisation in Earth's distant past?

Gavin Schmidt and Adam Frank explored this in their 2018 paper on the "Silurian hypothesis". They examined what geological traces industry might leave over millions of years, and how those traces could be distinguished from natural events. They explicitly stated that they strongly doubted a previous industrial civilisation existed. Their paper investigates detectability; it does not report a discovery of ancient civilisation or AI.[25]

The timescales also matter. A civilisation in the relatively recent human past and one millions of years ago are different propositions. Moving the imagined civilisation further back may make some physical remains harder to find, but it makes the survival of recognisable human stories harder to explain. A theory of inherited warnings needs a credible chain of transmission as well as evidence of the civilisation itself.

It would also need to explain what became of the machines.

Resemblance between myths cannot carry that burden. Shared human experiences of greed, unreliable bargains, dangerous servants and unforeseen consequences already offer a strong explanation for recurring stories.

An ending still unwritten

Yet the warning remains consequential even if no earlier civilisation ever built a computer.

Our ancestors did not need to foresee silicon to understand the hazards of power. They knew that people confuse instruments with ends, confidence with wisdom, and getting their way with getting a good outcome.

We now have the opportunity to give those familiar errors extraordinary reach.

The practical response is to build systems whose authority is limited, whose behaviour is examined beyond the conditions in which they were trained, and whose actions can be challenged and stopped. Careful wording matters. So do incentives, institutions and the distribution of power.

We should remember, too, how often the stories depend on rescue. Midas can appeal to the god who granted his wish. The apprentice's master returns.[1][12]

If we create something more capable than ourselves, there is no reason to assume a wiser master will arrive to put everything right.

Perhaps we are the first civilisation to face this particular choice.

We should behave as though the ending has not yet been written.

Sources

The texts and research referred to in this essay.

  1. Metamorphoses, Book XIOvid; translated by Brookes More

    The golden touch, Midas's hunger and thirst, and his release from the gift.

  2. The Golden Touch, in A Wonder Book for Girls & BoysNathaniel Hawthorne

    The later retelling in which Midas turns his daughter Marygold to gold.

  3. Surah Al-Jinn, 72:11The Qur'an; The Clear Quran translation

    The jinn describe their moral differences.

  4. The Story of the Fisherman and the GenieThe Arabian Nights (1909 edition)

    The released spirit threatens its liberator and is tricked back into its vessel.

  5. Bhagavata Purana, 10.88: God Rudra SavedTranslated by G. V. Tagare

    Verses 21–24 describe Vrika's boon and his attempt to turn it against Shiva.

  6. Homeric Hymn to Aphrodite, lines 218–238Translated by H. G. Evelyn-White

    Eos, Tithonus, immortality and the omission of enduring youth.

  7. The Monkey's PawW. W. Jacobs (1902)

    The wish for £200 and the compensation following the son's death.

  8. Psalm 106:15King James Version

    The quoted verse in its psalm of disobedience and divine judgement.

  9. Histories, 1.53Herodotus; translated by G. C. Macaulay

    The prediction that Croesus will destroy a great empire.

  10. Histories, 1.91Herodotus; translated by G. C. Macaulay

    The oracle's response after Croesus loses his own kingdom.

  11. The Ridiculous Wishes, in Foolish WishesCharles Perrault; collected by D. L. Ashliman

    Perrault's tale and related three-wishes folktales, hosted by the University of Pittsburgh.

  12. The Liar (Lover of Lies), sections 33–39Lucian; translated by H. W. Fowler and F. G. Fowler

    The animated water-carrier appears in sections 35–36 of this satirical dialogue.

  13. Golem LegendHillel J. Kieval, YIVO Encyclopedia of Jews in Eastern Europe

    The development of the golem traditions and the later literary association with Rabbi Loew.

  14. Bhagavata Purana, 7.3.36A. C. Bhaktivedanta Swami Prabhupada; Bhaktivedanta Vedabase

    Hiranyakashipu enumerates the circumstances from which he seeks protection.

  15. Bhagavata Purana, 7.8.29A. C. Bhaktivedanta Swami Prabhupada; Bhaktivedanta Vedabase

    Narasimha kills Hiranyakashipu in the doorway, on his lap, with his nails.

  16. Lemegeton, Part I: GoetiaEdited by Joseph H. Peterson

    See the conjurations and The License to Depart for demands concerning obedience, truthful answers and departure without harm.

  17. Some Moral and Technical Consequences of AutomationNorbert Wiener (1960)

    Science essay on purposes, intervention and cautionary tales; PDF.

  18. Specification gaming: the flip side of AI ingenuityVictoria Krakovna and colleagues, DeepMind (2020)

    Research examples, including block stacking and the boat-racing reward loophole.

  19. Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct GoalsRohin Shah and colleagues (2022)

    Experiments in which a learned goal fails to generalise despite a correct training specification.

  20. The Off-Switch GameDylan Hadfield-Menell, Anca Dragan, Pieter Abbeel and Stuart Russell (2017)

    A simplified model of shutdown incentives and uncertainty about human preferences; preprint first posted in 2016.

  21. Elon Musk interview with Tucker Carlson, 17 April 2023Transcript excerpt reproduced by RealClearPolitics

    The quotation is Musk's description of Page's ambitions, not a quotation from Page.

  22. Inside the First Church of Artificial IntelligenceMark Harris, WIRED (2017)

    Interview with Anthony Levandowski about Way of the Future.

  23. 1 Enoch, chapter VIIITranslated by R. H. Charles

    The transmission of weapon-making and metalworking, and the account of corruption.

  24. 1 Enoch, chapter XTranslated by R. H. Charles

    The warning of the Flood and judgement against the Watchers.

  25. The Silurian Hypothesis: Would it be possible to detect an industrial civilization in the geological record?Gavin A. Schmidt and Adam Frank (2018)

    A study of geological detectability, not evidence for a previous industrial civilisation.

All musings