Dan Hendrycks' Moral Theory Is Very Implausible
Does the supreme principle of morality say that you matter 360 billion times more than foreign strangers?
1 Introduction
Following in a venerable tradition of smart polymath non-philosophers, Dan Hendrycks, director of the Center for AI Safety, just released a paper claiming to have discovered the supreme principle of morality. The core claim of Hendrycks’ theory, dubbed Eigenism, is that moral concern should vary with similarity—you should be more concerned about the welfare of those who are more informationally similar to you.
Hendrycks explains the view: “The core idea is simple: to evaluate an outcome, an eigenist agent needs to look at how well a particular entity is doing, and weigh it by how heavily that entity carries the agent’s specific identity pattern.” Or, put even more simply, how much you should care about others depends on how similar they are to you (with similarity measured by Shapley mutual information). Moral weight is also adjusted for redundancy, so that the second algorithm like you gets half as much weight as the first. Another important feature of the view is that it conserves total concern—a constant degree of moral concern is divvied up based on similarity, rather than total concern growing as similarity grows.
I think this is an extremely implausible moral view. Furthermore, it does not help with any of the problems that Hendrycks discusses. Since the view has been getting fairly fawning reception, it seemed worth analyzing in detail.
2 Implausible results
Eigenism has many very implausible implications.
First of all, Eigenism cares intrinsically about similarity. But similarity does not seem to itself be morally important. To see this, imagine that there are two people who are otherwise identical. However, one of them has the same favorite color that I do, and this is implemented by a similar mental algorithm. On Eigenism, I would have more reason to care about the person who shares my favorite color. But this is crazy.
Along these lines, Eigenism gives significantly more reason to care about those who are more like you. Imagine that there are two alien races, both of whom have equal ability to experience pleasure and pain, equal knowledge, and so on. For any trait that might appear morally relevant, they possess it in equal quantities. However, one has a brain that is more similar to humans. According to Eigenism, we ought to care significantly more about that alien than the other. In fact, even beings with identical mental lives will matter different amounts on Eigenism, so long as one has a brain that is more informationally similar to yours.
Or, to give another example, suppose that there are two otherwise identical alien species, but one has three limbs, while the other has two. On Eigenism, you should care more about the interests of the aliens with two limbs, because that leads to greater psychological similarity with humans.
Second, Eigenism’s adjustment for redundancy is a big problem. As the number of copies of some algorithm approaches infinity, to adjust for redundancy, the amount it cares about any particular copy approaches zero. But this is crazy! It would be very bad to create 100,000 extra copies of someone being tortured, even if there are already a lot of copies.
Suppose that there are googol copies of John already suffering. Because Eigenism conserves total moral concern, it implies that creating 100 quadrillion copies of John isn’t bad at all, or is only barely bad! After all, they simply split up the concern divided among existing Johns.
In fact, because the universe is probably infinite, Eigenism likely implies that no one on Earth’s interests matter at all. We’re all redundant, with our algorithm split infinite ways.
Third, Eigenism also seems to recreate the population ethical problems of average utilitarianism. Hendrycks writes:
Adding a vast population of barely-happy strangers does not merely fail to overwhelm the existing community through sheer numbers. With respect to the generic overlap they share with the existing pattern, it reduces community eigeninterest by reallocating connectedness from flourishing carriers to barely-happy ones.
But if adding new people simply splits the share of concern from current people to the new people, rather than creating more to be morally concerned about, then it is bad so long as the new people are less well-off than existing people (adjusted for similarity). But this is very implausible:
Imagine that everyone in the world was extremely miserable. You could create a large number of miserable people, but they were slightly less miserable than those who already existed. By Eigenism’s lights, this would reallocate eigeninterest from existing people to future people. Because the future people would be better off, such an action would be very good. This is extremely implausible.
Similarly, it implies that creating happy people wouldn’t be good so long as they were less happy than existing people.
Fourth, Eigenism faces a dilemma. Ask: what is the system that we are measuring the informational similarity of. The two answers Hendrycks regards as potentially viable are all properties of an agent and psychological states. Either one has big problems:
If the answer is any physical property then this values overlap based on properties that obviously don’t matter. On this view, I would have slightly more reason to care about white people with brown eyes because they’re more like me. Hendrycks addresses this by suggesting that “Shared common characteristics such as race provide negligible reasons for rational concern,” but it seems wrong for these to provide any reason for rational concern. Or, to pick a perhaps less charged example, it doesn’t seem like I should care any more about someone just because their esophagus is shaped similarly to mine.
If the answer is any psychological property, then one needs some in-principle way to delineate psychological properties from non-psychological properties. I’m very skeptical that there’s a good way to do that. Even aside from this, this view ends up caring more about psychological mechanisms for implementing some trait than non-psychological mechanisms. Thus, imagine there’s some psychological process that produces white skin. Eigenism would imply that in such a world, I should care a bit more about those with white skin. However, if there was no such brain process, then I shouldn’t care more about those with white skin. Both the judgment itself and the odd asymmetry are very implausible.
Fifth, Eigenism totally ignores the interests of those without any informational overlap, even if they’re conscious, can suffer, and so on.
The view has a number of other problems, but these are probably best discussed in the context of the view’s alleged advantages. So let’s turn to those.
3 Why adopt the view?
Hendrycks seems to think the view offers a solution to almost every puzzle in ethics. I disagree and think it doesn’t offer a solution to any problem in ethics.
3.1 Copying and deleting AIs?
Hendrycks suggests that one advantage of the view is that it can explain why it’s not a huge loss if an AI makes 1,000 identical copies and 999 are deleted. He explains:
Because the Shapley connectedness function penalizes redundancy, the AI’s pattern is distributed across all the identical copies. Deleting redundant instances is rationally and ethically much closer to closing browser tabs than to killing a thousand distinct persons. What matters to an eigenist AI is whether anything unique is lost
This faces two problems:
The view implies that so long as there is a copy of Earth elsewhere in the universe, it wouldn’t be bad for everyone on Earth to die. This seems false. Even if there was a copy of you somewhere else, it seems bad if you were killed.
Insofar as the view splits moral weight across all copies, it implies that if Bob is being tortured, creating a million copies of Bob being tortured wouldn’t be bad at all. Bob’s weight would get divided across the million copies. This clearly is not right. On such a view, creating 10 billion copies of Bob being tortured would be less bad than creating one unique copy of Sarah in mild discomfort.
3.2 Impartiality explained?
Hendrycks suggests that the view gives a nice middle ground between egoism and complete impartiality. On Eigenism, how much others matter is a degreed property tracking similarity. So you won’t either behave with total impartiality nor as an egoist.
On this view, your reasons to help your wife and child come from the fact that they are psychologically similar to you. That doesn’t seem right! Your reasons to help them should be for their sake, not because they’re sort of you. Here is how Hendrycks explains why you should help others:
Once we see that, the old forced choice between selfishness and sacrifice begins to dissolve. Caring for others is no longer a saintly departure from self-interest; it is simply what self-interest looks like once you realize how far you actually extend.
But again, this is an egoistic explanation of why you should help others. If egoism is false, then so is Eigenism. Here’s another way to see this: Eigenism claims that egoism is wrong, not because it’s incorrect that you should only care about yourself, but because it is wrong about what you are. But now suppose that egoism was right about what you are—perhaps tomorrow we definitively established that humans are unique Cartesian souls. The Eigenist worldview seems to wrongly imply that in such a world, you should only care about yourself.
Similarly, it implies you have more reason to help those of your children who are more similar to you. Again, this doesn’t seem to be capturing the core partialist intuition. Hendrycks suggests:
Think about a parent standing before a burning building, forced to choose between saving their own child or two strangers. Utilitarianism delivers a harsh verdict here: if all wellbeing counts equally, saving your own child is a mistake
But Eigenism delivers the verdict that if by chance the stranger’s child is more psychologically similar to you than your own child, then you should save the stranger’s child. Not exactly a vindication of partiality! Similarly, if you are not very psychologically similar to your spouse, then you won’t generally have much more reason to save them than a stranger.
Next, Hendrycks suggests that the view “places a reasonable limit on moral demandingness.” Insofar as you get to care more about those who you are more similar to, you have a justification for not spending your whole life helping strangers. But this is of little help:
The common intuition is that it is good to help strangers even if not obligatory. Yet on Hendrycks’ view, helping them above the relevant threshold is actively wrong.
This implies that if the strangers were very similar to you, there’d be no bounds to what morality demands.
This fails to secure the core intuition which is that helping strangers shouldn’t be all or nothing. Suppose that you can incur some constant cost in well-being (say, one lost util) to help some stranger a large amount. The common intuition is that you are required to pay this cost sometimes but not to keep doing it. But on Hendrycks’ view, insofar as both cost and benefit are constant, you are either always required to do it or never required to do it.
3.3 The experience machine
Hendrycks suggests that the view explains why you shouldn’t plug into the experience machine, where you’d float in a tank being happy but have no connection to the outside world. Because it would diminish your relationships and make your life very different, on Hendrycks’ view, it wouldn’t be good for you. I found this response strange:
The experience machine isn’t some deep unsolved puzzle in ethics. It’s simply an objection to hedonism. If well-being is about more than pleasure, then the experience machine is no problem.
Hendrycks’ view only secures the result that plugging in is bad if it’s very different from what your life is currently like. But this isn’t the core intuition. If you plugged a baby into the experience machine, and made sure its mental states remained similar to that of a baby, that would seem very bad. Crucially, Hendrycks’ view seems to imply that it could be good to plug a baby into such an experience machine even insofar as doing so diminished its happiness—so long as this minimized the degree of change across its mental life. This is very implausible.
This additionally implies that plugging into the experience machine would be good so long as your mental life in the experience machine was more similar to your current mental life than it would naturally be. Imagine, for instance, that you’re moving to a new country. In this new country, you’ll have friends, achievements, etc, but your life will be very different. In such a world, Eigenism implies you should plug into an experience machine that keeps your life roughly the same, even if it lowers your happiness and prevents you from getting any non-happiness good. Nuts!
3.4 Population ethics?
Hendrycks suggests that Eigenism can avoid the result that it would be good to replace people with other, happier people. Insofar as those other people are very different from the original people, the resultant state of affairs would be worse. I don’t think this is right:
As we’ve already seen, Eigenism implies that if someone has many clones, it would be worth killing and replacing them with some unique non-cloned individual with a just barely worthwhile life.
Eigenism cares about how much similarity there is to you. Thus, if everyone on earth was replaced with people very similar to you (but each similar to you in different ways), on Eigenism this would be a good thing. Crucially, this might hold even if they had lower total welfare.
Next, Hendrycks writes:
The temptation of replacement is captured by what Derek Parfit famously called the “Repugnant Conclusion” (Parfit, 1984). What follows is an AI-themed version. Imagine a world with a billion humans living deeply fulfilled, joyful lives. Now imagine a second world with a hundred trillion AIs whose lives are just barely worth living, experiencing only enough minor comforts to prefer being switched on over being switched off. Because a hundred trillion is such a massive number, the absolute sum of happiness in the AI datacenters is mathematically larger. Utilitarianism seems to demand that we trade our human civilization for a sprawling, barely happy AI population.
The repugnant conclusion is not about replacement. It’s about comparing worlds. Whether the world flooded with barely happy people is better than the one with 10 billion very happy people is distinct from whether it should be replaced with the other world.
But crucially, Eigenism implies that killing 10 billion happy people and replacing them with no one could be neutral, so long as they have copies, or their Eigenconcern would flow to others with higher welfare. If it’s bad to imply that we should replace happy people, then surely it is worse to imply we should kill them and not replace them! This is much worse than the repugnant conclusion.
In addition, Eigenism still implies the additive repugnant conclusion. Suppose that everyone in the world is very miserable, and they exist in great numbers. You can either create 10 billion very happy people or an arbitrarily large number of barely happy people. Eigenism implies that the latter would be better, because the newcomers would soak up most of the moral concern—diminishing the badness distributed across the miserable people.
Hendrycks fails to mention the reason why the repugnant conclusion has been so sticky—it turns out to be very difficult to avoid. There are a number of other plausible premises that collectively entail the repugnant conclusion. One of these, and the one that Hendrycks gives up on, is called benign addition. It says that improving everyone’s welfare and adding new happy people is good.
Hendrycks’ view requires abandoning benign addition. If you make everyone better off, and then create new people at a lower level of welfare, the new people soak up a big share of moral concern. Because their welfare is lower, adjusted for this, value divided up by similarity goes down. This strikes me as worse than accepting the repugnant conclusion.
3.5 How much people matter
In an Appendix, Hendrycks tries to calculate how your moral concern should be divided across different people. “Count,” measures how much of the entity there is. “Central (c),” measures how much each entity matters. On this view, however, you matter around 360 billion times more than foreign strangers. This doesn’t seem right! I don’t think that my enjoying one meal is more morally important than feeding everyone in India, for example.
Similarly, the view implies that you matter one 1.2 quadrillion times more than a chicken or fish. Thus, stubbing my toe is much worse than a million chickens being tortured to death.
The numbers also seem a bit made up. Sure, you have some shared memories with your spouse, but what if she has dementia? Does she then matter barely more to you than any other dementia patient? And why does a baby have shared psychology with you? Is it because of similar genetics? Well then what about an adopted baby? Hendrycks says that the estimates “should be read as illustrative rather than definitive,” but even if they’re just illustrative, it is a big problem for the rough approximation of what the view says to yield such wrongheaded results.
4 Conclusion
Credit to Hendrycks for venturing out and proposing a bold new theory of ethics! But the theory he endorses is very implausible and has none of the upsides that he proposes. It faces a number of extremely serious counterexamples and is not useful for solving the problems he says it solves. Those looking for the supreme principle of morality ought to look elsewhere, for some other theory of ethics with more (eigen)value.



Thank you for dismantling this bonkers idea.
Setting aside Hendrycks' zookeeper/menagerist idea on redundancy for now, I will address the similarity idea.
I realize my thoughts are far from novel, but the way I see the moral responsibility question is in how we properly balance our care for those closest to us, including current family, community, city, region, country, and yes, our local animals, with our care for those furthest from us in proximity, relation, or time, who can be affected by our actions or already are.
Our natural emotional tendencies encourage us to care for those closest to us. However, because we recognize that our actions can and do affect people globally, we search for an organizing principle for extending our moral responsibility beyond those closest to us.
The evil choose race, and Hendrycks chooses something scarily like that: similarity. Some reasonable people limit their care to humans, and to animals only to the extent that they benefit humans, while the most enlightened, including BB, recognize their moral responsibility to all humans and animals, present and future.
There are those who argue that we are better able to assess the needs of those near us, related to us, similar to us, or alive today or soon, but those factors are already well addressed by our natural emotional tendencies.
Nice post! I agree Eigenism is very implausible.
>It says that improving everyone’s welfare and adding new happy people isn’t good.
I think this should say 'bad.'