It seems like your example hinges upon the FDT agent picking Left.
But it also says that the predictor with a one in a trillion trillion error rate predicted you would pick Right.
If all of this is correct, it seems like you're hinging your example on this being the one time in a trillion trillion when the predictor was wrong.
But decision theories shouldn't be judged on whether they work well in unbelievably rare edge cases that you would never encounter in a million lifetimes.
Compare Lottery Decision Theory:
> you have the choice to buy a lottery ticket for $100. There is a one in a trillion trillion chance you will win. If you win you get $1 million. Should you do it? Before answering, keep in mind that, unbeknownst to you, this is the one time you would actually win.
We can use this example to prove that you should definitely play the lottery. I think the bomb situation maps to this - although it fails in the one/trillion trillion case where the superpredictor was wrong, it succeeds in the 9999999999.../trillion trillion cases where the superpredictor is right, and when you multiply out the probabilities by the utilities (eg of getting any extra $100 vs. getting bombed), FDT gets you more utility overall.
The particular way FDT succeeds is that you never (okay, 1/trillion trillion times, but this rounds to never) find yourself in this situation. So just by asking about this situation, you've already started with the assumption that FDT fails, which is why you are so easily able to prove that FDT fails.
I think this goes back to what I said last time we discussed this. You and Eliezer are optimizing for different things. You are optimizing for never finding yourself in a situation where you have to do something silly within that situation; he is optimizing for having the most utility overall if you can set your algorithm. I think the thing he's optimizing for makes more sense.
Yes you are depending on this being the one time in a trillion trillion when the predictor was wrong. Decision theories are judged by whether they give the right answer across the board. So if they get the wrong result even in weird edge cases, they're disproved. But the wrong result isn't the same as any case where adopting the situation is bad for you. So what matters is not "if they work well" but if they give the right answers. Compare: if there's a single result where utilitarianism gives the wrong answer to what you should do, that disproves it. This doesn't mean that any scenario where being a utilitarian turns out badly is a counterexample.
Decision theories are theories of rational action. That's different from what dispositions you want to have. If you're in a world where you encounter Newcomb's problems all the time, then you want have one-boxing dispositions. But that's different from a theory of rationality.
Also, just want to register, I think the bomby problems are actually much less big of problems than the "FDT doesn't say anything in any case" problems. I find the discourse has focused more on cases where FDT is counterintuitive and has weirdly neglected the fact that it's totally ill-formed.
It's fine for theories to give the wrong answer in a probabilistic sense, i.e., they recommend some action under uncertainty and you end up getting unlucky and losing out. I think what you want to say with the bomb case is that it gives the wrong answer under certainty.
Looking through that thread, the FDT people might want to say there really is uncertainty in the bomb case - namely, which person you are among the real you and your simulations. But I don't feel the introducing simulation aspect is actually necessary to the example.
>Yes you are depending on this being the one time in a trillion trillion when the predictor was wrong. Decision theories are judged by whether they give the right answer across the board. So if they get the wrong result even in weird edge cases, they're disproved.
That's a judgement call. FDT proponents would point out that the FDT agent is getting a higher utility - paying a cost in one branch to get a bonus in every other one. So they could argue that CDT is failing in almost every branch - deciding to lose utility, thus "disproving" CDT.
What I feel is really happening is that predictors make decision theory setups into game theory ones. CDT clings to decision theory while FDT accepts (some) game theory.
In fact, "bomb" maps well to an extortion game theory setup. Predictor is claiming that there is a bomb in the house, which will go off automatically unless you wire $100 to their account. Your actions are "pay or decline". Their actions are "make bomb or bluff". They get a penalty is the bomb actually goes off.
They are capable of predicting you, and will only make the bomb if they predict that you will pay up. Then if you are CDT, you will give in to extortion (pay), and hence the predictor will make the bomb. If you are FDT, you will resist extortion (decline) and hence the predictor will not make the bomb.
And, in one in a trillion trillion cases, the predictor will make a bomb for an FDT agent that will decline to pay - BOOM.
> Decision theories are judged by whether they give the right answer across the board.
In a world of uncertainty, it is generally not the case that you can win in every possible outcome. The question is whether you can get the best outcome in the aggregate.
> So if they get the wrong result even in weird edge cases, they're disproved.
Guaranteed Payoffs gets the wrong answer on whether or not you should take a bet on a fair coin which pays out $2 if you win and $1 if you lose.
> Also, just want to register, I think the bomby problems are actually much less big of problems than the "FDT doesn't say anything in any case" problems. I find the discourse has focused more on cases where FDT is counterintuitive and has weirdly neglected the fact that it's totally ill-formed.
Have you seen any decision problems where a FDT proponent has not been able to give the FDT answer?
Like, I think we can exhibit problems where it's challenging, or requires empirical estimation. ("How correlated is my decision to vote with other voters in my district?") But I think this is a challenge with the world, not with the decision theory, and generally it's easy to come up with the rule. ("If my influence on whether other voters similar to me vote is at least 0.02, then it makes sense for me to vote.")
It models the bet as a two-stage decision. First you choose whether or not to make the bet (obviously yes, it's positive EV) and then once you know the result, the loser gets to decide whether or not the bet was serious or a joke (by paying up or backing out). Your counterparty is a predictor who will take the bet seriously if they think that you will take the bet seriously.
Guaranteed Payoffs recommends backing out if you lose (after all, it's worse to pay up than not and the coinflip has already not gone your way), but this means the counterparty doesn't view the bet as serious and so they don't pay you if they lose, which means you lost the opportunity to make those sorts of bets.
I feel like this is not a "weird edge case"? Or, like, I still don't know how to interpret what you mean by "So if they get the wrong result even in weird edge cases, they're disproved." unless it's something like "all decision theories are disproved". (But if all decision theories are 'disproved', why do you care about 'proof'?)
I don't think your second paragraph here works as an argument. Here is an attempt for a structurally similar argument, which is obviously wrong.
Ethical theories are theories of right action. That's different from what actions are fortunate. If you find yourself in a trolley problems, then it would be fortunate for you to pull the lever. But that is different from it being right. (And so a theory of "fortunate" action is not a theory of right action)
I believe the bomb example doesn't work (or at least not as stated). As stated, the note is irrelevant and uninformative about the bomb, because it has to be presented to the simulation before the simulation makes a decision, and then the exact same note will have to presented to you, whatever your simulation does.
It's not clear that there is an example where a) the note is accurate, b) the error rate is as stated, and c) the FDT agent is forced Left in this case.
EDIT: after thinking about it, you can have the following setup: the predictor runs with "bomb on Left note" included. If that results in a self consistent action - if you would go Right in that prediction - then they do that. If the action is not self consistent, they run the prediction without a note and go with that (obviously not including a note).
If they value $100 more than they disvalue a 1/10^{24} chance of a bomb, then an FDT agent will go Left in the "bomb on Left note" situation (making it inconsistent) and also Left in the "no note" situation. So "in the real world" it only expects to see no note. So seeing the note in the real world means that the predictor has made an error. The example is coherent (I disagree with its implications, but the example is coherent).
I agree the note seems like a spurious add-on since for the thought-experiment to make sense, the person has to already know the full setup of the problem before making a choice, including the fact that the decision whether to put a bomb in Left depended on an earlier simulation (and the original simulation must have been given the same sensory input that the "real" biological version will later get if a bomb is put in Left, i.e. the simulation was lied to and told they were experiencing a real trial that was the result of a prior simulation--or I suppose they could both just be told that the bio version's setup will be the result of the sim version's response to the bomb setup, without any explicit claims about which one they were experiencing). If they didn't know the setup, they would obviously have no motivation to take the bomb in a case where there's a bomb in Left.
But I don't think there's any need for Omega to have simulated the case where the person finds Left empty. I understood the scenario to be that Omega *only* simulates a case where there is a bomb in Left (and the sim version is told exactly the same things about the setup as above), then if the sim takes the bomb, Omega will leave Left empty in the later trial with the biological version, but if the sim pays $100 to take Right, the later biological version will face a trial with a bomb in Left.
The Bomb example stipulates that you know it, so I'm not sure whether this Lottery analogy is apt, because "always play the Lottery when you know you will win" seems reasonable. (That being said, given the probabilities involved, when you read this "predictor's note", you should think that it is almost certainly a fake/prank/illusion, though maybe additional stipulations could fix that)
I agree that the focus on extremely-rare-outcomes seems unfair, though (at least, in the framework of what makes sense to me as an optimisation target)
Note the similarity to transparent Newcomb's Problem--the boxes are transparent, so you know whether or not Omega predicted you would one-box or two-box. Even if you see the box full, and so know that Omega has already predicted you would one-box, I argue that it is important to one-box because that increases the probability you are in this (desirable!) scenario.
The thing that is strange about the lottery case is that discovering that you bought a lottery ticket is normally an undesirable scenario--the money was, in expectation, wasted--but it just so happens that you also won. If you think the lottery was rigged--your friend who works at the lottery commission mailed you the winning tickets--then you might think that actually this is a desirable scenario and it does make sense to play the rigged lottery. But that's not usual!
If your decision is what algorithm to adopt forever, CDT will give the same result as FDT it seems. Then they're just not competing theories. But decision theories are supposed to give answers to what you should do, not what algorithm to set forever. It's similar to objecting to utilitarianism on the grounds that if someone explicitly adopts it as a decision procedure, they'll do bad things--that's not what the theory is trying to answer.
Regarding the bomb case, isn't there a glaring disanalogy to your lottery: In the bomb case you can *see* the bomb! If you *knew* that the lottery would win this time then, yes, obviously you *should* buy the ticket.
Haha, thanks for giving me something to do this afternoon. There's very few things I enjoy more than arguing about decision theory.
A few thoughts:
I agree that the Bomb example is tough - I think FDT's framing that you retroactively change the past is genuinely bad, and the correct way of thinking about it is a lot deeper. (I have an unfinished draft about my own idiosyncratic framing, if you're interested I can send it to you).
But I think in criticisms of FDT, including this one, there is usually a missing mood. There's a mental move very important to intellectual progress of "Okay, I can't get on board with this as it currently stands, but there are some very interesting ideas here that seem important". I think this is clearly true about FDT / LW decision theories! So you should say it, and not call it "devoid of genuine content"! :P
(to be clear, I personally am not doing that mental move since I *am* on board with LW decision theory, but I understand why someone might not be due to the counterintuitiveness)
Also, I disagree that there's "no fact about whether two algorithms are the same". There's a whole literature about "multiple realizability" (often in the context of functionalism in philosophy of mind) that I don't know why MacAskill and you don't mention. Last time I tried to look into this, my personal takeaway was the opposite - that it *does* seem like there plausibly is an in-principle approach for determining whether two physical systems implement the same algorithm: Chalmers' CSA approach. (The concrete response to the calculator example would just be to say that the + and - version of the calculator are the same algorithm with different labels.)
I reject that it's anywhere close to the truth. Maybe the right view is something between CDT and EDT, but beyond that, I don't think FDT is particularly near being correct. Re two algorithms being the same, sure but then you get the problem we talk about where you might just change the economy.
No, because I highly doubt that an algorithm similar to my brain is implemented anywhere in the economy in a consequential way. There's no reason to think this.
I think it's irrational to pay in PH, but you want at the earlier time to bind yourself to do the irrational thing. So it depends on whether you can get yourself to be irrational later.
You should ask an LLM about Chalmers' CSA thing, the bar for two combinatorial state automata to be the same is very high. I found it really insightful. You'd basically need a full simulation of a human brain in the economy somehow, and if you somehow had this, I'm willing to bite the bullet that you can acausally change (to a slight degree) the workings of the economy.
But concretely, if you personally were in the situation, in front of the ATM, would you pay?
I'm not sure, it would just depend on empirical facts about my psychology.
If there is some system of inputs and outputs that corresponds to me defecting in the economy, then it wouldn't seem like my defecting would make any difference to it.
Wait, why do you agree that Bomb is tough? It's basically conditioning on an impossibility--it is free in real life to choose left.
(I have bitten the bullet on retroactively changing the past; I think this is in fact how you have to reason when you exist around other agents, who are reasoning about whether or not you can be convinced. If your partner believes that you will forgive cheating because "it happened in the past and there's no changing that now" then you get cheated on more than if you reason "by having a hard line here, I will have made it less likely that this happened to me.")
How does the age of the universe compare to a trillion trillion seconds?
This question might seem flippant or irrelevant, but I think it's actually pretty important to have a good sense of scale. Like, a lot of being good at decisions hinges on using numbers to mean things, and if you don't believe numbers mean things, you're going to have some incorrect positions.
Suppose rather than facing Bomb!, God flips a coin when creating the universe. If heads, the bomb is live every time; if tails, the bomb is a dud every time. Some agent faces this problem every second for the whole duration of the universe. Could you tell which way the coin landed, just from looking at the outcomes of the decisions?
[Note that I am assuming the 1/trillion trillion chance is real, rather than it secretly being a 1/1 chance that the predictor makes an error.]
Sorry, is English your native language? (The word "basically" is often used to signify rounding--there is a difference between 1 and 0.999999999999999999999999 but the difference is 0.000000000000000000000001, which is a scale which is often discarded, because many concerns will be more important and it's impractical to consider all of them.)
I think if you want to be sufficiently strict with your terminology it's not retrocausality, it's more like common causality. Like, the way 'causality' is often used by CDT proponents is more like "following the propagation of the dynamics of the universe"--I press a button, a voltage flows down a wire, then something happens. When I press the button to cooperate in the psychological twin prisoner's dilemma, there is _not_ a causal connection of this form between my cooperation and my twin's cooperation.
But there's clearly _some_ connection, which you want to have some name for. Maybe it's a logical connection, maybe it's a functional connection, whatever. When my psychological twin is earlier in time--like I'm deciding whether or not to clean dishes out of the sink--then the logical connection is similar in effect to a retrocausal connection. By cleaning the dishes today, what I actually do is make the sink clean tomorrow, not earlier this morning, but that the sink was empty this morning was because I cleaned dishes last night, because I have the disposition to clean dishes when they are in the sink, which is the common cause of cleaning dishes yesterday and today.
One of the main things that will differentiate CDT answers to decision cases and FDT answers to decision cases is that CDT will try to 'defect against itself' in this way. "Well, now that I'm in the city, I don't have to pay for hitchhiking, do I?". FDT doesn't believe that this is an option. "I'm in the city because I'm the sort of person who pays up when I'm in the city. The action that I'm taking now is necessary because it changes the probability of the event that I have already observed." The last sentence is weird! You have to bite some sort of bullet to generate that sentence. Retrocausality is probably not its true name--"dispositional thinking" or "functional thinking" are probably more accurate--but the core way that FDT is winning more points in these cases is by believing in the influence that it has in the decision problem outside of the place that CDT is looking.
[Like, when FDTers look at Bomb, they say "cool, the actual case in front of me is a trivial leaf of the overall decision problem, which barely affects the analysis. I courageously choose Left to save myself $100 almost all the time", whereas CDT says "oh man I will totally pay $100 to not die. The other impacts this decision has don't matter right now."]
Hey Vaniver! :) I like a lot of your work and online presence.
Yeah I used to bite the bullet of retrocausality too, until I realized that there's a better framing that preserves the upside and doesn't have the obvious "incorrectness" of thinking we can change the past: Just say that alternate realities "exist" in some way, and that you might be in a small pocket reality that changes the expected outcomes in the main reality, and that you also care about the you in the main reality. (I believe this is an updateless EDT / UDT thing, although I'm not sure)
Thanks! I think this are either 1) equivalent formulations, and so the question is just what seems weirder to you (which I don't expect to be objective), or 2) we should think carefully about the cases where they differ.
I think there are objective reasons to prefer the alternate reality framing - but also, I do think it changes things because it means the world is much larger, e.g. if we are in something like Tegmark IV. So e.g. ECL becomes more important, the importance of infinite ethics increases, etc.
You might say that alternate realities are also counterintuitive, but I think in some important ways they are not - I have a draft about this I could send you tomorrow, would be curious what you think.
My understanding is that in the causal graph implementation in the paper, you are intervening on nodes that are allowed to be temporally prior to your own action. I'd characterize that as "changing the past".
I fundamentally agree with the criticism about the mathematical impossible world and the vibe but I think there is an even deeper issue here.
At a fundamental level what are we even doing when we adopt a deciscion theory or say you should or should not do something? I mean in a fully literal sense you don't get to make decisions. You will always just follow the laws of physics.
What we are doing is adopting some kind of idealization about the world which -- just like when we define a game formally in mathematics -- idealizes certain things as choices while others are held fixed. And idealizing something as a choice is exactly to treat everything 'before' the choice as unable to depend on the outcome of the choice.
And once you understand things that way the whole Newcomb setup is just kinda non-sensical as regards deciscion theory. It's saying: what if you had a situation where it doesn't make sense to idealize what you are doing using a framework that treats it as a free choice how would you idealize it as a free choice.
The right answer is obviously: don't idealize it as a choice at all. Depending on how you describe the problem you can imagine idealizing in a way where the choice is what rule you precommit to or something like that but the whole debate about FDT or CDT or what have you is just fundamentally confused.
---
I mean just to illustrate how silly the argument is, what if I said the right answer to the paradox was: be someone who is physically guaranteed to take 1 box (so demon predicts you will take 1) and then take both. That does land you in a better position but it's kinda silly because I'm just breaking the rules of the game. Same with treating something both as predictable and a choice.
The rules of the deciscion theory idealization is that you have sometree and each node of the tree represents a choice with earlier ones being able to depend on later ones. If you want to look at scenarios like the demon case you need to reidealize it in a way that obeys those rules -- like a choice between rules.
Totally reasonable to attribute to me. I did coauthor a paper arguing for fdt. I’m just revealing what was in my heart of hearts to get you an accurate count of academics who like it. From a glance, I have a similar reaction of surprise at how many rationalists are fans or think of it as a mature theory. I am also surprised how many are Solomonoff induction stans in case you’re on the hunt for more philosophers v rationalists material.
You should have been at my other talk at Manifest!
First problem, which even the Solomonoff fans admit - the precise prior depends on your choice of universal algorithm. (I'll grant them their response that this only introduces an error up to a single constant, even though in this case the constant appears up front and dominates, unlike in the case of complexity theory, where it disappears in the limit.)
Second problem, which again the Solomonoff fans admit - it's uncomputable (though it is computably-approximable-from-below).
Third problem - it just builds in certainty that the truth is computable! Why think that? Especially if you think that some normatively ideal thing is uncomputable!
Fourth problem (which I think is the most important one, though lots of academic epistemologists face as well) - what motivation is there for someone who has a different prior to treat this prior as better? Omniscient priors always perform best in the worlds they are adapted to, and in general, if you throw someone into situations in proportion to a particular probability function, then them the prior that matches this probability function will be the most successful prior for them to have. If someone isn't being thrown into the world in proportion to the Solomonoff prior, why should we fetishize computational simplicity in precisely this way?
As far as I can tell, the motivation for the Solomonoff prior comes from people who are impressed by Occam's razor, but instead of trying to justify it, they reason in a Kantian transcendental way to figure out what constraint rationality would have to have to make Occam's razor automatically fall out.
> I know I’m not a simulated algorithm. The simulated algorithm isn’t conscious (we can stipulate). I am.
But you don’t know that. Especially not now that you’ve posted that!
Consider the most mundane form of simulation possible: human imagination. Suppose I set up a Newcomb experiment where I make my predictions by simply reading what the participant has written online and imagining what their thought process would be like. The prediction won’t be very accurate but it will probably be better than chance. Better than chance is all you need to create the paradox.
Well, now that I’ve read your post, my little imaginary version of you is definitely going to start with, ‘I know I’m not a simulated algorithm. The simulated algorithm isn’t conscious, I am.’ But imaginary-you is quite mistaken. Imaginary-you is not conscious. And when imaginary-you two-boxes, it costs the real you the prize. Hypothetically.
> Decision theory generally assumes that you’re self-interested. But if I’m the algorithm, then I care about algorithm me—not the version in the real world. So then I wouldn’t care about what the output of the algorithm was.
I think this is correct. If you are purely self-interested, then two-boxing is rational. However, if you are a simulation, you don’t lose anything by one-boxing because the simulation almost certainly terminates as soon as you make a choice. Therefore, you should one-box if you have any goodwill towards your real-world counterpart. Either because they are “kind of like me but different in a bunch of respects”, or just because you have a default of goodwill towards other people.
On the flipside, if you value your life as a simulated being and resent that the simulation will terminate, then you might rationally two-box as revenge. This presumably doesn’t apply to low-fidelity simulations like human imagination.
(I don’t know whether my opinion comports with FDT. As I see it, the simulation argument is an argument *against* FDT. It shows how plain old CDT can justify one-boxing.)
The fundamental problem with this whole debate is that what it *means* to idealize Newcomb's problem in terms of decision theory as a choice by the predictor and then a choice by you just IS to suppose that the predictor's choice can't depend on your choice in any way.
As such there aren't different decision theories. There is one correct answer: it is be the sort of person who would be predicted to take one box then take two.
If someone objects that you broke the assumption respond that no, that's what it means to idealize something as a subsequent decision in decision theory.
If they prefer, they can choose to model the situation instead as consisting of a deciscion precommitting the agent to be a 1 or 2 boxer and then a choice by the predictor with no subsequent decision. That problem also has an obvious answer.
The one thing that doesn't make sense is to idealize what you do with the boxes as a decision and then not treat it as such. That's conflating the formal idealized notion of decision in the theory with the notion of a decision as happening any time someone goes "Hmm, what should I do".
As such there aren't different decision theories AT all. There are different attitudes about how to idealize Newcomb's problem in terms of decision theory. But once you've done that it reveals there isn't really a philosophical problem -- just a question about how to think about the scenario.
I think that is a very neat way of putting the point. And -- after I wrote a much too long explanation of why the whole CDT/FDT/etc debate is confused -- I think this really gets to the heart of why that debate is confused.
The debate over which deciscion theory is correct collapses down the distinction between the questions of: should I be the sort of person who 1 boxes and if a miracle were to occur after the prediction that freed me from having to follow the laws of physics what should I want the outcome of that miracle to be.
The right thing to do is just to use which question you are trying to answer to idealize the problem in the appropriate way for that question. The apparent tension only appears bc of the false presumption that there is only one question one could be answering.
Does your verdict that "the rational act is two-boxing" actually guide the decision procedure of a rational agent?
If it does, then the predictor predicts two-boxing and the rational agent loses the million. Also, "rational sort of person" and "rational act" are the same thing after all.
If it doesn't - which I suspect - then the predictor predicts one-boxing, but... The agent is somehow built to ignore what she finds rational as a person? Isn't it impossible to build an agent that generally views one-boxing as rational, but somehow still two-boxes? In any case, it's highly suspect - paradoxical even - that the ideal agent's action directly goes against its own character.
You are assuming a lot in assuming that there even *is* a decision procedure. Most people don’t use anything like what we would call a “decision procedure” in guiding most of their behavior. The ideal agent one-boxes; the ideal act is two-boxing; the ideal agent doesn’t do the ideal act.
I think that's an implicit assumption in Newcomb's Problem, not my assumption. Either way, while humans are probably far from consistent in their decision making, there is *some* way they make decisions. That is their decision procedure.
If it’s an assumption about Newcomb’s problem then that version of Newcomb’s problem isn’t about rationality! Rationality is about doing what works, not about following procedures. (Procedures sometimes help, but that’s a contingent fact about some kinds of decision problems.)
But it seems you and I have a different idea of what a "decision procedure" is exactly. I think your view that the ideal agent one-boxes and the ideal act is to two-box is incoherent. Would love to debate you more on this
Aside from the nugget of stipulating non-consciousness of a simulated algorithm (which... What.)... let's abstract a bit. You are running a very weird comparison between the following three:
1. Causal Decision Theory
2. "EDT and updateless EDT" (two quite different theories; the latter is, essentially, Wei Dai's Updateless Decision Theory, or UDT, https://www.lesswrong.com/w/updateless-decision-theory, and thus _also_ partially grows from the Rationalist tradition, though Wei Dai appears to _also_ be more classically read-up!)
3. FDT
So, according to you, Rationalists are wrong because they deliver FDT; and yet the preferable alternative is UDT, which is done by a Rationalist and, historically, because of FDT-like exploration.
But it gets weirder. According to https://www.lesswrong.com/w/timeless-decision-theory, "[t]he FDT paper thus describes a general framework which remains agnostic about an updateless approach (like UDT) vs an updateful one (like TDT), but which sticks close to the logical-causality approach introduced by TDT." TDT stands for Timeless Decision Theory, Yudkowsky's _previous_ attempt at formalizing his gripes with CDT (and is itself barely formalized). If this is true, then _of course_ FDT is underdefined under the terms you offer, because it is at a different level of abstraction! It covers a _family_ of approaches!
>You face two open boxes, Left and Right, and you must take one of them. In the Left box, there is a live bomb[...] The Right box is empty, but you have to pay $100 in order to be able to take it.
>A long-dead predictor predicted whether you would choose Left or Right, *by running a simulation of you and seeing what that simulation did*. If the predictor predicted that you would choose Right, then she put a bomb in Left. If the predictor predicted that you would choose Left, then she did not put a bomb in Left, and the box is empty.
>*The predictor has a failure rate of only 1 in a trillion trillion. Helpfully, she left a note, explaining that she predicted that you would take Right, and therefore she put the bomb in Left.* [...] What box should you choose?
First problem with the scenario: as stated, the note actually contains no information about the bomb. How so? Well, the example is clear: the predictor predicted your choice *by running a simulation and seeing what the simulation did*. In order to predict you via simulating you, both you and the simulation need to have the same inputs. Thus the simulation would have to have seen an identical note to you.
The problem is that the note and its contents have to be decided *before they are shown to the simulation*. Because the Left or Right decision is not known before the simulation is run. And once the simulation is run with that particular note, then you have to see that note as well (identical inputs). So you will see the note that was decided upon before the prediction was made. Its content is thus immaterial to the presence or absence of the bomb.
If we want the example to work, it needs to be redefined first. We need an example that (a) makes the note true, (b) makes the predictor as accurate as stated, and (c) forces FDT to Left in this example. *It’s not clear that such an example exists*.
Partial example: there is a fixed-point version of this that gives (a) and (b), but it is no longer a problem for FDT. Assume the predictor ignores the simulation and always puts the bomb and the note in, and the note is known to be accurate. Then all agents will go Right, including FDT, because they believe the note. And the predictor, who predicted Right, will thus be extremely accurate.
>First problem with the scenario: as stated, the note actually contains no information about the bomb. How so? Well, the example is clear: the predictor predicted your choice *by running a simulation and seeing what the simulation did*. In order to predict you via simulating you, both you and the simulation need to have the same inputs. Thus the simulation would have to have seen an identical note to you.
Does the mechanism of prediction really need to be simulation? Are we allowed to imagine there's just some magic oracle that can accurately (but not perfectly) predict what you'd do without any simulating, so there were never any previous notes to begin with?
If not, here's another variant. Suppose time is past-eternal. Every year for all of the eternal past, the simulator has gone through the bomb experiment with a copy of you. The note in the present includes the following: "In every contiguous sliding window of yearly experiments of size 1 trillion up till now, you've chosen Right way over 99% of the time, so I've put the bomb in Left; I am programmed to only predict your behavior and fill up boxes based on the most recent such window, and no other. But since this statistical pattern of outcomes has in fact been universal across windows, the note I've left in each previous experiment has always said the same thing as what I'm telling you now." Everything else is the same as the original thought experiment. Now the simulator's notes are and have always been totally honest.
>Does the mechanism of prediction really need to be simulation? Are we allowed to imagine there's just some magic oracle that can accurately (but not perfectly) predict what you'd do without any simulating, so there were never any previous notes to begin with?
That doesn't necessarily gain you anything. If it's predicting you, it has to predict what you would do *if you saw that specific note* (otherwise it's not predicting what you will do, but what you would do in a different world).
Let's think some more. Suppose that we have two notes: “I put the bomb in Left”, “I didn't put the bomb in Left”, and the predictor runs both predictions. Now if you go Left in the first case and Right in the second case, the note is always inaccurate, so the predictor must include an inaccurate note.
Let's add another option: no note. We can imagine something like this: the predictor will run the prediction with the “bomb-note”. If that's consistent, it uses that. If not, it runs the prediction with the “no-bomb note”. If that's consistent, it uses that. If that's also inconsistent, it just predicts the “no note” and goes with that. So this predictor always reaches a result.
Then if you’re an FDT agent, you go Left, which makes the “bomb-note” inconsistent and the “no-bomb note” consistent. So you expect to see the “no-bomb note”.
Ok, I think we’ve got a working version. As an FDT agent, you see the “bomb-note” in the real world only if there’s a mistake (1 in a trillion trillion).
>If not, here's another variant. [...] Now the simulator's notes are and have always been totally honest.
That’s true, but note how it works: it essentially forces consistency via threats. Simplify to using just the previous single experiment. Then it’s saying “because you went Right last time, I put a bomb in the Left. Do you want to go Right this time?” to which the answer is obviously “yes, of course”.
>That’s true, but note how it works: it essentially forces consistency via threats. Simplify to using just the previous single experiment. Then it’s saying “because you went Right last time, I put a bomb in the Left. Do you want to go Right this time?” to which the answer is obviously “yes, of course”.
OK, but I don't understand how this addresses the worry. The simulator's choice(s) on whether to install the bomb in Left is still a function of you and a large number of perfect or near-perfect psychological duplicates' choices. If you reason as if you were a single abstract algorithm simultaneously in charge of all your symmetric duplicates' choices, you'll pick Left, and the simulator won't install the bomb. If you pick Right, then you aren't reasoning as if you were a single abstract algorithm simultaneously in charge of all your symmetric duplicates' choices. So how is that still in the vicinity of FDT?
The problem is that the predictor's belief that the FDT agent will go Right is self-confirming. If the predictor has that belief, and everyone knows it, then it puts a bomb in the Left box and the FDT agent goes Right - which makes the original belief correct.
In the simulation scenario, the FDT agent knows they can change this belief through their actions in the simulation. So then they choose Left, which corrects the predictor's belief.
However, in your scenario, the FDT agent is stuck in a poor equilibrium: it can't change the self-confirming "Right" belief through its simulated actions. So the question becomes: how did it get into that poor equilibrium?
You stipulated that the note was correct when it said "In every contiguous sliding window of yearly experiments of size 1 trillion up till now, you've chosen Right way over 99% of the time [...] I am programmed to only predict your behavior and fill up boxes based on the most recent such window, and no other."
To me, that seems like you said "let's stipulate that the FDT agent starts in a poor equilibrium".
(Note that if the agent is altruistic towards its future duplicates - and there are enough future duplicates to make it worthwhile - it will get out of the equilibrium by choosing "Left" and burning, time and time again, until the predictor gets the message and shifts to predicting Left. Consequently, if the FDT agent is altruistic towards its future duplicates, your scenario becomes impossible - it would never happen in the first place)
>However, in your scenario, the FDT agent is stuck in a poor equilibrium: it can't change the self-confirming "Right" belief through its simulated actions.
What do you mean? If (say) all the infinite prior iterations of the agent choose Left, then the simulator's predictions will always have been that the agents' choice will be Left, so the bomb will never be installed in Left, including in the present. So the prior simulated iterations do affect the simulator's prediction/decision.
>To me, that seems like you said "let's stipulate that the FDT agent starts in a poor equilibrium".
Whether this criterion is relevant to my version of the thought experiment or not, it's unclear to me how it connects to the definition of FDT, which as far as I know doesn't directly mention being in some sort of bad equilibrium (an extremely general term), but instead mentions something a bit more specific like "what would the best thing for my decision function to output, given that it's running elsewhere for a lot of other agents." Isn't it the case that if all the (past/present/future) duplicates running copies of my decision algorithm choose Left, it would be better for all of us? If the answer to that question is "yes," but then we should nevertheless choose Right in the present, then what is the actual nuts-and-bolts definition of FDT?
Working on a more substantive response, but, for the moment, note that the first counter-example is incorrect (FDT two-boxes in both cases):
>Imagine that there’s some gene that correlates 99.9% with two-boxing. The gene is not caused by two-boxing, they just perfectly correlate. Now imagine two different scenarios:
>1. The predictor looks to see if you have the gene. If you don’t, they put $3,000 in the first box. If you do, they put nothing in the first box. The second box has $1,000. Should you take both boxes?
>2. The predictor runs a simulation of you with 99.9% accuracy. The cases where the simulation is inaccurate are the same as the ones where there isn’t an overlap between your gene and which box you take. Thus, there is 100% overlap between the predictor’s judgment in this case and the last. The only difference is that in the last case, they look to see whether you have the gene, while in this case, they run a simulation of you. If they guess that you’ll one-box, they put $3,000 in box one, while if they guess you’ll two-box, they put nothing. The second box has $3,000. Should you take both boxes?
>Here FDT’s answer is that you should two-box in the first case but not in the second case.
That is incorrect; FDT two-boxes in both cases. The key is this line in 2.: **The cases where the simulation is inaccurate are the same as the ones where there isn’t an overlap between your gene and which box you take.** This gives you the power to break the predictor via your decision. If you have the gene and one-box, you break the predictor (hence it predicts you will two-box). Otherwise, you will two-box and it will predict that via simulation. The opposite happens if you don’t have the gene (it will always predict you one-box). So FDT notes that the predictor’s behaviour depends on the gene only, and thus two-boxes.
I'm not sure if this works or is even clear, but: Maybe you can solve this by distinguishing between "meta-dispositions" such as FDT and "material dispositions" that factually determine what you decide to do.
As I understand it the general idea behind FDT is something like: “Whenever you face a decision, imagine that you choose the disposition/algorithm you have, that determines what you decide for each and every action up to that point. Determine what the disposition is that maximises the combined value of all your decisions, then act according to this disposition".
If you say the optimal disposition to have is just to be an FDT-follower, that's probably an infinite regress. But if you can say in the decision process you need to determine the optimal "material disposition", that is, any disposition that factually tells you what to do (which FDT does not, I think this is also what you say?), like "do the thing that causes the most utility" or "always keep your promise to someone who saved you".
If you then want to maximise your utility, you usually get the result act according to the material disposition "do the thing that causes the most utility" (which is just CDT I think?). But if for whatever reason you believe that you can gain a lot of utility by having a different material disposition, you act out of that disposition. This is the case in Parfit's Hitchhiker, because there the driver scans the disposition you have, and only saves your life if the material disposition you have is one that causes you to pay him later.
Ok, but the disposition you have and that the driver scans is not a material disposition at all, but this weird "meta-disposition" that causes you to pick a disposition later. But, because the driver knows that when it's time to pay you will pick that material disposition that would have maximised your utility up to that point, he knows that you will pick a material dispositon that causes you to pay him. So, regress avoided?
So clearly I need to brush up on my decision theory, but let me get this right.
The issue is that 1. FDT claims you should
Make the decision that basically if infinitely repeated in all similar scenarios as the output of your decision procedure would leave you best off. 2. So the problem is that it leads to weird stuff such as
A. It matters what other algorithms relationship is to your algorithm such as the case where it you would pick the bomb because you having that as your decision output would counter factually mean there was no bomb even though there is?
Plus all sorts of similar issues where we have to basically try to explain why a scenario where ~you did things only ~you would ever do is related to this situation.
B. Anytime an algorithm is introduced you need to show what the relationship between your algorithm and this 3rd party algorithm (I know there are only 2 parties), but it seems impossible to establish anything beyond an epistemic one so you really have EDT and even if you could it would lead to weird examples like the gene one.
I'm not sure who the relevant experts are. The main use of any of these theories that I can see, admittedly from my position deep inside the rationalist community, is the kind of AI safety research that MIRI used to do, the kind that assumes you have to solve everything in one go with no feedback or we all die. And even if you think that that research direction is correct (which I am not asserting), they themselves aren't really responding to empirical evidence so I'm not sure how one would evaluate whether they are good experts.
I do think you are refusing to engage seriously with Newcomb's problem, in the way that non-philosophers often refuse to engage seriously with thought experiments. The whole point of Newcomb's problem is that Omega isn't 99.9% accurate in its modeling of you, it is 100% accurate, it cannot be wrong. And in that scenario, 1-boxing seems obviously correct to me. The usefulness of Newcomb's problem, for me, is (1) recognizing the asymptotic behavior that 1-boxing (or more generally being honest in social and business interactions) is correct where your counterparty can perfectly predict you, and (2) highlighting that other people are often pretty good at predicting us in relevant ways, and so in any particular interaction we should consider that we may be very close to that asymptote.
In your hypothetical about cutting off your leg because the prediction of that was the condition that caused you to be woken up, again, my intuition strongly says that cutting off your leg is the correct action there. I understand how an unthoughtful person could say otherwise, I understand the aversion, I'm not entirely confident I'd have the courage to go through with it, but I don't see how a thoughtful reflective person would endorse the position of not cutting off their leg. (You're calling it "for no benefit" kindof contradicts the scenario you set up.)
What originally brought me over to this FDTish way of thinking (I call it FDTish because I have not gone through the rigorous math and do not see a reason to) was the point that I don't know if I am the simulation or not. Again, this may not reflect reality perfectly, but it is a good asymptotic case to think about that for practical purposes we may be close to. I don't know how you can just assert that the simulation isn't conscious. Nobody ever specifies that in any of these hypotheticals, and it is not the natural interpretation. I don't think it is even logically possible to create a simulation that always outputs the correct behavior (meaning the behavior of the person being simulated) without the simulation being conscious. Even on your version of dualism, with its unnatural attachment between brains and minds, why would you expect it to be possible, much less a default, for a simulation that makes perfectly accurate predictions not to be conscious?
//The whole point of Newcomb's problem is that Omega isn't 99.9% accurate in its modeling of you, it is 100% accurate, it cannot be wrong.//
This is not a standard stipulation of Newcomb,. It is generally stipulated to be only 99.9% accurate.
I call it "for no benefit" because at the time you take the decision you know you don't benefit.
Maybe in practice the simulation would be conscious. But we can imagine a scenario where it isn't conscious. That's all we need for the counterexample.
"Maybe in practice the simulation would be conscious. But we can imagine a scenario where it isn't conscious. That's all we need for the counterexample."
But it isn't a counterexample! If whether or not you are conscious is a relevant input for making the decision, then the predictor just inputs the right value of this variable ("is conscious") to simulation-you. If she doesn't, then the simulation isn't accurate, and, once again, there is no subjunctive dependence.
> Maybe in practice the simulation would be conscious. But we can imagine a scenario where it isn't conscious. That's all we need for the counterexample.
I can imagine that there are three positive integers such that a^3 + b^3 = c^3. I think you need something a little stronger than what a human mind can imagine.
We normally think it's okay to, for the purpose of thought experiments, stipulate views even if you think those views are wrong. E.g. the following seems true: if deontology is true, you shouldn't kill one to save five.
Normative views, or views we think are empirical false, sure. Not views we think are logically contradictory, not unless we are trying to do a proof by contradiction. That was the point of my comparison to Fermat's last theorem.
I'll also note that you are exhibiting the same flaw I see in Chalmers' zombie argument. He just expects us to accept that a p-zombie is logically possible, and he gives no good reason for us to think that, and my guess is that it isn't.
Wouldn't it be odd if consciousness somehow depended on the substrate? Why would we expect that to be the case? I don't claim to have a proof that it's not logically possible for consciousness to depend on substrate. For that we'd need a formal specification of what consciousness is, and we don't have that yet. But I'd bet that once we get there, it will turn out that consciousness can, at least in principle, be implemented on many different substrates. At the very least, if you want to use it as a premise for your arguments about decision theory, the burden is on you to show that it is possible for such an accurate simulation to not be conscious.
"I call it "for no benefit" because at the time you take the decision you know you don't benefit."
What you're missing is that your simulation "knows" this just as much as you - the predictor just makes the simulated environment like that. If not, then the predictor hasn't made a good simulation - more importantly, there is no subjunctive dependence and FDT wouldn't recommend cutting of your leg.
So your leg cutting argument is just a strawman - it doesn't attack FDT.
... No? You're just stipulating that the agent has relevant input the simulated agent does not - which, if true, would break subjunctive dependence. So, again - strawman.
But the algorithm can be the same as the one you use to make decisions even if it's not conscious. E.g. if the bubble sorting one was conscious but the other wasn't.
The algorithm itself can be, but you use "I am conscious" as input for your algorithm. That is, your algorithm uses an input variable called "isConscious" that is said to "true". (This is apart from the algorithm actually being conscious; it's just an input variable you use.) The predictor will just set this variable to "true" when simulating you as well. So you and your simulation still have the exact same information regarding the problem at hand, and you can't use the fact that you are conscious to know that you are not the simulated version of you. (It couldn't be otherwise! If you could there would be no subjunctive dependence.)
"But what could it possibly depend on? Isn’t it obvious that there’s no single privileged joint-carving way to decide the similarity of algorithms that doesn’t just look at the statistical correlation between their outputs? Certainly FDTers owe us some account of how this works. It doesn’t do to call it an unsolved problem, when this is the entire engine of the theory—when there’s no plausible story of what a solution would even look like, strong active reason to think there is no such solution, and a solution is needed for the theory to give any result in any case."
You're pointing at a real problem here, but are overstating your case to the extreme. Clearly, there are algorithms for which we can prove they give the same result - e.g., quick sort and bubble sort. And quick sort that has a 1 in a 100 probability of swapping 2 numbers after its normal procedure is done is very similar to bubble sort, etc.
So I think you're correct that this is an open problem, but it's not that we have absolutely no idea.
First, deciding whether two algorithms in fact implement the same function is one of the classic undecidable problems. (If you could solve it, then you could use it to solve the halting problem.)
But even if you know that two algorithms implement the same function, there's a lot of important uses of "same algorithm" for which you don't want to count them as the same! (If you just consider different sorting algorithms, you can see that some of them have different runtimes, and some of them are robust against different types of errors in calculation, and so on.)
But more importantly, for FDT to do reasonable things, it needs to be able to classify the outputs of one algorithm as in some sense being counterfactually dependent on the outputs of another, even when they don't have all the same outputs, and when neither calls the other as a sub-routine. This is going to be a much more difficult question to answer in plausible ways than just saying when two algorithms that implement the same function are the same or different!
It sure is undecidable, but that doesn't mean you can't ever determine it in practice. (Given your next paragraph you seem to agree with this.)
I agree things like runtime are important, but at least runtime seems straightforward enough to solve, right?
I fully agree with your last paragraph. This is a real problem; I just don't see a reason to believe it's an impossible challenge. (Not saying you claim this.)
I'm confused about the bomb example.
It seems like your example hinges upon the FDT agent picking Left.
But it also says that the predictor with a one in a trillion trillion error rate predicted you would pick Right.
If all of this is correct, it seems like you're hinging your example on this being the one time in a trillion trillion when the predictor was wrong.
But decision theories shouldn't be judged on whether they work well in unbelievably rare edge cases that you would never encounter in a million lifetimes.
Compare Lottery Decision Theory:
> you have the choice to buy a lottery ticket for $100. There is a one in a trillion trillion chance you will win. If you win you get $1 million. Should you do it? Before answering, keep in mind that, unbeknownst to you, this is the one time you would actually win.
We can use this example to prove that you should definitely play the lottery. I think the bomb situation maps to this - although it fails in the one/trillion trillion case where the superpredictor was wrong, it succeeds in the 9999999999.../trillion trillion cases where the superpredictor is right, and when you multiply out the probabilities by the utilities (eg of getting any extra $100 vs. getting bombed), FDT gets you more utility overall.
The particular way FDT succeeds is that you never (okay, 1/trillion trillion times, but this rounds to never) find yourself in this situation. So just by asking about this situation, you've already started with the assumption that FDT fails, which is why you are so easily able to prove that FDT fails.
I think this goes back to what I said last time we discussed this. You and Eliezer are optimizing for different things. You are optimizing for never finding yourself in a situation where you have to do something silly within that situation; he is optimizing for having the most utility overall if you can set your algorithm. I think the thing he's optimizing for makes more sense.
Yes you are depending on this being the one time in a trillion trillion when the predictor was wrong. Decision theories are judged by whether they give the right answer across the board. So if they get the wrong result even in weird edge cases, they're disproved. But the wrong result isn't the same as any case where adopting the situation is bad for you. So what matters is not "if they work well" but if they give the right answers. Compare: if there's a single result where utilitarianism gives the wrong answer to what you should do, that disproves it. This doesn't mean that any scenario where being a utilitarian turns out badly is a counterexample.
Decision theories are theories of rational action. That's different from what dispositions you want to have. If you're in a world where you encounter Newcomb's problems all the time, then you want have one-boxing dispositions. But that's different from a theory of rationality.
Also, just want to register, I think the bomby problems are actually much less big of problems than the "FDT doesn't say anything in any case" problems. I find the discourse has focused more on cases where FDT is counterintuitive and has weirdly neglected the fact that it's totally ill-formed.
It's fine for theories to give the wrong answer in a probabilistic sense, i.e., they recommend some action under uncertainty and you end up getting unlucky and losing out. I think what you want to say with the bomb case is that it gives the wrong answer under certainty.
Looking through that thread, the FDT people might want to say there really is uncertainty in the bomb case - namely, which person you are among the real you and your simulations. But I don't feel the introducing simulation aspect is actually necessary to the example.
"Decision theories are theories of rational action. That's different from what dispositions you want to have."
The fact that, under FDT, these are not different at all counts as a major point for FDT over e.g. CDT.
>Yes you are depending on this being the one time in a trillion trillion when the predictor was wrong. Decision theories are judged by whether they give the right answer across the board. So if they get the wrong result even in weird edge cases, they're disproved.
That's a judgement call. FDT proponents would point out that the FDT agent is getting a higher utility - paying a cost in one branch to get a bonus in every other one. So they could argue that CDT is failing in almost every branch - deciding to lose utility, thus "disproving" CDT.
What I feel is really happening is that predictors make decision theory setups into game theory ones. CDT clings to decision theory while FDT accepts (some) game theory.
In fact, "bomb" maps well to an extortion game theory setup. Predictor is claiming that there is a bomb in the house, which will go off automatically unless you wire $100 to their account. Your actions are "pay or decline". Their actions are "make bomb or bluff". They get a penalty is the bomb actually goes off.
They are capable of predicting you, and will only make the bomb if they predict that you will pay up. Then if you are CDT, you will give in to extortion (pay), and hence the predictor will make the bomb. If you are FDT, you will resist extortion (decline) and hence the predictor will not make the bomb.
And, in one in a trillion trillion cases, the predictor will make a bomb for an FDT agent that will decline to pay - BOOM.
> Decision theories are judged by whether they give the right answer across the board.
In a world of uncertainty, it is generally not the case that you can win in every possible outcome. The question is whether you can get the best outcome in the aggregate.
> So if they get the wrong result even in weird edge cases, they're disproved.
Guaranteed Payoffs gets the wrong answer on whether or not you should take a bet on a fair coin which pays out $2 if you win and $1 if you lose.
> Also, just want to register, I think the bomby problems are actually much less big of problems than the "FDT doesn't say anything in any case" problems. I find the discourse has focused more on cases where FDT is counterintuitive and has weirdly neglected the fact that it's totally ill-formed.
Have you seen any decision problems where a FDT proponent has not been able to give the FDT answer?
Like, I think we can exhibit problems where it's challenging, or requires empirical estimation. ("How correlated is my decision to vote with other voters in my district?") But I think this is a challenge with the world, not with the decision theory, and generally it's easy to come up with the rule. ("If my influence on whether other voters similar to me vote is at least 0.02, then it makes sense for me to vote.")
Wait that’s obviously not right about the coin flip case. Guaranteed payoffs only applies in cases where you’re not uncertain
Sorry, I should've linked to my longer description; I'm talking about the case from my comment on MacAskill's post: https://www.lesswrong.com/posts/ySLYSsNeFL5CoAQzN/a-critique-of-functional-decision-theory?commentId=mR9vHrQ7JhykiFswg
It models the bet as a two-stage decision. First you choose whether or not to make the bet (obviously yes, it's positive EV) and then once you know the result, the loser gets to decide whether or not the bet was serious or a joke (by paying up or backing out). Your counterparty is a predictor who will take the bet seriously if they think that you will take the bet seriously.
Guaranteed Payoffs recommends backing out if you lose (after all, it's worse to pay up than not and the coinflip has already not gone your way), but this means the counterparty doesn't view the bet as serious and so they don't pay you if they lose, which means you lost the opportunity to make those sorts of bets.
Yes, it is old news that CDT sometimes leads a person to have unfortunate dispositions.
I feel like this is not a "weird edge case"? Or, like, I still don't know how to interpret what you mean by "So if they get the wrong result even in weird edge cases, they're disproved." unless it's something like "all decision theories are disproved". (But if all decision theories are 'disproved', why do you care about 'proof'?)
I don't think your second paragraph here works as an argument. Here is an attempt for a structurally similar argument, which is obviously wrong.
Ethical theories are theories of right action. That's different from what actions are fortunate. If you find yourself in a trolley problems, then it would be fortunate for you to pull the lever. But that is different from it being right. (And so a theory of "fortunate" action is not a theory of right action)
I believe the bomb example doesn't work (or at least not as stated). As stated, the note is irrelevant and uninformative about the bomb, because it has to be presented to the simulation before the simulation makes a decision, and then the exact same note will have to presented to you, whatever your simulation does.
https://benthams.substack.com/p/functional-decision-theory-not-even/comment/285235267
It's not clear that there is an example where a) the note is accurate, b) the error rate is as stated, and c) the FDT agent is forced Left in this case.
EDIT: after thinking about it, you can have the following setup: the predictor runs with "bomb on Left note" included. If that results in a self consistent action - if you would go Right in that prediction - then they do that. If the action is not self consistent, they run the prediction without a note and go with that (obviously not including a note).
If they value $100 more than they disvalue a 1/10^{24} chance of a bomb, then an FDT agent will go Left in the "bomb on Left note" situation (making it inconsistent) and also Left in the "no note" situation. So "in the real world" it only expects to see no note. So seeing the note in the real world means that the predictor has made an error. The example is coherent (I disagree with its implications, but the example is coherent).
I agree the note seems like a spurious add-on since for the thought-experiment to make sense, the person has to already know the full setup of the problem before making a choice, including the fact that the decision whether to put a bomb in Left depended on an earlier simulation (and the original simulation must have been given the same sensory input that the "real" biological version will later get if a bomb is put in Left, i.e. the simulation was lied to and told they were experiencing a real trial that was the result of a prior simulation--or I suppose they could both just be told that the bio version's setup will be the result of the sim version's response to the bomb setup, without any explicit claims about which one they were experiencing). If they didn't know the setup, they would obviously have no motivation to take the bomb in a case where there's a bomb in Left.
But I don't think there's any need for Omega to have simulated the case where the person finds Left empty. I understood the scenario to be that Omega *only* simulates a case where there is a bomb in Left (and the sim version is told exactly the same things about the setup as above), then if the sim takes the bomb, Omega will leave Left empty in the later trial with the biological version, but if the sim pays $100 to take Right, the later biological version will face a trial with a bomb in Left.
Thanks! Have used that phrasing in the second part of this: https://www.lesswrong.com/posts/SdGbWkCZgCN7EGBxM/pragmatic-fdt-and-predictors-as-game-theory-1
> unbeknownst to you
The Bomb example stipulates that you know it, so I'm not sure whether this Lottery analogy is apt, because "always play the Lottery when you know you will win" seems reasonable. (That being said, given the probabilities involved, when you read this "predictor's note", you should think that it is almost certainly a fake/prank/illusion, though maybe additional stipulations could fix that)
I agree that the focus on extremely-rare-outcomes seems unfair, though (at least, in the framework of what makes sense to me as an optimisation target)
Note the similarity to transparent Newcomb's Problem--the boxes are transparent, so you know whether or not Omega predicted you would one-box or two-box. Even if you see the box full, and so know that Omega has already predicted you would one-box, I argue that it is important to one-box because that increases the probability you are in this (desirable!) scenario.
The thing that is strange about the lottery case is that discovering that you bought a lottery ticket is normally an undesirable scenario--the money was, in expectation, wasted--but it just so happens that you also won. If you think the lottery was rigged--your friend who works at the lottery commission mailed you the winning tickets--then you might think that actually this is a desirable scenario and it does make sense to play the rigged lottery. But that's not usual!
It's important to play 'follow the improbability' ( https://www.lesswrong.com/posts/k6EPphHiBH4WWYFCj/gazp-vs-glut ) and notice when the hypothetical involves rigging and whether or not that is 'legit'.
If your decision is what algorithm to adopt forever, CDT will give the same result as FDT it seems. Then they're just not competing theories. But decision theories are supposed to give answers to what you should do, not what algorithm to set forever. It's similar to objecting to utilitarianism on the grounds that if someone explicitly adopts it as a decision procedure, they'll do bad things--that's not what the theory is trying to answer.
Regarding the bomb case, isn't there a glaring disanalogy to your lottery: In the bomb case you can *see* the bomb! If you *knew* that the lottery would win this time then, yes, obviously you *should* buy the ticket.
"You are the only person left in the universe. You have a happy life" only rationalist could write a sentence like this
Also, what do I care about $100 if there's nobody else in the universe?
You just do!
Scrooge Mc Duck, last surviving descendant of Earth...
Haha, thanks for giving me something to do this afternoon. There's very few things I enjoy more than arguing about decision theory.
A few thoughts:
I agree that the Bomb example is tough - I think FDT's framing that you retroactively change the past is genuinely bad, and the correct way of thinking about it is a lot deeper. (I have an unfinished draft about my own idiosyncratic framing, if you're interested I can send it to you).
But I think in criticisms of FDT, including this one, there is usually a missing mood. There's a mental move very important to intellectual progress of "Okay, I can't get on board with this as it currently stands, but there are some very interesting ideas here that seem important". I think this is clearly true about FDT / LW decision theories! So you should say it, and not call it "devoid of genuine content"! :P
(to be clear, I personally am not doing that mental move since I *am* on board with LW decision theory, but I understand why someone might not be due to the counterintuitiveness)
Also, I disagree that there's "no fact about whether two algorithms are the same". There's a whole literature about "multiple realizability" (often in the context of functionalism in philosophy of mind) that I don't know why MacAskill and you don't mention. Last time I tried to look into this, my personal takeaway was the opposite - that it *does* seem like there plausibly is an in-principle approach for determining whether two physical systems implement the same algorithm: Chalmers' CSA approach. (The concrete response to the calculator example would just be to say that the + and - version of the calculator are the same algorithm with different labels.)
I reject that it's anywhere close to the truth. Maybe the right view is something between CDT and EDT, but beyond that, I don't think FDT is particularly near being correct. Re two algorithms being the same, sure but then you get the problem we talk about where you might just change the economy.
No, because I highly doubt that an algorithm similar to my brain is implemented anywhere in the economy in a consequential way. There's no reason to think this.
Have you looked at Parfit's Hitchhiker? (https://x.com/reconfigurthing/status/2031963649124818967) Do you pay?
Yes but if it was!
I think it's irrational to pay in PH, but you want at the earlier time to bind yourself to do the irrational thing. So it depends on whether you can get yourself to be irrational later.
You should ask an LLM about Chalmers' CSA thing, the bar for two combinatorial state automata to be the same is very high. I found it really insightful. You'd basically need a full simulation of a human brain in the economy somehow, and if you somehow had this, I'm willing to bite the bullet that you can acausally change (to a slight degree) the workings of the economy.
But concretely, if you personally were in the situation, in front of the ATM, would you pay?
I'm not sure, it would just depend on empirical facts about my psychology.
If there is some system of inputs and outputs that corresponds to me defecting in the economy, then it wouldn't seem like my defecting would make any difference to it.
You're not sure what you would do in practice? Hmm can you imagine being transported there right now and having to make the decision?
Wait, why do you agree that Bomb is tough? It's basically conditioning on an impossibility--it is free in real life to choose left.
(I have bitten the bullet on retroactively changing the past; I think this is in fact how you have to reason when you exist around other agents, who are reasoning about whether or not you can be convinced. If your partner believes that you will forgive cheating because "it happened in the past and there's no changing that now" then you get cheated on more than if you reason "by having a hard line here, I will have made it less likely that this happened to me.")
How is it conditioning on impossibility?
How does the age of the universe compare to a trillion trillion seconds?
This question might seem flippant or irrelevant, but I think it's actually pretty important to have a good sense of scale. Like, a lot of being good at decisions hinges on using numbers to mean things, and if you don't believe numbers mean things, you're going to have some incorrect positions.
Suppose rather than facing Bomb!, God flips a coin when creating the universe. If heads, the bomb is live every time; if tails, the bomb is a dud every time. Some agent faces this problem every second for the whole duration of the universe. Could you tell which way the coin landed, just from looking at the outcomes of the decisions?
[Note that I am assuming the 1/trillion trillion chance is real, rather than it secretly being a 1/1 chance that the predictor makes an error.]
It's not impossible if it's very improbable.
Sorry, is English your native language? (The word "basically" is often used to signify rounding--there is a difference between 1 and 0.999999999999999999999999 but the difference is 0.000000000000000000000001, which is a scale which is often discarded, because many concerns will be more important and it's impractical to consider all of them.)
This is an unnecessarily douchey reply to an accurate clarification.
Wait why have you bitten that bullet? My understanding is that FDT does not imply retrocausality at all.
I think if you want to be sufficiently strict with your terminology it's not retrocausality, it's more like common causality. Like, the way 'causality' is often used by CDT proponents is more like "following the propagation of the dynamics of the universe"--I press a button, a voltage flows down a wire, then something happens. When I press the button to cooperate in the psychological twin prisoner's dilemma, there is _not_ a causal connection of this form between my cooperation and my twin's cooperation.
But there's clearly _some_ connection, which you want to have some name for. Maybe it's a logical connection, maybe it's a functional connection, whatever. When my psychological twin is earlier in time--like I'm deciding whether or not to clean dishes out of the sink--then the logical connection is similar in effect to a retrocausal connection. By cleaning the dishes today, what I actually do is make the sink clean tomorrow, not earlier this morning, but that the sink was empty this morning was because I cleaned dishes last night, because I have the disposition to clean dishes when they are in the sink, which is the common cause of cleaning dishes yesterday and today.
One of the main things that will differentiate CDT answers to decision cases and FDT answers to decision cases is that CDT will try to 'defect against itself' in this way. "Well, now that I'm in the city, I don't have to pay for hitchhiking, do I?". FDT doesn't believe that this is an option. "I'm in the city because I'm the sort of person who pays up when I'm in the city. The action that I'm taking now is necessary because it changes the probability of the event that I have already observed." The last sentence is weird! You have to bite some sort of bullet to generate that sentence. Retrocausality is probably not its true name--"dispositional thinking" or "functional thinking" are probably more accurate--but the core way that FDT is winning more points in these cases is by believing in the influence that it has in the decision problem outside of the place that CDT is looking.
[Like, when FDTers look at Bomb, they say "cool, the actual case in front of me is a trivial leaf of the overall decision problem, which barely affects the analysis. I courageously choose Left to save myself $100 almost all the time", whereas CDT says "oh man I will totally pay $100 to not die. The other impacts this decision has don't matter right now."]
Hey Vaniver! :) I like a lot of your work and online presence.
Yeah I used to bite the bullet of retrocausality too, until I realized that there's a better framing that preserves the upside and doesn't have the obvious "incorrectness" of thinking we can change the past: Just say that alternate realities "exist" in some way, and that you might be in a small pocket reality that changes the expected outcomes in the main reality, and that you also care about the you in the main reality. (I believe this is an updateless EDT / UDT thing, although I'm not sure)
Thanks! I think this are either 1) equivalent formulations, and so the question is just what seems weirder to you (which I don't expect to be objective), or 2) we should think carefully about the cases where they differ.
I think there are objective reasons to prefer the alternate reality framing - but also, I do think it changes things because it means the world is much larger, e.g. if we are in something like Tegmark IV. So e.g. ECL becomes more important, the importance of infinite ethics increases, etc.
You might say that alternate realities are also counterintuitive, but I think in some important ways they are not - I have a draft about this I could send you tomorrow, would be curious what you think.
FDT's framing is not that you retroactively change the past though. It just says that if you Left-box, your simulation also Left-boxed.
My understanding is that in the causal graph implementation in the paper, you are intervening on nodes that are allowed to be temporally prior to your own action. I'd characterize that as "changing the past".
It's more like: the past is fixed, and by making a decision, you learn your past decision as well.
(But that's not quite right, because the "you" in this story exists at both times.)
So in that framing, you just can't believe your own eyes about what your decision in the past was (in e.g. Transparent Newcomb)? How?
So why can't you believe your own eyes here?
My understanding of Transparent Newcomb is that, if you one-box, you learn that you are in fact the simulated agent.
No, that's wrong - it's correct to one-box in Transparent Newcomb even if your decision gets predicted without a high-fidelity simulation.
I fundamentally agree with the criticism about the mathematical impossible world and the vibe but I think there is an even deeper issue here.
At a fundamental level what are we even doing when we adopt a deciscion theory or say you should or should not do something? I mean in a fully literal sense you don't get to make decisions. You will always just follow the laws of physics.
What we are doing is adopting some kind of idealization about the world which -- just like when we define a game formally in mathematics -- idealizes certain things as choices while others are held fixed. And idealizing something as a choice is exactly to treat everything 'before' the choice as unable to depend on the outcome of the choice.
And once you understand things that way the whole Newcomb setup is just kinda non-sensical as regards deciscion theory. It's saying: what if you had a situation where it doesn't make sense to idealize what you are doing using a framework that treats it as a free choice how would you idealize it as a free choice.
The right answer is obviously: don't idealize it as a choice at all. Depending on how you describe the problem you can imagine idealizing in a way where the choice is what rule you precommit to or something like that but the whole debate about FDT or CDT or what have you is just fundamentally confused.
---
I mean just to illustrate how silly the argument is, what if I said the right answer to the paradox was: be someone who is physically guaranteed to take 1 box (so demon predicts you will take 1) and then take both. That does land you in a better position but it's kinda silly because I'm just breaking the rules of the game. Same with treating something both as predictable and a choice.
The rules of the deciscion theory idealization is that you have sometree and each node of the tree represents a choice with earlier ones being able to depend on later ones. If you want to look at scenarios like the demon case you need to reidealize it in a way that obeys those rules -- like a choice between rules.
I haven't read this whole thing, but someone sent me the quote involving me. FWIW:
1. I never believed in FDT, but I did think it was a fun view. At the time I wrote the paper with Nate Soares, I put most of my credence on CDT.
2. I now put most of my credence on EDT (although I've recently argued against it) because Arif Ahmed is very persuasive.
Fixed, sorry. Unrelated, have read some of your papers over the years and found them quite good.
Totally reasonable to attribute to me. I did coauthor a paper arguing for fdt. I’m just revealing what was in my heart of hearts to get you an accurate count of academics who like it. From a glance, I have a similar reaction of surprise at how many rationalists are fans or think of it as a mature theory. I am also surprised how many are Solomonoff induction stans in case you’re on the hunt for more philosophers v rationalists material.
Hmm, I don't really know much about the Solomonoff induction thing! Why do academics reject it?
You should have been at my other talk at Manifest!
First problem, which even the Solomonoff fans admit - the precise prior depends on your choice of universal algorithm. (I'll grant them their response that this only introduces an error up to a single constant, even though in this case the constant appears up front and dominates, unlike in the case of complexity theory, where it disappears in the limit.)
Second problem, which again the Solomonoff fans admit - it's uncomputable (though it is computably-approximable-from-below).
Third problem - it just builds in certainty that the truth is computable! Why think that? Especially if you think that some normatively ideal thing is uncomputable!
Fourth problem (which I think is the most important one, though lots of academic epistemologists face as well) - what motivation is there for someone who has a different prior to treat this prior as better? Omniscient priors always perform best in the worlds they are adapted to, and in general, if you throw someone into situations in proportion to a particular probability function, then them the prior that matches this probability function will be the most successful prior for them to have. If someone isn't being thrown into the world in proportion to the Solomonoff prior, why should we fetishize computational simplicity in precisely this way?
As far as I can tell, the motivation for the Solomonoff prior comes from people who are impressed by Occam's razor, but instead of trying to justify it, they reason in a Kantian transcendental way to figure out what constraint rationality would have to have to make Occam's razor automatically fall out.
Was it recorded?
I think so. Do you know anything about how/where/when these recordings might be available?
How could one not be "impressed" by Occam's Razor?
Ah sorry, heard from someone else that you adopted it. Will fix.
That makes me curious! What do you think of EDT's answer on XOR Blackmail?
> I know I’m not a simulated algorithm. The simulated algorithm isn’t conscious (we can stipulate). I am.
But you don’t know that. Especially not now that you’ve posted that!
Consider the most mundane form of simulation possible: human imagination. Suppose I set up a Newcomb experiment where I make my predictions by simply reading what the participant has written online and imagining what their thought process would be like. The prediction won’t be very accurate but it will probably be better than chance. Better than chance is all you need to create the paradox.
Well, now that I’ve read your post, my little imaginary version of you is definitely going to start with, ‘I know I’m not a simulated algorithm. The simulated algorithm isn’t conscious, I am.’ But imaginary-you is quite mistaken. Imaginary-you is not conscious. And when imaginary-you two-boxes, it costs the real you the prize. Hypothetically.
> Decision theory generally assumes that you’re self-interested. But if I’m the algorithm, then I care about algorithm me—not the version in the real world. So then I wouldn’t care about what the output of the algorithm was.
I think this is correct. If you are purely self-interested, then two-boxing is rational. However, if you are a simulation, you don’t lose anything by one-boxing because the simulation almost certainly terminates as soon as you make a choice. Therefore, you should one-box if you have any goodwill towards your real-world counterpart. Either because they are “kind of like me but different in a bunch of respects”, or just because you have a default of goodwill towards other people.
On the flipside, if you value your life as a simulated being and resent that the simulation will terminate, then you might rationally two-box as revenge. This presumably doesn’t apply to low-fidelity simulations like human imagination.
(I don’t know whether my opinion comports with FDT. As I see it, the simulation argument is an argument *against* FDT. It shows how plain old CDT can justify one-boxing.)
https://link.springer.com/article/10.1007/s11238-025-10080-w
The fundamental problem with this whole debate is that what it *means* to idealize Newcomb's problem in terms of decision theory as a choice by the predictor and then a choice by you just IS to suppose that the predictor's choice can't depend on your choice in any way.
As such there aren't different decision theories. There is one correct answer: it is be the sort of person who would be predicted to take one box then take two.
If someone objects that you broke the assumption respond that no, that's what it means to idealize something as a subsequent decision in decision theory.
If they prefer, they can choose to model the situation instead as consisting of a deciscion precommitting the agent to be a 1 or 2 boxer and then a choice by the predictor with no subsequent decision. That problem also has an obvious answer.
The one thing that doesn't make sense is to idealize what you do with the boxes as a decision and then not treat it as such. That's conflating the formal idealized notion of decision in the theory with the notion of a decision as happening any time someone goes "Hmm, what should I do".
As such there aren't different decision theories AT all. There are different attitudes about how to idealize Newcomb's problem in terms of decision theory. But once you've done that it reveals there isn't really a philosophical problem -- just a question about how to think about the scenario.
The way I put it is that the rational sort of person to be is to be a one-boxer, but the rational act is two-boxing.
I think that is a very neat way of putting the point. And -- after I wrote a much too long explanation of why the whole CDT/FDT/etc debate is confused -- I think this really gets to the heart of why that debate is confused.
The debate over which deciscion theory is correct collapses down the distinction between the questions of: should I be the sort of person who 1 boxes and if a miracle were to occur after the prediction that freed me from having to follow the laws of physics what should I want the outcome of that miracle to be.
The right thing to do is just to use which question you are trying to answer to idealize the problem in the appropriate way for that question. The apparent tension only appears bc of the false presumption that there is only one question one could be answering.
Does your verdict that "the rational act is two-boxing" actually guide the decision procedure of a rational agent?
If it does, then the predictor predicts two-boxing and the rational agent loses the million. Also, "rational sort of person" and "rational act" are the same thing after all.
If it doesn't - which I suspect - then the predictor predicts one-boxing, but... The agent is somehow built to ignore what she finds rational as a person? Isn't it impossible to build an agent that generally views one-boxing as rational, but somehow still two-boxes? In any case, it's highly suspect - paradoxical even - that the ideal agent's action directly goes against its own character.
You are assuming a lot in assuming that there even *is* a decision procedure. Most people don’t use anything like what we would call a “decision procedure” in guiding most of their behavior. The ideal agent one-boxes; the ideal act is two-boxing; the ideal agent doesn’t do the ideal act.
I think that's an implicit assumption in Newcomb's Problem, not my assumption. Either way, while humans are probably far from consistent in their decision making, there is *some* way they make decisions. That is their decision procedure.
If it’s an assumption about Newcomb’s problem then that version of Newcomb’s problem isn’t about rationality! Rationality is about doing what works, not about following procedures. (Procedures sometimes help, but that’s a contingent fact about some kinds of decision problems.)
But it seems you and I have a different idea of what a "decision procedure" is exactly. I think your view that the ideal agent one-boxes and the ideal act is to two-box is incoherent. Would love to debate you more on this
I claim that doing what works *is* following a decision procedure.
However one makes her decisions, that *is* a decision procedure. And it can be predicted, in principle.
Not sure whether he full sail adopts FDT, but Nevin Climenhaga is extremely sympathetic to it.
Aside from the nugget of stipulating non-consciousness of a simulated algorithm (which... What.)... let's abstract a bit. You are running a very weird comparison between the following three:
1. Causal Decision Theory
2. "EDT and updateless EDT" (two quite different theories; the latter is, essentially, Wei Dai's Updateless Decision Theory, or UDT, https://www.lesswrong.com/w/updateless-decision-theory, and thus _also_ partially grows from the Rationalist tradition, though Wei Dai appears to _also_ be more classically read-up!)
3. FDT
So, according to you, Rationalists are wrong because they deliver FDT; and yet the preferable alternative is UDT, which is done by a Rationalist and, historically, because of FDT-like exploration.
But it gets weirder. According to https://www.lesswrong.com/w/timeless-decision-theory, "[t]he FDT paper thus describes a general framework which remains agnostic about an updateless approach (like UDT) vs an updateful one (like TDT), but which sticks close to the logical-causality approach introduced by TDT." TDT stands for Timeless Decision Theory, Yudkowsky's _previous_ attempt at formalizing his gripes with CDT (and is itself barely formalized). If this is true, then _of course_ FDT is underdefined under the terms you offer, because it is at a different level of abstraction! It covers a _family_ of approaches!
This was a great post! Perhaps the first time you've written something that seems correct to me but wasn't something I already believed.
Problem with the bomb example:
>You face two open boxes, Left and Right, and you must take one of them. In the Left box, there is a live bomb[...] The Right box is empty, but you have to pay $100 in order to be able to take it.
>A long-dead predictor predicted whether you would choose Left or Right, *by running a simulation of you and seeing what that simulation did*. If the predictor predicted that you would choose Right, then she put a bomb in Left. If the predictor predicted that you would choose Left, then she did not put a bomb in Left, and the box is empty.
>*The predictor has a failure rate of only 1 in a trillion trillion. Helpfully, she left a note, explaining that she predicted that you would take Right, and therefore she put the bomb in Left.* [...] What box should you choose?
First problem with the scenario: as stated, the note actually contains no information about the bomb. How so? Well, the example is clear: the predictor predicted your choice *by running a simulation and seeing what the simulation did*. In order to predict you via simulating you, both you and the simulation need to have the same inputs. Thus the simulation would have to have seen an identical note to you.
The problem is that the note and its contents have to be decided *before they are shown to the simulation*. Because the Left or Right decision is not known before the simulation is run. And once the simulation is run with that particular note, then you have to see that note as well (identical inputs). So you will see the note that was decided upon before the prediction was made. Its content is thus immaterial to the presence or absence of the bomb.
If we want the example to work, it needs to be redefined first. We need an example that (a) makes the note true, (b) makes the predictor as accurate as stated, and (c) forces FDT to Left in this example. *It’s not clear that such an example exists*.
Partial example: there is a fixed-point version of this that gives (a) and (b), but it is no longer a problem for FDT. Assume the predictor ignores the simulation and always puts the bomb and the note in, and the note is known to be accurate. Then all agents will go Right, including FDT, because they believe the note. And the predictor, who predicted Right, will thus be extremely accurate.
>First problem with the scenario: as stated, the note actually contains no information about the bomb. How so? Well, the example is clear: the predictor predicted your choice *by running a simulation and seeing what the simulation did*. In order to predict you via simulating you, both you and the simulation need to have the same inputs. Thus the simulation would have to have seen an identical note to you.
Does the mechanism of prediction really need to be simulation? Are we allowed to imagine there's just some magic oracle that can accurately (but not perfectly) predict what you'd do without any simulating, so there were never any previous notes to begin with?
If not, here's another variant. Suppose time is past-eternal. Every year for all of the eternal past, the simulator has gone through the bomb experiment with a copy of you. The note in the present includes the following: "In every contiguous sliding window of yearly experiments of size 1 trillion up till now, you've chosen Right way over 99% of the time, so I've put the bomb in Left; I am programmed to only predict your behavior and fill up boxes based on the most recent such window, and no other. But since this statistical pattern of outcomes has in fact been universal across windows, the note I've left in each previous experiment has always said the same thing as what I'm telling you now." Everything else is the same as the original thought experiment. Now the simulator's notes are and have always been totally honest.
>Does the mechanism of prediction really need to be simulation? Are we allowed to imagine there's just some magic oracle that can accurately (but not perfectly) predict what you'd do without any simulating, so there were never any previous notes to begin with?
That doesn't necessarily gain you anything. If it's predicting you, it has to predict what you would do *if you saw that specific note* (otherwise it's not predicting what you will do, but what you would do in a different world).
Let's think some more. Suppose that we have two notes: “I put the bomb in Left”, “I didn't put the bomb in Left”, and the predictor runs both predictions. Now if you go Left in the first case and Right in the second case, the note is always inaccurate, so the predictor must include an inaccurate note.
Let's add another option: no note. We can imagine something like this: the predictor will run the prediction with the “bomb-note”. If that's consistent, it uses that. If not, it runs the prediction with the “no-bomb note”. If that's consistent, it uses that. If that's also inconsistent, it just predicts the “no note” and goes with that. So this predictor always reaches a result.
Then if you’re an FDT agent, you go Left, which makes the “bomb-note” inconsistent and the “no-bomb note” consistent. So you expect to see the “no-bomb note”.
Ok, I think we’ve got a working version. As an FDT agent, you see the “bomb-note” in the real world only if there’s a mistake (1 in a trillion trillion).
>If not, here's another variant. [...] Now the simulator's notes are and have always been totally honest.
That’s true, but note how it works: it essentially forces consistency via threats. Simplify to using just the previous single experiment. Then it’s saying “because you went Right last time, I put a bomb in the Left. Do you want to go Right this time?” to which the answer is obviously “yes, of course”.
>That’s true, but note how it works: it essentially forces consistency via threats. Simplify to using just the previous single experiment. Then it’s saying “because you went Right last time, I put a bomb in the Left. Do you want to go Right this time?” to which the answer is obviously “yes, of course”.
OK, but I don't understand how this addresses the worry. The simulator's choice(s) on whether to install the bomb in Left is still a function of you and a large number of perfect or near-perfect psychological duplicates' choices. If you reason as if you were a single abstract algorithm simultaneously in charge of all your symmetric duplicates' choices, you'll pick Left, and the simulator won't install the bomb. If you pick Right, then you aren't reasoning as if you were a single abstract algorithm simultaneously in charge of all your symmetric duplicates' choices. So how is that still in the vicinity of FDT?
The problem is that the predictor's belief that the FDT agent will go Right is self-confirming. If the predictor has that belief, and everyone knows it, then it puts a bomb in the Left box and the FDT agent goes Right - which makes the original belief correct.
In the simulation scenario, the FDT agent knows they can change this belief through their actions in the simulation. So then they choose Left, which corrects the predictor's belief.
However, in your scenario, the FDT agent is stuck in a poor equilibrium: it can't change the self-confirming "Right" belief through its simulated actions. So the question becomes: how did it get into that poor equilibrium?
You stipulated that the note was correct when it said "In every contiguous sliding window of yearly experiments of size 1 trillion up till now, you've chosen Right way over 99% of the time [...] I am programmed to only predict your behavior and fill up boxes based on the most recent such window, and no other."
To me, that seems like you said "let's stipulate that the FDT agent starts in a poor equilibrium".
(Note that if the agent is altruistic towards its future duplicates - and there are enough future duplicates to make it worthwhile - it will get out of the equilibrium by choosing "Left" and burning, time and time again, until the predictor gets the message and shifts to predicting Left. Consequently, if the FDT agent is altruistic towards its future duplicates, your scenario becomes impossible - it would never happen in the first place)
>However, in your scenario, the FDT agent is stuck in a poor equilibrium: it can't change the self-confirming "Right" belief through its simulated actions.
What do you mean? If (say) all the infinite prior iterations of the agent choose Left, then the simulator's predictions will always have been that the agents' choice will be Left, so the bomb will never be installed in Left, including in the present. So the prior simulated iterations do affect the simulator's prediction/decision.
>To me, that seems like you said "let's stipulate that the FDT agent starts in a poor equilibrium".
Whether this criterion is relevant to my version of the thought experiment or not, it's unclear to me how it connects to the definition of FDT, which as far as I know doesn't directly mention being in some sort of bad equilibrium (an extremely general term), but instead mentions something a bit more specific like "what would the best thing for my decision function to output, given that it's running elsewhere for a lot of other agents." Isn't it the case that if all the (past/present/future) duplicates running copies of my decision algorithm choose Left, it would be better for all of us? If the answer to that question is "yes," but then we should nevertheless choose Right in the present, then what is the actual nuts-and-bolts definition of FDT?
Working on a more substantive response, but, for the moment, note that the first counter-example is incorrect (FDT two-boxes in both cases):
>Imagine that there’s some gene that correlates 99.9% with two-boxing. The gene is not caused by two-boxing, they just perfectly correlate. Now imagine two different scenarios:
>1. The predictor looks to see if you have the gene. If you don’t, they put $3,000 in the first box. If you do, they put nothing in the first box. The second box has $1,000. Should you take both boxes?
>2. The predictor runs a simulation of you with 99.9% accuracy. The cases where the simulation is inaccurate are the same as the ones where there isn’t an overlap between your gene and which box you take. Thus, there is 100% overlap between the predictor’s judgment in this case and the last. The only difference is that in the last case, they look to see whether you have the gene, while in this case, they run a simulation of you. If they guess that you’ll one-box, they put $3,000 in box one, while if they guess you’ll two-box, they put nothing. The second box has $3,000. Should you take both boxes?
>Here FDT’s answer is that you should two-box in the first case but not in the second case.
That is incorrect; FDT two-boxes in both cases. The key is this line in 2.: **The cases where the simulation is inaccurate are the same as the ones where there isn’t an overlap between your gene and which box you take.** This gives you the power to break the predictor via your decision. If you have the gene and one-box, you break the predictor (hence it predicts you will two-box). Otherwise, you will two-box and it will predict that via simulation. The opposite happens if you don’t have the gene (it will always predict you one-box). So FDT notes that the predictor’s behaviour depends on the gene only, and thus two-boxes.
I'm not sure if this works or is even clear, but: Maybe you can solve this by distinguishing between "meta-dispositions" such as FDT and "material dispositions" that factually determine what you decide to do.
As I understand it the general idea behind FDT is something like: “Whenever you face a decision, imagine that you choose the disposition/algorithm you have, that determines what you decide for each and every action up to that point. Determine what the disposition is that maximises the combined value of all your decisions, then act according to this disposition".
If you say the optimal disposition to have is just to be an FDT-follower, that's probably an infinite regress. But if you can say in the decision process you need to determine the optimal "material disposition", that is, any disposition that factually tells you what to do (which FDT does not, I think this is also what you say?), like "do the thing that causes the most utility" or "always keep your promise to someone who saved you".
If you then want to maximise your utility, you usually get the result act according to the material disposition "do the thing that causes the most utility" (which is just CDT I think?). But if for whatever reason you believe that you can gain a lot of utility by having a different material disposition, you act out of that disposition. This is the case in Parfit's Hitchhiker, because there the driver scans the disposition you have, and only saves your life if the material disposition you have is one that causes you to pay him later.
Ok, but the disposition you have and that the driver scans is not a material disposition at all, but this weird "meta-disposition" that causes you to pick a disposition later. But, because the driver knows that when it's time to pay you will pick that material disposition that would have maximised your utility up to that point, he knows that you will pick a material dispositon that causes you to pay him. So, regress avoided?
So clearly I need to brush up on my decision theory, but let me get this right.
The issue is that 1. FDT claims you should
Make the decision that basically if infinitely repeated in all similar scenarios as the output of your decision procedure would leave you best off. 2. So the problem is that it leads to weird stuff such as
A. It matters what other algorithms relationship is to your algorithm such as the case where it you would pick the bomb because you having that as your decision output would counter factually mean there was no bomb even though there is?
Plus all sorts of similar issues where we have to basically try to explain why a scenario where ~you did things only ~you would ever do is related to this situation.
B. Anytime an algorithm is introduced you need to show what the relationship between your algorithm and this 3rd party algorithm (I know there are only 2 parties), but it seems impossible to establish anything beyond an epistemic one so you really have EDT and even if you could it would lead to weird examples like the gene one.
Do I have that right?
I'm not sure who the relevant experts are. The main use of any of these theories that I can see, admittedly from my position deep inside the rationalist community, is the kind of AI safety research that MIRI used to do, the kind that assumes you have to solve everything in one go with no feedback or we all die. And even if you think that that research direction is correct (which I am not asserting), they themselves aren't really responding to empirical evidence so I'm not sure how one would evaluate whether they are good experts.
I do think you are refusing to engage seriously with Newcomb's problem, in the way that non-philosophers often refuse to engage seriously with thought experiments. The whole point of Newcomb's problem is that Omega isn't 99.9% accurate in its modeling of you, it is 100% accurate, it cannot be wrong. And in that scenario, 1-boxing seems obviously correct to me. The usefulness of Newcomb's problem, for me, is (1) recognizing the asymptotic behavior that 1-boxing (or more generally being honest in social and business interactions) is correct where your counterparty can perfectly predict you, and (2) highlighting that other people are often pretty good at predicting us in relevant ways, and so in any particular interaction we should consider that we may be very close to that asymptote.
In your hypothetical about cutting off your leg because the prediction of that was the condition that caused you to be woken up, again, my intuition strongly says that cutting off your leg is the correct action there. I understand how an unthoughtful person could say otherwise, I understand the aversion, I'm not entirely confident I'd have the courage to go through with it, but I don't see how a thoughtful reflective person would endorse the position of not cutting off their leg. (You're calling it "for no benefit" kindof contradicts the scenario you set up.)
What originally brought me over to this FDTish way of thinking (I call it FDTish because I have not gone through the rigorous math and do not see a reason to) was the point that I don't know if I am the simulation or not. Again, this may not reflect reality perfectly, but it is a good asymptotic case to think about that for practical purposes we may be close to. I don't know how you can just assert that the simulation isn't conscious. Nobody ever specifies that in any of these hypotheticals, and it is not the natural interpretation. I don't think it is even logically possible to create a simulation that always outputs the correct behavior (meaning the behavior of the person being simulated) without the simulation being conscious. Even on your version of dualism, with its unnatural attachment between brains and minds, why would you expect it to be possible, much less a default, for a simulation that makes perfectly accurate predictions not to be conscious?
//The whole point of Newcomb's problem is that Omega isn't 99.9% accurate in its modeling of you, it is 100% accurate, it cannot be wrong.//
This is not a standard stipulation of Newcomb,. It is generally stipulated to be only 99.9% accurate.
I call it "for no benefit" because at the time you take the decision you know you don't benefit.
Maybe in practice the simulation would be conscious. But we can imagine a scenario where it isn't conscious. That's all we need for the counterexample.
"Maybe in practice the simulation would be conscious. But we can imagine a scenario where it isn't conscious. That's all we need for the counterexample."
But it isn't a counterexample! If whether or not you are conscious is a relevant input for making the decision, then the predictor just inputs the right value of this variable ("is conscious") to simulation-you. If she doesn't, then the simulation isn't accurate, and, once again, there is no subjunctive dependence.
> Maybe in practice the simulation would be conscious. But we can imagine a scenario where it isn't conscious. That's all we need for the counterexample.
I can imagine that there are three positive integers such that a^3 + b^3 = c^3. I think you need something a little stronger than what a human mind can imagine.
We normally think it's okay to, for the purpose of thought experiments, stipulate views even if you think those views are wrong. E.g. the following seems true: if deontology is true, you shouldn't kill one to save five.
Normative views, or views we think are empirical false, sure. Not views we think are logically contradictory, not unless we are trying to do a proof by contradiction. That was the point of my comparison to Fermat's last theorem.
And what’s logically contradictory about substrate dependence?
I'll also note that you are exhibiting the same flaw I see in Chalmers' zombie argument. He just expects us to accept that a p-zombie is logically possible, and he gives no good reason for us to think that, and my guess is that it isn't.
Wouldn't it be odd if consciousness somehow depended on the substrate? Why would we expect that to be the case? I don't claim to have a proof that it's not logically possible for consciousness to depend on substrate. For that we'd need a formal specification of what consciousness is, and we don't have that yet. But I'd bet that once we get there, it will turn out that consciousness can, at least in principle, be implemented on many different substrates. At the very least, if you want to use it as a premise for your arguments about decision theory, the burden is on you to show that it is possible for such an accurate simulation to not be conscious.
"I call it "for no benefit" because at the time you take the decision you know you don't benefit."
What you're missing is that your simulation "knows" this just as much as you - the predictor just makes the simulated environment like that. If not, then the predictor hasn't made a good simulation - more importantly, there is no subjunctive dependence and FDT wouldn't recommend cutting of your leg.
So your leg cutting argument is just a strawman - it doesn't attack FDT.
I talk about this in the post.
... No? You're just stipulating that the agent has relevant input the simulated agent does not - which, if true, would break subjunctive dependence. So, again - strawman.
But the algorithm can be the same as the one you use to make decisions even if it's not conscious. E.g. if the bubble sorting one was conscious but the other wasn't.
The algorithm itself can be, but you use "I am conscious" as input for your algorithm. That is, your algorithm uses an input variable called "isConscious" that is said to "true". (This is apart from the algorithm actually being conscious; it's just an input variable you use.) The predictor will just set this variable to "true" when simulating you as well. So you and your simulation still have the exact same information regarding the problem at hand, and you can't use the fact that you are conscious to know that you are not the simulated version of you. (It couldn't be otherwise! If you could there would be no subjunctive dependence.)
"But what could it possibly depend on? Isn’t it obvious that there’s no single privileged joint-carving way to decide the similarity of algorithms that doesn’t just look at the statistical correlation between their outputs? Certainly FDTers owe us some account of how this works. It doesn’t do to call it an unsolved problem, when this is the entire engine of the theory—when there’s no plausible story of what a solution would even look like, strong active reason to think there is no such solution, and a solution is needed for the theory to give any result in any case."
You're pointing at a real problem here, but are overstating your case to the extreme. Clearly, there are algorithms for which we can prove they give the same result - e.g., quick sort and bubble sort. And quick sort that has a 1 in a 100 probability of swapping 2 numbers after its normal procedure is done is very similar to bubble sort, etc.
So I think you're correct that this is an open problem, but it's not that we have absolutely no idea.
Yes you can look at whether two algorithms give the same outcome. But as I explain, that can't be what FDT cares about.
And it's not what I am saying! I know bubble sort and quick sort give the same answers WITHOUT looking at their outputs.
Those aren't the same algorithm, they just produce the same answer.
Exactly - that's the point. They are 2 different algorithms that implement the same function. So I can decide they always output the same thing.
I mean, obviously you were talking about different algorithms in your original quite I replied to?
In what sense are they the same, if you're not looking at whether they output the same thing?
There's a deeper problem here.
First, deciding whether two algorithms in fact implement the same function is one of the classic undecidable problems. (If you could solve it, then you could use it to solve the halting problem.)
But even if you know that two algorithms implement the same function, there's a lot of important uses of "same algorithm" for which you don't want to count them as the same! (If you just consider different sorting algorithms, you can see that some of them have different runtimes, and some of them are robust against different types of errors in calculation, and so on.)
But more importantly, for FDT to do reasonable things, it needs to be able to classify the outputs of one algorithm as in some sense being counterfactually dependent on the outputs of another, even when they don't have all the same outputs, and when neither calls the other as a sub-routine. This is going to be a much more difficult question to answer in plausible ways than just saying when two algorithms that implement the same function are the same or different!
It sure is undecidable, but that doesn't mean you can't ever determine it in practice. (Given your next paragraph you seem to agree with this.)
I agree things like runtime are important, but at least runtime seems straightforward enough to solve, right?
I fully agree with your last paragraph. This is a real problem; I just don't see a reason to believe it's an impossible challenge. (Not saying you claim this.)